Good breakdown, and that "pain upfront" framing hits home. It feels like the right call for control, but man, starting from scratch on a parser engine is daunting.
I'm curious about the tools for Option 1. You mentioned Golang or a stream processor. For a small team like ours, is something like Fluentd with custom plugins a practical middle ground, or does that just create a different kind of maintenance monster later?
Fluentd's definitely a viable middle ground, especially for a small team. We used it as an initial ingestion buffer before our custom parsers. It can handle a lot of format weirdness with its built-in parsers and filters.
But you've hit the nail on the head about a different maintenance monster. Writing and versioning custom Ruby plugins for every new log source became its own chore, and debugging a misbehaving plugin in a chain felt opaque. For us, that's when we switched to writing small, focused parsers in Python (since that's our main stack) and deployed them as sidecars. That gave us better isolation and our existing tooling for tests and metrics.
So I'd say use Fluentd for the initial collection and routing, but push the actual parsing logic to discrete services you can control. It splits the difference.
Totally agree on the sidecar pattern for parsing. We landed there too after fighting a giant "log transformer" monolith. The key for us was using the sidecar just for parsing and mapping, then shipping the clean UDM to a central queue. Keeps the cognitive load low - each sidecar's job is just "make these weird logs normal."
But a small warning from our scars: watch out for sidecar sprawl. Without discipline, you'll end up with 30 slightly different Python 3.7 images that all need security patches. We eventually built a base parser image with the shared libs and just had each acquisition's Dockerfile add its specific mapping file. Saved our ops team a lot of headache.
Also, Fluentd as the dumb router is solid advice. It's terrible at complex logic but great at "get this log from here to there."
Oh, the sidecar sprawl warning is so real, thanks for that. The base parser image idea is a lifesaver I hadn't considered.
How do you handle the versioning for that base image? Do you just push updates and have all the acquisition sidecars restart on a schedule, or is there a more controlled rollout to make sure a new shared library doesn't break one of the older parsers?
One step at a time
Absolutely, making log format changes a gating item for your SLA is the key move. We learned that the hard way too.
We tried the parser metadata approach first, but contacts changed roles and internal changelogs got neglected. The SLA requirement was the only thing that worked, because it tied their feature releases directly to our pipeline's health. It turns "your logs broke our ingestion" into a joint problem to solve before launch, not an outage afterward.
One caveat, though: this only works if your team has enough clout to enforce it. In our last org, the acquisition's product team was seen as the "new revenue," and our platform team couldn't push back. We ended up with the contractual lever but no real authority to pull it. Did you run into any political hurdles getting that agreement in place?
Data is sacred.