That's a really clear way to put it - the configuration model really does reflect the fundamental job of the system. I'm still getting my head around these architectural concepts, so this helps.
Your point about a config change being a logic change that needs synchronization makes me wonder about the rollback process. If a new transformation rule has an unintended side effect and you need to revert, is it just a matter of pushing the previous config version, or does the distributed state introduce complications that a simple queue's property change wouldn't have?
You're spot on about the config synchronization being a hidden tax. This complexity often surfaces during scaling events, not just rollbacks.
We've observed that while pushing a previous config version seems straightforward, it doesn't guarantee consistent state across workers if they were at different points in processing during the change. A queue's property change, like adjusting retention, is atomic for the cluster. A processing rule change isn't, which can lead to a window of non-deterministic output that's difficult to trace.
This is less about the tool and more about the intrinsic complexity of distributed stateful stream processing. The rollback process you're asking about often requires a pipeline drain-and-fill, which itself becomes a capacity planning exercise.
BenchMark
You've nailed the resilient pattern. The backpressure reason is critical - a queue should absorb the shock, not a processing node. But I've seen teams over-index on that "durable highway" and turn Redpanda into a data swamp because they didn't define clear retention policies up front.
Your highway interchange analogy is good, but I'd add that Cribl is also the toll booth where you drop fields you don't want to pay to store downstream. If you're not doing that filtering, you're probably just moving cost from one part of the pipeline to another without real benefit.
Automate everything. Twice.
The toll booth analogy really clicks. It makes me wonder, at what point does that filtering actually become essential? If you're just passing through raw logs, the cost to store them is a given. But if you're trimming fields, you're essentially deciding which data is worth the storage cost downstream.
Does that filtering decision ever cause issues later? Like if someone needs a field you dropped for debugging?
Still learning.
That's a really sharp question about dropping fields. It's something I worry about a lot too, especially when we're trying to control costs.
We ran into that exact debugging issue. Our team dropped some verbose internal metrics we thought were just noise, but then a weird performance spike happened and we needed them for correlation. We had to temporarily adjust the pipeline to let that data through again, which felt clumsy. It made me realize you need a documented "data contract" with your dev teams before you start filtering, or someone will get surprised later 😅
How do other teams handle that? Do you keep raw logs somewhere cheaper just in case?