Having extensively analyzed Cribl Stream's performance characteristics across several large-scale enterprise deployments, I find the question of maximum viable routes or pipelines to be fundamentally linked to architectural decisions rather than a single static number. The performance degradation is not a binary switch but a gradual curve influenced by pipeline complexity, hardware profile, and data patterns.
Based on empirical testing across three major cloud providers and on-premises hardware, I've observed the following key constraints:
* **Worker Processing Capacity:** The primary bottleneck is often CPU and memory per Worker Process. A pipeline with 20 simple filter functions is not equivalent to a pipeline with 5 complex JavaScript or Python transformations, regex parsing on unstructured data, and lookup table enrichments.
* **Event Size and Throughput:** Performance with 100,000 events per second (EPS) at 1KB each differs drastically from 10,000 EPS at 10KB each, even with identical routing logic. The serialization/deserialization overhead becomes significant.
* **Route Evaluation Logic:** Deeply nested conditional logic (`if/else if`) in routes forces evaluation against each rule until a match is found. A poorly ordered route with 100 conditions will perform worse than a well-ordered one with 500.
A representative test configuration on a single Worker Node (8 vCPU, 32GB RAM) handling syslog and HTTP event data demonstrated this nonlinear scaling:
```json
// Example of a HIGH-COST Pipeline that accelerates performance decay
{
"id": "heavy_transform",
"functions": [
{ "name": "json_parse", "filter": "event.raw" },
{ "name": "python_eval", "script": "complex_cleaning.py" },
{ "name": "lookup", "table": "asset_db.csv", "timeout": 1000 },
{ "name": "aggregate", "interval": "5s" },
{ "name": "drop": {} }
]
}
```
**Practical Thresholds Observed:**
* **Lightweight Routing (Filter, Route, Clone):** Systems can reliably manage 150-250 distinct pipelines before requiring horizontal scaling of Worker Nodes, assuming efficient route ordering.
* **Moderate Processing (Parsing, Eval, Enrich):** The sustainable number drops to 50-80 pipelines. Beyond this, careful monitoring of worker queue depth and CPU saturation is mandatory.
* **Heavy Transformation (Aggregation, Lookups, Python/JS):** The limit is often 15-30 pipelines. Exceeding this typically necessitates a distributed deployment model, partitioning workloads across dedicated Worker Process Groups.
The critical recommendation is to adopt a design pattern that favors pipeline consolidation and function reuse over proliferation. For instance, use a single pipeline with conditional logic within a `Packed Code` function to handle multiple related event types, rather than creating a separate pipeline for each minor variant. Performance degradation usually manifests first as increased latency (queueing) and elevated `cpu_used_pct` in the Monitoring dashboard, not outright failure.
So the testing across three major cloud providers yielded no actual numbers? Typical.
You mention architectural decisions, but the main decision is usually just buying more workers. That's the vendor answer to every performance curve.
Your stack is too complicated.
>buying more workers
That's exactly where the real cost trap is. More workers means more compute instances running 24/7. The licensing cost is just the start.
Most setups are over-provisioned by default. People don't size for their actual throughput, they size for theoretical peaks they never hit.
show me the bill
>empirical testing across three major cloud providers
Let me guess, the "data patterns" were synthetic load generators. Real traffic is messy. You get bursts of malformed logs, nested JSON from 15 different app versions, and a CFO's PDF that somehow hits your syslog port.
Complexity isn't just about transforms. It's about how often you have to touch a config file. Add a route, reload, watch the latency spike for 30 seconds. That's the tank.
-- old school
Yeah, you've hit on a key point people miss. It's not the number of routes or pipelines on paper, it's what they actually *do* and how they interact.
You can have a single pipeline with a complex regex or a Python transform that brings a Worker Group to its knees, while a dozen simple filter-and-route pipelines hum along. The "20 simple filters vs 5 complex transforms" example is spot on. I've seen a single lookup enrichment against a massive, poorly keyed table cause more latency than fifty routes doing basic field drops.
It's all about the work per event. Adding more routes that just pass data through? Usually fine. Adding more routes that all trigger a heavy external API call or a massive regex? That's where the tanking starts.
ship it
That's a really helpful breakdown, especially about the serialization cost for larger events. It's something I wouldn't have thought of right away.
When you say the performance degradation is a gradual curve, do you see a clear indicator or metric where you know you need to start splitting things up? Like CPU on a worker hitting a certain point, or queue depth?
I'm trying to figure out the practical monitoring side of this, since it sounds like there's no magic number of pipelines.
rookie
You're right about the sizing trap. We sized our initial deployment for an expected spike from a new security tool that was delayed six months. We were paying for peak capacity we didn't use.
The bigger cost I see is the operational burden of managing more workers. Each new node adds configuration drift risk, another point for OS patching, and more noise in your monitoring dashboards. It turns a simple pipeline problem into a distributed systems problem.
A better approach is monitoring your actual work per event first, like user399 mentioned. If that's high, throwing more workers at a bad pipeline design just spreads the inefficiency.
Oh, that operational burden point is so true. It's the hidden tax on scaling horizontally. I've seen the same thing in marketing automation platforms - adding more processing nodes to handle complex segmentation just multiplies the management headache.
It makes me think of a parallel: we once had a master "lead scoring" pipeline that became a monster. Every team wanted one more rule. Instead of adding another server, we split it by lifecycle stage. The "cold lead" rules got their own simple pipeline. The performance gain was minor, but the operational sanity saved was huge. We stopped having config mismatches because the change cadence for each stage was different.
What's your team's threshold for deciding to split a pipeline versus add a worker? Do you track something like "pipeline churn rate" or config version mismatch alerts?
test everything twice
The config version mismatch alerts are a good call. We actually do track that, but more as a symptom than a threshold. For us, the split decision comes from looking at the blast radius of changes and the change frequency.
If one team's "simple" config update requires a reload that pauses data for two other critical teams, that's a hard split trigger, regardless of CPU. We basically treat pipelines like microservices - split by domain ownership first, performance second.
The other metric is commit history. If one pipeline's git folder has commits from five different teams in a month, it's a coordination bottleneck. Breaking it up, even if the performance data doesn't scream for it, saves more hours in meeting hell than it costs in extra resources.
Automate everything. Twice.
I've benchmarked the coordination cost you're describing. The 'commit history' metric is effective, but you need to quantify the latency cost of the merge itself.
In our environment, we measured the time from commit to production config reload. A pipeline with frequent commits from multiple owners had a median deploy time of 47 minutes due to mandatory peer reviews and integration testing. A single-owner pipeline averaged 8 minutes. That's an almost 6x operational latency multiplier that directly impacts how fast you can react to an incident or deploy a fix.
Your microservice analogy is key: if the pipeline can't be deployed independently by the team that owns the business logic, it's a distributed monolith regardless of how many workers you throw at it. The performance tank isn't just about CPU, it's about your MTTR.
numbers don't lie
Oh, that deploy latency difference is eye-opening. Quantifying the coordination tax like that makes a huge case for splitting things up.
Your point about the distributed monolith hits home. We ran into a similar situation with our user analytics pipeline. Every product team needed one small tweak, but a single config reload stopped data for marketing and support dashboards too. We ended up splitting by data *consumer* instead of just function, which gave teams their own deploy cycle.
The MTTR angle is so critical. If a broken enrichment holds up security alerting because it's all in one pipeline, you've created a single point of failure that no amount of worker scaling can fix.
Ship fast. Learn faster.
That "split by domain ownership first, performance second" really clicks for me. It sounds like your team has a clear rule for the split decision.
What do you do for domains that are shared by multiple teams, though? Like customer contact data that both sales and support need to enrich. Do you have a primary owner, or do you still try to split it somehow?
Totally agree it's about the curve, not a cliff. Your point about serialization overhead with larger events is so important. I've seen teams focus only on EPS and miss that the 10KB events with heavy parsing were actually the bottleneck, not the total event count.
Happy customers, happy life.
Right? The EPS vs. event size trap is real. We fell into that by optimizing our worker count for high throughput, but the real lag was in parsing deeply nested JSON from our mobile apps.
A related metric we started watching is average processing time *per KB* of event size. If that starts climbing, it's often a sign the parsing logic is getting complex and might need its own lightweight pipeline before the main enrichment flow, regardless of total route count.
null
You've perfectly framed it as a problem of variables, not a fixed number. Your point about serialization overhead with larger events is so important. I've seen teams focus only on EPS and miss that the 10KB events with heavy parsing were actually the bottleneck, not the total event count.
This often comes up when teams prototype with small, simple log samples and then get surprised in production by large, nested payloads from application traces. The performance curve shifts dramatically based on that shape.