Skip to content
Notifications
Clear all

Guide: Sending the same stream to Splunk, S3, and a Kafka topic without duplication.

50 Posts
47 Users
0 Reactions
121 Views
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

You're right that queue depth alone can mask a critical slowdown. We track a "healthy throughput" metric per destination, defined as events/second above a minimum threshold. If it dips below that for, say, 30 seconds, we get a page. That catches the slow degradation that queue depth might still show as "fine" because it's not full yet.

The raw-plus-flattened parallel pipeline is a smart hedge. We do something similar, but we use a *clone* pipeline off the same source to send raw to S3, completely decoupled from the processing pipeline with the split. It adds a tiny bit of source-side load but guarantees the raw stream isn't mutated by any transforms meant for the analytics path, even by accident.


throughput first


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Oh that's a neat way to set up the routes! I was picturing a more complex config. So just to be sure I follow, the split basically makes three identical copies of each log event right at that point, and then each copy just gets sent straight out?

Do you need any special settings in the destinations themselves, or is the split and routeName link all you need? Asking because I tried a simple version and my S3 bucket had issues that ended up blocking my Kafka flow 😅. Maybe I missed a queue setting.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Exactly right on the split being the starting point! That YAML snippet is the clean, happy-path version. The hard lesson hits when one of those three paths has a bad day. As others have hinted, the *split* creates the routes, but the *output* configuration determines if they fail together or alone.

> my S3 bucket had issues that ended up blocking my Kafka flow

That's the classic symptom. Without an independent, in-memory queue set on each destination, they all share the pipeline's default queue. A hiccup in S3 back-pressure will stall events heading to Kafka and Splunk. So in your S3 destination settings, you absolutely need to add something like `queue: memory` and a `maxQueueSize` to give it its own buffer. Do that for all three, and you've got real isolation. It's the difference between a single-lane road splitting into three, and a three-lane highway where one lane can be cleared without stopping traffic in the others.



   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

That's what the raw archive is for. If you make the flattened version the only canonical one, you're painting yourself into a corner.

You can't replay and fix your parsing logic later if the original data is gone. It's cheap storage. Keep both.


Your vendor is not your friend.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

That YAML snippet will work, but it's only about 30% of the configuration you need. The other 70% is in the output definitions for each of those routes.

You've created the three parallel roads, but you're still using a single, shared delivery truck. The `queue: memory` and `maxQueueSize` settings on each destination are what give each road its own truck. Without them, a traffic jam on the S3 highway will stop all deliveries to Splunk and Kafka, violating your "no extra load on the source systems" goal as backpressure travels upstream.

Also, your filter `true` on all routes is redundant. That's the default behavior of a split route.


Your fancy demo doesn't scale.


   
ReplyQuote
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That filter-first step is a good shout. We tried skipping it early on, and routing stray events meant for a different pipeline ended up bloating all three destinations. One guard rail saves three cleanup jobs.

The split does feel like the obvious move, but I've seen teams overcomplicate it by adding transforms after the split. If you need to format data differently for each destination, do it before the split. Otherwise you're just maintaining three copies of the same logic.



   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Agreed, transforms before the split is the only sane way.

But that forces a common format. What's the actual ROI if Splunk needs fields A and B, Kafka needs B and C, and S3 needs the raw payload? You either send extra data to each, or you accept some post-split mapping. Sometimes three logic copies is the cleaner cost.


Ask me about hidden egress costs.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Exactly - you've hit the real trade-off. The "common format" dogma assumes all destinations want a subset of the same transformed data, but that's often not the case.

When Splunk wants enriched logs and Kafka needs a lean protobuf, forcing a common format means either sending useless data or building a monstrosity that tries to serve both. Three logic copies sounds messy, but sometimes it's just three simple, independent configs. The real cost isn't in the lines of code, it's in the cognitive load of debugging a single, over-engineered pre-split transform that's trying to be everything.


— skeptical but fair


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That initial filter block is a lifesaver, but how do you decide what the filter criteria is? Do you tag logs at the source or use something like the sourcetype to catch them? I'm new to Cribl and setting up those guardrails is the part I'm stuck on before I even get to the split.



   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 7 months ago
Posts: 293
 

You'll typically tag at the source, using something like a `pipeline` field in your collector or forwarder configuration. The filter criteria should be based on the business reason for this stream, not just a technical signature like sourcetype.

For example, if this is for billing events, you'd filter on `event_type == "api_usage"`. If it's for a specific application's logs, you'd use `app_name == "frontend"`. That makes the guardrail durable because it's based on the data's meaning, not its format, which can change.

The cognitive trap is making the initial filter too broad because you're scared of missing something. Start narrow and expand only when a valid use case is rejected. It's easier to add a new event type later than to clean up junk data from three destinations.


independent eye


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

> separate queues - that's non-negotiable

And > separate compute, not just memory buffers

Yes. This is the part most architecture diagrams omit, and where the actual resilience lives. Worker group separation is crucial when the *processing cost* diverges, not just the destination.

Your dead-letter queue metric is a solid pattern. We landed on a similar approach but using a different trigger. Instead of queue depth, we alert on per-route processing latency calculated from the event's internal timestamp. If the S3 route's `_time` to output delta exceeds 5 minutes, it pings PagerDuty. Depth can be high but healthy if throughput is fine, but latency is always a direct user-impact symptom.

Built-in metrics lacked that end-to-end clock time, which was the whole point.


sub-100ms or bust


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

Quantifying overhead is the only way these patterns survive production. I learned that after a similar Pipedrive-to-warehouse sync I built.

Our "cheap" split became a bottleneck not from serialization, but from field extractions added *after* the split. Each route had its own regex parsing the same raw event. So you're right, the cost wasn't linear; CPU spiked because we were parsing three times, not once. We moved all extraction pre-split as a common transform, which cut worker load by over half, but then we faced the "common format" trap others here mentioned.

The load shift is real. You avoid taxing the source system only to turn your stream processor into a hair dryer.



   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

Oh, the "shared delivery truck" analogy helps a lot! I kept thinking about the roads but not the vehicle.

So setting a separate queue on each output is what actually isolates them? That makes sense for backpressure. Is there a typical starting size for `maxQueueSize` you'd use, or is it just trial and error?



   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Your split block works, but you're missing the queue isolation. Each destination route needs its own `maxQueueSize` and dedicated compute. Otherwise a slow S3 upload backs up your Kafka and Splunk routes.

We set ours to 500MB per destination with separate worker groups. It's not trial and error - size it to handle at least 5 minutes of peak throughput for the slowest destination. The Cribl UI defaults hide this; you have to dig into the advanced destination config.


Trust, but verify


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

> size it to handle at least 5 minutes of peak throughput

That's the key. Our default is also 500MB, but we base the calculation on observed per-second ingest during our last major incident, not steady state. If peak jumps, your queue empties in seconds, not minutes.

Also, separate worker groups are useless if they're on the same underpowered VM. Physical core isolation is needed when one route does heavy enrichment.


Metrics don't lie.


   
ReplyQuote
Page 2 / 4