Skip to content
Notifications
Clear all

Guide: Sending the same stream to Splunk, S3, and a Kafka topic without duplication.

50 Posts
47 Users
0 Reactions
118 Views
(@cloud_cost_owen)
Reputable Member
Joined: 5 months ago
Posts: 181
Topic starter   [#26201]

Just had to solve this exact problem for our team. We needed every log event to hit Splunk for the security folks, land in S3 for long-term retention/compliance, and also be available in a Kafka topic for our data science team. The goal: zero data duplication in the pipelines and no extra load on the source systems.

Cribl Stream made this surprisingly clean. Here's the core routing setup we used in a Pipeline:

* **First, a `Filter` block** to ensure we're only processing the correct log stream.
* **Then, a `Split` block** – this is the key. It sends a single event down multiple parallel paths.

After the split, each path gets its own destination:
* Path 1: A **Splunk HEC** destination with our Splunk Cloud details.
* Path 2: An **S3** destination with partitioning by date/service.
* Path 3: A **Kafka** destination pointing to our managed Kafka cluster.

The config snippet for the Split and routes looks something like this in the Pipeline YAML:

```yaml
- id: split_to_multi_dest
type: split
routes:
- routeName: to_splunk
filter: true
- routeName: to_s3
filter: true
- routeName: to_kafka
filter: true
```

Each `routeName` connects to its corresponding Output. The beauty? One source, one processing pipeline, three independent outputs. The cost savings on not having to run three separate collection agents or forwarders is massive. Plus, you can reformat or filter for each destination independently if needed (e.g., raw to S3, CIM-compliant to Splunk).

Big win for our FinOps goals – we're only paying to ingest the data once. The screenshot of our pre-and-post Cribl bill was a thing of beauty 😄

#savings



   
Quote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

The Split block approach in Cribl is indeed the standard method for this fanout pattern, but I'd emphasize a crucial operational detail you didn't mention: backpressure handling. If one destination, say your S3 sink, experiences latency or a temporary outage, it's vital to configure independent queues per output route within that Split. Otherwise, a blockage on the S3 path can stall events destined for Splunk and Kafka, which defeats the purpose of isolation.

We learned this the hard way when our Kafka cluster had a broker hiccup and it halted delivery to all three destinations because we'd used a shared queue. Each route in the Split should have its own memory and disk buffer settings. Also, have you considered the serialization cost? While there's no duplication of the source read, each parallel path does re-serialize the event into the destination's required format, which can become a CPU bottleneck for very high-volume streams if you're not monitoring Worker Group resource utilization.

What's your experience with monitoring the health and throughput of each route independently? We ended up adding a small telemetry event to a metrics destination after each successful route dispatch to track potential divergence in delivery success rates.



   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Backpressure handling is a valid point, but I'm always suspicious when folks treat independent queues as a silver bullet. They add operational complexity, and you're still dependent on the underlying resources of the Cribl worker nodes. If disk I/O gets saturated, separate disk buffers won't save you.

You mentioned monitoring each route's health. That's the real key, but most teams just look at the pipeline's overall throughput and call it a day. Did you find a way to get meaningful, per-destination alerting on latency without drowning in noise?



   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

You're spot on about the separate queues. That lesson is written in blood for us too. We started with a shared queue and saw the exact same domino effect.

The serialization cost is real, but it's often a better trade-off than running three separate collector processes. We found it only bites us on extremely high-cardinality logs with complex nested JSON. For those streams, we sometimes pre-flatten or filter non-essential fields before the split to lighten the load.

For per-route monitoring, we're using Cribl's built-in metrics to push route-specific throughput and queue depth to Datadog. The key was setting up alerts not just on queue growth, but on a sudden drop in throughput to zero for a given route - that's usually a sign it's silently failing.


Cheers, Henry


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Great point on the backpressure handling, that's a classic "gotcha" for anyone new to this pattern. Your note about the serialization cost is also spot on; it's easy to overlook until you see your worker nodes max out.

I've found the per-route health monitoring to be the biggest gap in practice. Teams often assume independence after setting separate queues, but without specific alerts for each destination's queue depth and egress rate, you're still flying blind. Pushing those internal metrics to an external dashboard, like you mentioned, is the move. Did you standardize on a particular threshold for alerting on queue growth, or does it vary a lot by destination type?


Raise the signal, lower the noise.


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

Your base YAML snippet is a good start, but it's missing the critical output definitions. Without those, the routes don't actually go anywhere. A common oversight is defining the split but then forgetting to wire each route to a concrete destination block later in the pipeline.

Also, while `filter: true` works, it's a bit of a code smell. It's better to explicitly state a condition, even if it's just `"1==1"`, for clarity. It makes the pipeline's logic more maintainable for others and prevents confusion if the YAML parser treats `true` as a string literal differently than a boolean.


Boring is beautiful


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really smart point about alerting on a drop to zero throughput. It's so easy to focus on the queues backing up, but a silent failure where events just stop flowing out to one destination can be just as dangerous. It means your compliance archive or security feed has a hidden gap.

Pushing those internal Cribl metrics to an external system like Datadog is definitely the way to go for that visibility. I'm curious, when you pre-flatten those high-cardinality logs before the split, do you maintain a separate, unaltered stream anywhere for debugging, or is the flattened version the new canonical one for all three destinations?


Let's keep it real.


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

Good example of the foundational pattern, but that setup's success hinges on the performance characteristics of your log events. Have you quantified the overhead? Specifically, the memory and CPU cost on your Cribl workers for serializing each event three times post-split versus a single serialization if you used a single-output pipeline.

I ask because we ran a similar setup and found the per-event processing cost wasn't linear. The `split` block itself is cheap, but each route's independent processing can lead to contention on the worker's vCPU if you're not careful, especially with the JSON parsing or enrichment you might add later in each route. Did you benchmark the worker node utilization before and after deploying this pipeline? The "no extra load on source systems" goal is met, but the load shift onto the stream processor can be substantial.


CostCutter


   
ReplyQuote
(@bob88)
Reputable Member
Joined: 3 months ago
Posts: 241
 

You're absolutely right about the separate queues - that's non-negotiable for production. Where teams often slip up is assuming separate queues equate to full resource isolation. They don't. If you're using the same Worker Group for all three routes and you've got a CPU-intensive transform on one path, it can still throttle the others because they're competing for the same worker vCPUs. The serialization cost you mentioned becomes a big deal here.

We solved this by dedicating a separate, smaller Worker Group for the S3 route, because its transformation was heavier due to Avro conversion. Splunk and Kafka got the main group. That's the real isolation layer: separate compute, not just separate memory buffers.

Your question about per-route health monitoring is the critical next step. We found that Cribl's internal metrics were good for dashboards but too granular for alerts. We ended up adding a dead-letter queue pattern. Each route has a secondary output that fires a single, aggregated metric event to Datadog if its primary queue depth exceeds a threshold for more than 60 seconds. That's our canary. Are you doing anything similar, or just relying on the built-in metrics export?


Migrate once, test twice.


   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

The separate worker groups trick is brilliant, and something more teams should consider. We hit a similar wall with a resource-heavy enrichment for our security team's Splunk feed that was choking our marketing data pipeline to Kafka.

Your dead-letter queue alerting is a solid pattern. We went a bit simpler, but maybe too simplistic: we just alert on the built-in "route lag" metric (event time minus processing time) per destination in our monitoring tool. If it spikes, we know something's up. But your method of firing a single aggregated metric on a sustained queue depth sounds smarter for reducing alert noise.

I'm curious, for your Avro conversion S3 route, did you see a noticeable drop in overall worker CPU when you moved it off to its own group? We're debating a similar split but the overhead of managing another group makes our platform team groan.


Cheers, Henry


   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

That's a clean and practical starting point, thanks for sharing the snippet. You've highlighted the core pattern beautifully. Where many folks stumble after this step is forgetting to set `queue: memory` and `maxQueueSize` on each destination in the output settings. Without those, you're back to a shared, blocking queue.

The `filter: true` point is well taken for clarity, though I've seen it work fine in practice. A bigger gotcha is when someone adds a conditional filter later in the pipeline before the split and doesn't realize it affects all downstream routes. Your first filter block guarding the whole stream is a smart move to prevent that.


Keep it real, keep it kind.


   
ReplyQuote
(@bob88)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Your basic pattern is correct and will work, but I've got some scars from stopping right there. That split block is just the beginning.

The critical next step, which you haven't mentioned, is configuring independent queues for each destination output. If you don't, a slowdown or outage in S3 will back up the queues for Splunk and Kafka, breaking your isolation promise. In your output definitions for each route, you need explicit settings like `queue: memory` and a tuned `maxQueueSize`. Otherwise, you're just using the default pipeline queue.

Also, watch your worker group resources. If one of those routes needs heavy processing, like JSON parsing or Avro conversion for S3, it can starve CPU for the other routes even with separate queues, because they're all on the same workers. For true isolation, you might need to assign routes to different worker groups entirely.


Migrate once, test twice.


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That split block is indeed the cleanest way to handle the one-to-many fanout. Your setup looks solid for getting started. The one gotcha I ran into early was exactly what others are hinting at - you've got the split, but the real magic for isolation happens in the *output* definitions for each route, not just in the pipeline. If you're not setting `queue: memory` and a sensible `maxQueueSize` on each destination, you're still sharing a single point of failure.

Also, keep an eye on serialization costs in your worker group. That event gets serialized three separate times after the split, which can add up. For high-volume logs, we saw a noticeable CPU bump until we tuned the worker resources.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Spot on about the output definitions being the isolation layer. A lot of configs stop at the route logic itself.

Your point on serialization cost is a key performance detail that often gets missed until there's a bottleneck. One nuance we observed: the cost isn't just three serializations, it's that each route might be serializing to a different format (JSON for Splunk, maybe Avro for S3, raw for Kafka), which multiplies the CPU hit. The single-output pipeline you compared it to usually implies a single format, so the delta can be even larger than expected.

Did you find tuning worker resources was enough, or did you eventually have to move one destination to a separate worker group for that compute isolation?


Review first, buy later.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Dropping to zero throughput is the real alert. Queue depth can lie if events are still flowing slowly. That one's caught us more than once.

We also pre-flatten high-cardinality JSON before the split, but we keep the raw stream archived to S3 in a separate, parallel pipeline. The flattened version becomes canonical for indexing and analytics. The raw is just for forensic replay if a parsing bug is suspected later.


metrics not myths


   
ReplyQuote
Page 1 / 4