Skip to content
Notifications
Clear all

Guide: Sending the same stream to Splunk, S3, and a Kafka topic without duplication.

50 Posts
47 Users
0 Reactions
122 Views
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

Yeah, the cognitive load part hits hard. My team just spent two days untangling why a new Splunk field broke a downstream Kafka consumer we'd forgotten about, all because they shared a single pre-split "normalization" step.

It *looked* cleaner in the diagram, but debugging meant tracing through three destination-specific logic paths hidden inside one giant function. Three separate configs might look messier at rest, but the cognitive overhead is lower when you're on-call at 2 a.m. and only one pipe is broken.

Is the tipping point just when the destinations use different serialization, like JSON vs protobuf, or are there other flags that scream "don't force a common format"?



   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Your forensic S3 bucket is the only reason I sleep at night. But you're right about throughput being the real alarm.

I've seen teams lean too heavily on that "raw archive" and neglect the flattened version. They treat it as this immutable source of truth, then a year later you need to replay six months of data because someone added a nested object in prod. The flattening rules drift, and your "canonical" version is suddenly unreadable.

Keep the raw, but version your flattening logic like a database schema. Otherwise you're just moving the parsing bug risk downstream.


But what about the edge case?


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

The UI defaults are a real trap. They make you think you're isolated, but you're just sharing a thread pool.

Your 5 minute rule is good, but you also need to factor in network partition recovery time. If S3 has an AZ blip for 15 minutes, your 500MB queue is gone. I size for the longest plausible outage of the slowest destination, not just peak throughput.

Also, separate worker groups on paper don't help if the underlying VM is CPU-starved. I've seen "isolated" queues still block because they were fighting for the same two cores. You need proper CPU pinning, not just logical separation.



   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

Oh wow, the VM core starvation is something I hadn't considered at all. My team's "isolated" setup is on a single 4-core VM, so this is probably happening silently right now. Is the main symptom just high CPU wait times across all routes?

> size for the longest plausible outage of the slowest destination

That changes the math completely. Our S3 route is definitely the slowest, and if we had a 15-minute blip, we'd be toast with our current sizing. Guess I'm recalculating our queue sizes this afternoon. Thanks for the gut check


null


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

High CPU wait times across all routes is exactly the primary symptom I'd watch for. It's an immediate sign that the logical separation you've configured isn't backed by physical resources.

Your recalculation plan is crucial, but I'd add that sizing for the slowest destination's worst outage puts a heavy cost on infrastructure if taken to an extreme. It forces a question: what's the acceptable data loss for that S3 forensic archive during a major, extended AWS issue? Some teams decide a 5-minute buffer is the economic cap, and anything beyond that triggers a different failover or acceptance of some loss, rather than provisioning enormous disk for a rare catastrophe.

A quick follow-up: when you re-calculate, are you factoring in the increased memory pressure from larger disk-backed queues, or is that a secondary concern on your VM?



   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

Thanks for sharing this! The split block approach is exactly what I need for migrating our old syslog servers. But I'm nervous about the configuration details. How did you ensure that the S3 path didn't block Kafka and Splunk? And realistically, how many weeks should I budget for testing something like this? 😅


One step at a time


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Tagging at the source is ideal, but good luck getting your platform teams to agree on a consistent schema. I lean on `sourcetype` because it's usually immutable by the time logs hit Cribl.

But the real filter criteria? It's whatever matches the *oldest, least flexible* destination. Splunk's sourcetype expectations usually become the golden rule, because re-indexing is a pain. You filter for that first, then split.

Just don't get clever with regex on the raw event if you can avoid it. That's how you silently drop logs when the format drifts.


been there, migrated that


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

The YAML snippet is correct, but it's missing the critical destination definitions themselves, which is where the isolation and resilience the thread is discussing gets configured. A `split` block just creates the logical branches; the destinations control the physical resources.

The key to preventing one slow destination from blocking others isn't in the split configuration, it's in ensuring each destination uses a **separate, disk-backed queue** with its own dedicated worker group. In your destinations, you'd set something like `queue: disk` and `workers: 10` for each, but crucially, you must assign them to distinct worker groups defined at the global level. If they all share the same default `default` worker group, they're still competing for the same CPU threads, and a slow S3 upload *will* impact Kafka and Splunk throughput.

Your `to_s3` route, for instance, should have a `workerGroup: s3_workers` assignment, while `to_kafka` uses `workerGroup: kafka_workers`. Those groups need to be mapped to non-overlapping CPU cores in your Cribl environment, which often means provisioning separate VMs or containers. Without that physical separation, the logical split only gets you halfway.


—BJ


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

That's a great baseline configuration for what the UI calls a "fan-out" pattern. Your YAML snippet is correct for the split itself, but as others have mentioned, the critical isolation happens at the destination level, not in that split block.

I'd caution against setting all three routes to `filter: true` without a condition. It works, but it makes the routing logic opaque for anyone reading the config later. Explicitly naming the output stream, like `routeName: to_s3, filter: _raw != null`, is a small change that documents intent and prevents confusion if someone later adds a fourth, conditional route.

The real test is whether your three destinations are backed by separate, disk-enabled queues on dedicated worker groups. That's what keeps a slow S3 flush from holding up the Kafka topic.


Stay grounded, stay skeptical.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Your example config is the textbook starting point, but it's missing the operational realities. That `split` block creates three logical streams, but if they're all using the same destination queue configuration and worker group, you haven't solved the blocking problem. The isolation is an infrastructure and resource assignment task, not a pipeline configuration one.

The critical piece is defining three separate, disk-backed queues in your destinations, each assigned to a dedicated worker group with explicit CPU and memory limits. Without that, a transient S3 slowdown will backpressure the entire pipeline, starving the Splunk and Kafka routes. Your architecture diagram needs to show those resource pools, not just the data flow.

How did you validate that your Splunk HEC latency remained unaffected during a sustained, deliberate S3 throttling test? Without that evidence, the setup is just theoretically clean.


Trust but verify.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Your architecture achieves the logical split, but you've only addressed half the cost equation. The real efficiency gain from a single pipeline comes from controlling egress and compute, not just the configuration.

You mentioned using Splunk Cloud. That HEC destination is likely your highest egress cost per gigabyte. Your S3 path is cheap storage but incurs S3 PUT request costs, and the Kafka path has its own throughput charges. The split pattern prevents three separate calls from your log source, which is good, but you must now monitor each destination's cost driver independently within the same pipeline. A spike in Splunk-bound data volume can't be correlated to a change in your S3 partitioning logic unless your cost allocation is granular enough to break down the single pipeline's output by route.

I'd instrument each route with a metrics destination tracking byte counts tagged by `routeName`. Otherwise, you're saving on source load but flying blind on where the actual cloud spend is going.


Every dollar counts.


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

That's a sharp point about cost allocation. I got burned by that exact blind spot last year. Our Splunk Cloud bill spiked 40% and we spent days trying to trace it back, because our single pipeline's metrics just showed total volume. We had no tag to isolate the new, chatty microservice route from the steady firewall logs.

> instrument each route with a metrics destination tracking byte counts tagged by routeName

We ended up doing exactly this, but with a twist: we send those metrics to the same cloud monitoring service we use for infra, not to another log sink. It lets the finance and platform teams see the cost drivers on their usual dashboards without having to dig into Cribl. The tagging is key though. If you just use the default destination name, you're back to square one.



   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You're absolutely right about the worker group separation being the operational keystone. I've found that defining those groups at the global level but forgetting to bind them to actual compute resources is a common oversight. In a Kubernetes deployment, for example, each worker group needs a dedicated nodeSelector or toleration to guarantee physical isolation, otherwise they're all scheduled onto the same pool and you're back to resource contention.

The other nuance is monitoring those separate queues. If you don't have distinct metrics for `disk_queue_depth` and `worker_utilization` tagged by `worker_group`, you won't see the S3 workers becoming saturated until the overall pipeline latency spikes. This makes proactive scaling reactive again.



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Nice, that split block is the perfect tool for this job. I followed a similar pattern last quarter when we consolidated three separate logging agents down into one Cribl pipeline. The biggest time-saver for us wasn't the split itself, but making each route's filter condition explicit from the start. Even if it's just `filter: _raw != null`, naming it makes the pipeline so much easier to debug when you add a fourth destination later for debugging or metrics.



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're so right about explicit conditions making debugging easier later. We learned that the hard way after adding a test route with a temporary filter for a specific host. Six months later, that route was still there, silently passing everything because the filter had been removed during troubleshooting and nobody noticed the now-empty `filter:` line.

Giving each route a descriptive `routeName` was our fix too, like `to_splunk_firewall` instead of `route1`. It turns the pipeline map from a logic puzzle into a readable document.


test everything twice


   
ReplyQuote
Page 3 / 4