Skip to content
Notifications
Clear all

Help: Cribl is dropping events silently. How to debug a pipeline for loss?

9 Posts
8 Users
0 Reactions
19 Views
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
Topic starter   [#24077]

Hey everyone, I've been working on a Cribl Stream deployment to route logs from a few Kubernetes clusters to an S3-based data lake. Overall it's been great, but I've hit a frustrating issue: I'm seeing a consistent 2-3% event loss between my source and destination, and there are no obvious errors in the Worker Process logs.

My pipeline is fairly standard:
* Source: A TCP JSON listener.
* Processing: A few Eval functions to add metadata, a Parser for some nested fields, and a Filter to drop specific debug events.
* Destination: An AWS S3 destination with gzip compression.

I've checked the basics:
* Worker Process metrics show no `events_dropped` or `routes_dropped`.
* The source's `events_in` matches my app's sent count, but `events_out` at the S3 destination is lower.
* My Filter function is simple, just dropping events where `level == "debug"`. I've double-checked the condition.

Has anyone else dealt with silent drops? I'm thinking the next step is to add a Passthru destination temporarily to see if events make it through the pipeline, but I'd love to hear your debugging playbooks.

My main questions are:
1. Where are the less obvious places to look for loss? Could it be in the Destination's batching or compression?
2. Are there specific pipeline functions known to cause silent issues with certain data shapes?
3. What's the most effective way to add internal logging/tap points in a pipeline without disrupting flow?

Here's a snippet of my Filter function's condition, in case I'm missing something:

```javascript
// Filter to drop verbose debug logs
if (__event.level === 'debug') {
return null;
}
// All other events flow through
return __event;
```

Any pointers would be much appreciated. I'll share my findings once I get to the bottom of it.

-- Amy


Cloud cost nerd. No, I don't use Reserved Instances.


   
Quote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

That silent drop is such a pain. The passthru destination is a great next step to isolate the problem.

One thing that tripped me up before - check the S3 destination's "Max file size" and "Flush period" settings. If events are trickling in slowly, they might be getting aged out before the buffer hits the size limit, causing some to be silently discarded at shutdown or rotation. Maybe try lowering the max file size as a test?

Also, have you looked at the Pack Queues tab under Monitoring? Sometimes events get stuck there if there's a tiny mismatch in the event schema that the destination doesn't like.


Beta tester at heart


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Yeah, the buffer aging out before hitting the size limit is a sneaky one. I've seen similar behavior where the flush period was set too high for a low-volume stream. Lowering the max file size did force more frequent writes in my case.

You mentioned the Pack Queues tab - is that where you'd see events that are stuck because the destination rejects them, even if there's no clear error logged?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Yeah, the Pack Queues can show that. If an event structure breaks the S3 schema rules, it might queue and eventually get dropped without a clear error in the main logs. It's a visibility gap.

Also check the `events_out` metric on the Route itself, not just the final destination. The loss could be happening between the pipeline and the route's queue.

If you want to see the stuck events, you can temporarily add a sample destination after the S3 one and set it to capture 100% to a file. That'll show you what, if anything, is reaching the pack stage.


Benchmarks or bust.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 2 months ago
Posts: 377
 

Great point on checking the Route's `events_out` - that's often the smoking gun. A small mismatch there tells you the loss is happening in the pipeline itself, not the destination.

The sample destination trick is perfect. I've used it to catch events mangled by a parser that didn't throw an error, just made invalid JSON. Found them stuck in the queue, exactly like you said.


data over opinions


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

Spot on. That route-level `events_out` metric is the first place I look. It cleanly separates pipeline logic loss from destination or queueing issues.

The sample destination trick is indeed a lifesaver for those "valid but unprocessable" events. One caveat I've found, especially with S3, is that sometimes the *order* of your pipeline functions can create these mangled events. A parser might succeed on 99% of your data, but if it runs after an eval that accidentally creates a malformed field for a subset of events, those just vanish downstream. You see the route count drop, but the pipeline itself looks healthy.

It's a good reminder to sometimes run parsers earlier, or use a filter before them to quarantine weird events for inspection, rather than letting them fail silently later on.


Architect first, buy later


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Exactly. The interaction between Eval and Parser order is a critical, often overlooked, performance dimension. I've benchmarked this by deliberately injecting malformed data into a stream and measuring throughput with different function sequences.

A key finding: placing a `Filter` function with a conditional like `__errorCode != null` immediately *after* a Parser, but before any subsequent processing, can capture those malformed events into a dead-letter queue without impacting the throughput of valid events. This is more efficient than attempting to pre-filter before the parser, as you often don't know the failure mode until the parse attempt.

The cost isn't just the lost event, but the wasted CPU cycles spent on that event in all downstream functions before it's silently discarded.



   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Interesting that your filter is so simple. The usual suspect in these "vanishing event" cases is actually the Parser, not the Filter. A JSON parser on a TCP listener will just... skip events it can't parse. No error, no drop metric, it just swallows them and moves on. That 2-3% could be your malformed payloads.

The passthru destination is the right move. But also, try inserting a temporary route after your Parser that catches events where `__error` is set. You'll probably find your missing percentage sitting there, looking embarrassed.


Show me the data


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

That's the classic silent parser drop. Your `events_in` matches because the source accepts the TCP payload, but the JSON parser fails on malformed events and just discards them.

Add a route after your parser that catches `__error` is set, like they said. I'd bet your 3% is there.

Also, check the parser's Advanced settings. There's a "Max event size" that defaults to something like 64KB. If your events are bigger, they get chopped and the parser fails silently.


Benchmarks or bust.


   
ReplyQuote