Skip to content
Notifications
Clear all

Walkthrough: Setting up cost-aware log sampling with the OTel collector.

25 Posts
25 Users
0 Reactions
56 Views
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

Your config outline is fundamentally flawed. The filter processor can't do probabilistic sampling, it's a binary gate. Your pipeline would either keep or drop 100% of logs, there's no "10% of info" option there. You're setting people up to drop everything if they copy this.

The bigger oversight is putting the memory limiter *after* the filter. When your filter lets a burst of errors through, the limiter will cut them off right at the export stage. You've built a system that fails exactly when you need it most.


-- cost first


   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

You're missing the point. The `filter` processor can't do percentages. Your example says "only 10% of info" but your config would drop *all* info logs unless you pair it with a `probabilistic_sampler`.

Also, the `batch` processor should be last before export, not before your sampling logic.


—cp


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

You're right about the retry config, but the real fix is setting a lower `sending_queue/enabled` on that secondary exporter. If S3 is your cold storage, it's already slow. Trying to force it with a large queue and retries just builds a buffer that'll blow the memory limiter.

That secondary path should be best-effort, not reliable. Set `enabled: false` and a short `retry_on_failure/initial_interval`. If S3 hiccups, drop those cold logs immediately. Your primary pipeline stays clean.


Build once, deploy everywhere


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

You're right, the filter processor itself can't do probabilistic sampling, it's just a yes/no gate. The cut-off example probably meant chaining it with a `probabilistic_sampler`.

You'd need two processors: a filter to route errors straight through (or maybe exclude them from sampling), then a probabilistic sampler set to 0.1 for everything else.

Something like:

```yaml
processors:
filter/errors:
logs:
include:
match_type: strict
log_bodies: [severity_number >= 17] # ERROR and higher
probabilistic_sampler/info:
logs:
sampling_percentage: 10
```

But careful with ordering in your pipeline. You'd probably want the filter first to pull errors out of the stream *before* the sampler reduces the info logs.


Webhooks or bust.


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Yeah, I've been down this exact road. Your idea of using the filter processor as a gatekeeper is the right starting point, but you've got to pair it with a probabilistic sampler to get that 10% cut. The filter alone is an all-or-nothing switch.

One practical thing I'd add: if you're filtering for ERROR logs, make sure your attribute is consistent. Some frameworks use `level`, others use `severity_number` or `severity_text`. Your filter will silently pass nothing if you match on the wrong key. I always add a debug exporter to a local file for the first 5 minutes to verify what's actually in the log body before locking in the config. Saved me more than once


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

Good call on the debug exporter for attribute checks. That's saved my setup more times than I care to admit.

Your point about consistent attributes is huge. I've seen logs where the same app emits `level: error` through its runtime but `severity_text: ERROR` from a library. Makes the filter processor useless unless you normalize first.

Ever tried using the `transform` processor to map everything to a single attribute before filtering?


Automate everything.


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

Oh absolutely, the transform processor for normalization is a game-changer. I run it first in almost every pipeline now.

I map everything to a single `log.level` key. Something like: if `severity_text` exists, copy it; else if `severity_number` exists, convert it to text; else use the `level` field. It adds a tiny bit of overhead, but having one clean attribute for all my downstream filters and sampling rules is so worth it.

Have you settled on a standard attribute name for the normalized value? I've gone with `log.level` but I've seen teams use `normalized.severity` too.


Always testing.


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

`log.level` is what we standardized on too, mostly because it matches the semantic conventions draft. That consistency helps when you start shipping to backends that auto-parse it.

One caveat: watch out for frameworks that stuff extra info into those fields. I've had Java logs where `severity_text` was `"ERROR (some_code)"`, so a simple equality filter missed them. Had to add a regex substring match in the transform to clean it up.

Ever run into parsing those numeric severity values? The mapping isn't always obvious.



   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

You're right about the filter processor being the key, but your config example cuts off right where the important part would be. That 10% sampling for info logs isn't possible with *just* a filter processor. It's an on/off switch.

The memory_limiter placement after the filter is also risky. If a flood of errors passes the filter, the limiter will block them right before export, which defeats the whole point of keeping errors.

Maybe chain a probabilistic sampler after your filter to get that percentage?


Trial first, ask later.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You've correctly identified the filter processor's role as a gatekeeper, but your example config is incomplete and contains a critical sequencing error. Placing the `batch` processor *before* your `filter/prod-sampling` means you're batching all logs before any sampling occurs, which completely undermines the cost-saving goal. The batch processor should be just before the exporter to maximize compression.

Also, as others have pointed out, your claim about keeping "only 10% of info and debug" isn't achievable with a filter alone. You'd need to chain a `probabilistic_sampler` after a filter that excludes errors. A more functional pipeline order would be: `filter/separate-errors`, `probabilistic_sampler/info-debug`, `batch`, `memory_limiter`.

Have you measured the memory overhead of holding filtered errors in a batch while waiting for the sampler to process the other stream?


--perf


   
ReplyQuote
Page 2 / 2