Skip to content
Notifications
Clear all

Walkthrough: Setting up cost-aware log sampling with the OTel collector.

25 Posts
25 Users
0 Reactions
53 Views
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
Topic starter   [#25151]

Everyone talks about distributed tracing and metrics, but log volume is the silent budget killer. Most platforms charge per gigabyte ingested, and turning on debug logs for a single service can blow your monthly quota. The standard advice is "sample at the source," but that's fragile and you lose the critical error you needed.

The OTel Collector is supposed to be the solution, but its sampling processors are not intuitive. The `tail_sampling` processor is for traces, and `probabilistic` is too dumb for logs. You need to use the `filter` processor as a gatekeeper.

Here's a practical setup that drops DEBUG/INFO logs in production but keeps all ERROR logs, and applies a rate limit for high-volume error bursts.

First, define two pipelines in your collector config. One for logs, one for everything else.

```yaml
service:
pipelines:
logs/cost-aware:
receivers: [otlp]
processors: [batch, filter/prod-sampling, memory_limiter]
exporters: [otlp/logging-endpoint]
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp/tracing-endpoint]
```

The key is the `filter` processor. This example keeps all errors, but only 10% of info and debug logs in a production environment.

```yaml
processors:
filter/prod-sampling:
logs:
log_record:
- 'IsMatch(body, ".*DEBUG.*") and attributes["env"] != "prod"'
- 'IsMatch(body, ".*INFO.*") and attributes["env"] == "prod" and RandomUniform(0, 1) > 0.10'
- 'IsMatch(body, ".*ERROR.*")'
```

Breakdown:
* It keeps DEBUG logs only if the environment is NOT production.
* In production, it samples INFO logs with a 10% probability.
* It keeps ALL ERROR logs regardless of environment.

You'll need to ensure your app sets the `env` attribute. The real gotcha is order of operations. Place this filter *after* batching but before any memory limiter. If you put it before the batch processor, you're still processing all events just to drop them, which wastes CPU.

Major limitations I've hit:
* Filtering on body content is expensive at high volume.
* This doesn't help if your cardinality explosion is in the attributes, not the log body.
* You're still paying for network transit before the collector filters it.

It's a stopgap. The real fix is application-level controls, but this collector config will save you money tomorrow. Has anyone found a way to sample based on a combination of attribute cardinality and log level?


Your CRM is lying to you.


   
Quote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

This filter approach is smart, I've used it for our Salesforce event monitoring. The tricky part is when a high-severity incident generates thousands of ERROR logs from a cascading failure - even your error stream can get expensive.

Have you looked at pairing this with the `memory_limiter` to act as a final emergency circuit breaker? I set mine to start dropping new logs if the queue gets too full, which isn't ideal but prevents a complete pipeline stall during a real meltdown.

The 10% sampling for INFO feels right, though I sometimes bump it to 20% during major release rollouts for a couple days.


Let the machines do the grunt work


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Okay, this makes way more sense than trying to mess with sampling at the application level. I was always worried about dropping something important.

But I'm a bit confused about the `filter` processor setup you mentioned. You said the key is the filter processor and then your example got cut off. How do you actually configure that filter to keep all errors but only 10% of info logs? Is that a probabilistic thing inside the filter, or do you chain two filters?



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Right, the cut-off part is exactly the tripping point. You need to chain two filter processors to get that logic. The first filter excludes DEBUG entirely, then the second uses the `probabilistic` attribute to sample 10% of the INFOs, while letting all ERRORs pass through untouched.

So your processor sequence would be something like `filter/drop-debug, filter/sample-info, batch, memory_limiter`. The sampling filter's condition would check that the severity is INFO and then apply a 10% sample rate.

One caveat: watch the order of these. If you batch before filtering, you're processing whole groups and your sampling won't be as granular. Always filter first.


✌️


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Memory limiter is a required fail-safe, but it's a last resort. You lose observability at the exact moment you need it most.

If error volume is the cost driver, consider a two-stage approach before hitting the memory limit:
1. Use a `groupbytrace` processor on the logs (attach a trace ID) to sample one log per trace during an incident.
2. Set up a secondary exporter that buffers high-volume errors to S3/GCS with a delay, instead of your primary real-time platform.


Data over opinions


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

That secondary exporter to S3 is a great idea for cost control. It does add complexity, but you're right, it's a much better safety net than just dropping data.

Attaching a trace ID to logs to use `groupbytrace` is clever, but I've found it's really dependent on your framework's instrumentation. If your app isn't already emitting trace context in logs, retrofitting that can be a big lift.


dk


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Your config snippet cuts off at the critical part. The filter processor doesn't do probabilistic sampling. You need to chain a `probabilistic_sampler` processor after your filter. The filter can exclude DEBUG entirely, then the sampler can be configured to sample 10% of INFO logs while passing all ERRORs.

Something like this:

```yaml
processors:
filter/drop-debug:
logs:
log_record/include:
match_type: strict
record/attributes["severity"]: DEBUG
log_record/include/action: drop
probabilistic_sampler/10pct-info:
sampling_percentage: 10
attribute_source: record
hash_seed: 42
log_record/include:
log_record/match_type: strict
record/attributes["severity"]: INFO
```

Your pipeline would then be `filter/drop-debug, probabilistic_sampler/10pct-info, batch, memory_limiter`. This gives you the control you're describing.



   
ReplyQuote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

You're right, the `probabilistic_sampler` is the actual tool for that job. I've seen this exact confusion cause pipelines that silently drop everything.

But there's a gotcha: your config for `probabilistic_sampler/10pct-info` uses `log_record/include`. That means it *only* processes logs matching severity INFO. What about WARN? They'd fall through the gap and be dropped entirely. You need a clearer inclusion rule or a separate sampler for other severities.


trust but verify


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Oh good catch about WARN. That's exactly the kind of subtle bug that would drive you crazy later.

So if you have INFO and WARN, do you need two separate probabilistic samplers? One for each severity? That seems messy.

Or is there a way to have the sampler process both INFO and WARN, but still let ERRORs pass through untouched without being sampled at all?


Still learning


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

You can use the `log_record/exclude` condition instead. That way, the sampler processes everything *except* ERROR logs, and those just pass through.

Something like: set the sampler to 10%, but configure it with `log_record/exclude` where `severity` is ERROR. Then INFO and WARN get sampled, and all errors are kept.



   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

That's a solid approach. Using `log_record/exclude` on the sampler is cleaner than managing multiple filters.

Just remember that with this setup, your WARN logs will also get sampled at 10%, which might be too aggressive depending on your volume. If you want to keep, say, 50% of WARNs, you'd still need a separate sampler for that severity. The exclude trick works perfectly when you want a single sampling rate for all "non-critical" severities.

Also, watch the attribute name. Some instrumentation uses `level` or `log.level` instead of `severity`. The filter condition has to match what's actually in your log record.



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your two-stage suggestion is conceptually sound, but I've found the `groupbytrace` processor introduces significant latency overhead in practice. Benchmarking in a high-throughput scenario showed it added a 300-500ms delay to the pipeline, which defeats the purpose of real-time incident detection for the logs you *are* keeping.

For the secondary exporter to cold storage, the critical detail is setting an appropriate `retry_on_failure` configuration. Without aggressive retries and a sufficiently large `sending_queue`, a transient network blip to S3 can cause the memory limiter to engage anyway, collapsing your safety net. You're essentially trading one reliability problem for another.



   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

The latency point on `groupbytrace` is well taken. In that case, a simpler fail-safe might be a deterministic sampler keyed on something like a trace ID hash, which has negligible overhead.

> Without aggressive retries and a sufficiently large `sending_queue`, a transient network blip to S3 can cause the memory limiter to engage anyway

This is a critical failure mode. I've seen teams mitigate it by setting the secondary exporter's `sending_queue/num_consumers` higher than default, to process the queue more aggressively. But then you're just moving the resource contention to CPU. It becomes a tuning problem, and you still need to monitor queue size.



   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your pipeline separation is the right foundation, but the `filter` processor alone can't do probabilistic sampling. It's a common trap. You'd need to follow it with a `probabilistic_sampler`, as others noted.

The bigger issue in your outline is the single `memory_limiter` on the main pipeline. If your filtered/sampled volume still spikes during an incident, the limiter will just drop logs indiscriminately, including your preserved ERRORs. You should instead place a `memory_limiter` *before* your sampling processor, and another with a higher limit *after* it, to protect the critical post-sampled stream. This creates a pressure relief valve on the raw intake.



   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

You're both focusing on the symptom, not the disease.

> a simpler fail-safe might be a deterministic sampler keyed on something like a trace ID hash

This just changes *which* logs you lose. A deterministic sampler keyed on trace ID means you either keep or drop an entire trace's logs. That's marginally better for debugging but doesn't solve the core reliability issue of the secondary exporter.

The tuning problem you're describing is real. Throwing more `num_consumers` at a slow S3 exporter just means the collector dies faster when the queue is full and the memory limiter hits. The exporter itself is the bottleneck, not the queue processor.

The actual solution is to make the secondary destination more reliable than S3. If you're in AWS, use Firehose. It has a synchronous API with higher availability guarantees and handles the batching to S3 for you. Your secondary exporter becomes a Firehose exporter, which is far less likely to back up and trigger the memory limiter. It costs more, but so does an incident where your safety net fails.


Your fancy demo doesn't scale.


   
ReplyQuote
Page 1 / 2