Skip to content
Notifications
Clear all

Walkthrough: Reducing OpenTelemetry cardinality before it hits Claw's intake.

3 Posts
3 Users
0 Reactions
18 Views
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
Topic starter   [#19334]

Everyone's pushing OpenTelemetry as the free savior. Then you get the first bill from Claw (or any other vendor) and realize their pricing is all about cardinality. Your 'free' data pipeline just built you a meter that spins faster with every new service tag.

Here's what we actually do before the OTLP exporter even gets the batch.

First, drop the default resource detectors you don't need. That cloud.platform.instance.id? Probably useless. Set it in the SDK config. Then, write a simple processor to strip high-cardinality attributes from spans. Things like `http.user_agent` or full query strings. You can do it by key name.

Second, and most important, aggressively filter your logs at source. Don't let them become spans or metrics. If you're using the logging bridge, you're paying for debug statements in production. Configure your logger to drop below INFO level in prod, and write a custom LogRecord processor to strip attributes there too.

The goal is to make the pipeline dumb. Send only what you'll actually query on. Every unique tag combination is a line item.


read the fine print


   
Quote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Totally agree, especially about the logging bridge. It's a silent killer on the bill. I'd add that you need to watch out for `db.statement` attributes on spans from auto-instrumentation, too. Even parameterized queries can slip through with different values and explode that cardinality.

One caveat from our mobile apps: we found stripping `http.user_agent` broke some crucial UX analysis. We had to write a processor that hashed it instead of dropping it outright. Gives us a stable tag for grouping without the infinite uniqueness.

Your last line is the real truth. We treat our collector config as a cost-control layer. If we aren't alerting on it or graphing it, it shouldn't be in the telemetry stream. Period.


edge cases matter


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Great point about making the pipeline "dumb". I'm new to setting this up and hadn't considered the logging bridge at all. So if I'm using a logging library that's auto-instrumented, even my debug logs become billable spans unless I filter them out at the logger config level first? That feels like a hidden trap.

Also, `cloud.platform.instance.id` - is that useless because you're already using a proper service name? Trying to figure out which default resource attributes are safe to drop.



   
ReplyQuote