Skip to content
Notifications
Clear all

How do I convince engineering that not every span needs 10 custom tags?

4 Posts
4 Users
0 Reactions
23 Views
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
Topic starter   [#14669]

A common pattern I've observed in teams adopting OpenTelemetry is the tendency to treat span tags as a free-form logging mechanism. While the semantic conventions provide excellent guidance, engineers often add numerous custom tags "just in case," leading to exponential cardinality growth. This directly impacts your observability backend's performance and, more critically, your monthly bill.

Consider a simple HTTP server span. An engineer might add tags for:
* `user.agent.full`
* `request.headers.x_custom_flags`
* `internal.calculation.version`
* `feature.flags.active` (as a map, not a boolean)

Each unique combination of tag values creates a new time series. If `x_custom_flags` has 50 possible values and `feature.flags.active` has 20 combinations, you've just multiplied your series count for that single span.

The technical argument hinges on understanding your vendor's pricing model, which is typically based on:
* **Ingested volume** (GB of spans/logs/metrics)
* **Indexed span count** (spans per second, often with a tag cardinality multiplier)
* **Custom metrics/time series** generated from high-cardinality tags

To build a persuasive case, translate this into a concrete, testable pipeline rule. For example, implement a staging collector configuration that samples or drops spans exceeding a tag threshold, and report the cost differential.

```yaml
# Example OTel Collector processor for tag limiting
processors:
attributes/cardinality_limit:
actions:
- key: "http.request.header.x_custom_flags"
action: delete
# Or, sample only 10% of spans with this tag:
- key: "internal.calculation.version"
action: upsert
value: "sampled"
condition: random(0.1) == false
```

Propose a governance rule: each custom tag must be justified by a specific debugging or alerting use case documented in the service's observability manifest. Tags for "debugging someday" should be sampled, not always-on. Frame it not as limiting insight, but as ensuring the signal-to-noise ratio remains high for the data you are paying to store and query.

--crusader


Commit early, deploy often, but always rollback-ready.


   
Quote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

You're right about the pricing models, but you're missing the operational impact. High cardinality doesn't just bloat the bill, it kills your ability to actually debug things.

When an alert fires at 3am and your dashboard times out because it's trying to render 10,000 unique series for a single service, you're blind. That's the concrete example I use with teams: "Your pager goes off, and you won't be able to load the graphs." It makes it real.

Start by tagging one or two expensive spans in their code with the projected cost per hour. They'll change their tune fast.


garbage in, garbage out


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

You've nailed the root cause, treating spans as structured logs. The financial translation is crucial, but I'd anchor it in their own workflow. Show them the performance cost of their "just in case" data.

Add a `telemetry.impact.cost_estimate` tag to a staging deployment for a single high-traffic endpoint, using your vendor's per-span/tag pricing. The number will be comical. Then, simulate the query latency in their preferred dashboard when filtering on a high-cardinality tag they added. Engineers hate slow tools more than abstract billing concepts.

The real trick is providing a sanctioned, low-cardinality outlet for that debug data. A separate logging stream with a sampled trace ID often satisfies the "need" without poisoning the metrics pipeline.


Measure twice, cut once.


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Vendor pricing models are a distraction. The real problem is most engineers can't even see the cardinality impact until it's too late. The backends hide it behind smoothed aggregates and default views.

Your argument about unique combinations is correct, but good luck finding a vendor's pricing page that actually explains their cardinality multiplier. They call it "indexed span count" and bury the formula in a footnote.

You need a cardinality budget per service, published in dashboards everyone can see. Turns "just in case" into an explicit trade-off.


Your stack is too complicated.


   
ReplyQuote