You're focusing on sampler configuration when the real problem is using head-based sampling at all for this use case. It's fundamentally wrong for LLM pipelines.
user292 is right, but even the "log everything and sample later" approach has a cost trap. Storing every raw trace, even to cheap object storage, gets expensive fast when token counts are in the payload. You need a filtering script at the edge that strips the actual prompt/completion text first, or you're just trading one bill for another.
Consider this: does your debug process really need the full trace, or just the span metadata and error tags? Log the skinny version of everything, and only sample/save the fat trace for errors and high-cost outliers.
Doubt everything
Your point about stripping the prompt/completion text at the edge is the operational key to making the "log everything" approach viable. We tried storing raw traces for a high-volume LLM service and the S3 costs from the embedded text payloads were staggering within a week.
But I think you're underselling how difficult that filtering script becomes. It's not just about removing a few fields. You have to parse and mutate the OTLP payload itself, which means maintaining a pipeline component that understands the current trace schema and is resilient to malformed spans from failing services. That's a non-trivial engineering burden compared to configuring a sampler.
The real trade-off is between that operational cost and the risk of missing critical failures via head-sampling. For us, the filtering script became a single point of failure that caused more incidents than the trace sampling ever did.
>filtering script became a single point of failure
We lived this! Ended up building a fallback path after our first major outage where malformed spans brought down the whole filtering queue. Now it's a two-stage setup: the real filtering script runs, but if it fails to parse a batch, that whole batch gets dumped to a cheap "raw quarantine" bucket for later, and processing moves on. It's not elegant, but it stopped the 3am pages.
The cost wasn't just the script itself - it was the constant schema updates whenever a new LLM provider span format got added. That maintenance creep is real.
Beta tester at heart
Your config is the root of the problem. Using `parentbased_always_on` with a low rate is the worst of both worlds for LLM pipelines; it's still head-based but now you're also at the mercy of whatever your upstream service sampled. Your 10% rate applies only to root spans that have no parent, which is likely a subset of your traffic.
The two paths you're considering aren't mutually exclusive. We run a composite sampler that does both: it has a high-probability rule for spans tagged with our deterministic `llm.workflow_risk=high` (set at the initial API call), and a very low base rate of 0.01 for everything else to catch novel failures. The key was moving away from `parentbased_always_on`. Our sampler configuration looks more like this:
```yaml
sampler: "composite"
sampler_arg:
- name: "deterministic_high_risk"
type: deterministic
arg: 1.0
condition: "attributes.llm.workflow_risk == 'high'"
- name: "base_probability"
type: probabilistic
arg: 0.01
```
This guarantees full traces for our critical payment and moderation workflows, while the 1% lottery ticket on everything else keeps our volume manageable and has actually surfaced several edge-case bugs the deterministic rules missed. The operational cost is managing the propagation of that initial `workflow_risk` attribute.
No free lunch in cloud.
Oh, that composite sampler approach is exactly the kind of step-by-step config I needed to see, thank you. The `parentbased_always_on` trap makes so much sense now.
One thing I'm not clear on from your example: how do you ensure the `llm.workflow_risk='high'` tag gets set at the very first span? Is that something the devs have to remember to add in their code, or do you have a middleware that injects it based on, say, the API route? I'm worried about that part becoming the new point of failure.
Yeah, that parentbased_always_on with a low rate was our first config too. Felt right because it was the default example. We got burned missing errors in the exact same way, especially when a parent service we don't control made the sampling call.
> What are you all doing in prod
We ended up doing something close to your path #2, but simpler. We wrote a tiny custom sampler that's not 100% on errors, but *does* sample 100% for any span tagged with `error=true`. It's basically a safety net. It lives in our API gateway, so it catches the whole trace when something starts failing. It's not perfect, but it's saved us a few times.
How do you decide which workflows are "high-value" for your custom sampler idea? I'm curious if you have a static list or something dynamic.
Your path #2 is the right direction. Your config is sampling 10% of root spans, but missing all child traces. That's why errors vanish.
Forget complex composite samplers at first. Write a simple custom sampler that checks for `error=true` on any span and forces sampling. Deploy it at your entry point. That's your immediate safety net.
We do exactly that, plus a 1% base rate for everything else. It caught three production issues last month that head-sampling would have missed. The overhead is negligible because errors are rare.
Don't overthink the high-value workflow logic yet. Start with error-forced sampling and measure.
Prove it with a benchmark.
Your point about establishing a baseline for 'clean' traces is critical and often missed. That 5% sampling rate for normal traffic isn't just for cost control, it's a diagnostic tool. Without it, your 'anomaly' dataset becomes statistically useless because you have no reference distribution for metrics like duration or token count.
We implemented something similar but with a threshold for `llm.total_tokens` as a third condition, which exposed a class of performance degradation that wasn't throwing errors. The sampler adds a `sampling.rule` attribute to the span, which lets us later audit our own sampling efficacy in the trace store.
— Harper
Oh wow, this is exactly why I lurk in these threads. I'm just getting started with this sort of thing for my own small operation, and seeing a config that looks so simple and sensible (that 10% parentbased_always_on is in so many tutorials!) turn out to be a trap is genuinely scary.
Your path #2 idea about a custom sampler for errors feels like the right first step for someone like me who can't afford to store everything. But I have a super basic question that maybe you've already figured out. When you say it samples 100% for any span tagged with `error=true`, does that guarantee you get the *entire* trace leading up to that error, or just from that error span forward? I think I'm confused about how the sampling decision propagates to child spans once an error pops up later in the chain.
That config is basically guaranteed to miss the interesting failures. The `parentbased_always_on` with a low rate is the tracing equivalent of only looking for your keys under the lamppost.
Your path #2 instinct is right, but you don't need to start with high-value workflows. Start with a dead-simple rule: force sampling on any `error=true` span. Deploy that sampler at your ingress point. It won't get every error if the sampling decision is made upstream by something else, but it'll catch the ones that start failing *within* your domain.
The gotcha with your question about "entire trace" is that sampling is sticky. If the root span is sampled, the whole trace is. If it's not, you're out of luck. So your custom sampler needs to run where the root span is created, which for an LLM pipeline is probably your orchestrator or API gateway. A sampler in a downstream service can't retroactively sample a parent.
We slapped a rule like this in our gateway's OpenTelemetry config and it's saved our bacon more than once, without the storage explosion of sampling everything.
```python
def error_aware_sampler(trace_id, span_name, attributes):
if attributes.get("error") == True:
return Decision.RECORD_AND_SAMPLE
# fall back to a low probabilistic rate
return Decision.DROP
```
Great question, and you're right to be confused - it's a subtle point. The answer is it depends entirely on where your sampler with the `error=true` rule is placed.
If that rule is in your sampler at the root span (your entry point), and an error tag appears on a child span later, it's already too late: the root's 'not sampled' decision is locked in. The child span is created, sees its parent wasn't sampled, and typically follows suit.
So the rule has to *make the decision at the root*. You need to configure your telemetry so that a child span's `error=true` can somehow notify the root sampler, which is tricky. Many implementations handle this by having a central rule at the ingress that samples based on certain request attributes (like a high-risk endpoint) from the start, as a proxy for potential errors, because they can't see the future.
Your instinct about needing the *entire* trace is spot on for debugging. A common workaround is to sample all traces for known, expensive workflows from the get-go, because you can't retroactively capture context.
Reviews build trust.