You're absolutely right about async tasks breaking context propagation. We got burned by that with our bulk document processing queue. The root span got sampled, but each async summarization job spawned its own root, and because the trace context wasn't in the job payload, those children were lost. We had to manually inject the parent trace ID into the queue message metadata.
Your absurdly low base rate is smart, though. We use a similar trick but tied to high token counts instead of pure random. Anything over a certain token threshold forces a sample, even if the workflow wasn't pre-tagged. It catches those weirdly expensive "successes" that are actually failures in disguise.
Everyone's focusing on the sampling algorithm and missing the real problem. You're going to blow your budget storing all these traces.
You said it yourself: worried about volume and cost. Adjusting the sampler to 30% or adding custom rules without a storage strategy first is how you get a six-figure observability bill. Your LLM traces are huge.
Skip the custom sampler rabbit hole. Do this:
- Force sample 100% for 48 hours on a single, critical workflow.
- Analyze the actual span composition and size. Find the high-cardinality, high-cost attributes.
- Then implement trace-level filtering *before* storage to strip out the bloated payloads you don't need for 99% of queries.
Otherwise your "battle-tested pattern" will be explaining a cost overrun to your finance team.
show me the bill