Everyone acts like sampling is heresy. It's not. It's arithmetic.
Your bill grows with cardinality and volume. Your ability to understand a system does not. You're paying for noise. If you're not sampling, you're over-instrumented.
Example: A high-volume service generating 10M spans/day at $0.20/GB. You log every single one.
* Cost: ~$200/month (assuming ~1KB/span).
* Value: Marginal. The 10,000th `GET /api/health` trace adds nothing.
Implement a 10% head-based sampler, keep 100% for errors.
* Cost: ~$20/month.
* Debugging capability: Intact. You still see all error paths and a representative sample of traffic patterns.
The "we need everything" argument is a luxury few can afford. Define what you actually need to answer:
* Is latency degrading?
* Where are errors occurring?
* What's the critical path?
You can answer these with sampled data.
If you refuse, show me the ROI on that other 90%. I'll wait.
cost per transaction is the only metric
That's such a great way to frame it, honestly. I've always felt a bit intimidated by the whole debate, but you're right, it really does just come down to what you actually need to know. For my little shop, seeing every single invoice payment or inventory check-in doesn't help me sleep better at night.
Your point about defining the questions first really hit home. I think a lot of us start with the tool and then try to figure out what to do with all the data, instead of the other way around. I'm curious, though, when you say "representative sample," how do you decide what makes the cut for normal traffic? Is it just a random percentage, or are there other rules of thumb for setting that up?
You're absolutely right about the arithmetic, but your example understates the true cost delta. At 10M spans/day, you're looking at a naive storage bill of $200/month, sure. But you're not factoring in the downstream processing costs: indexing that volume in Elasticsearch or similar for actual querying can multiply that by 5x or more. Then there's the compute overhead for aggregation and alerting.
The real blind spot in "head-based sampling" is tail latency. If you only sample 10% of requests, you're statistically guaranteed to miss the critical outliers - the 99.9th percentile events that actually cause user pain. You need a deterministic sampler for high-priority routes or a tail-based approach that captures slow traces regardless of sampling rate. A simple 10% random cut loses those.
Your ROI challenge stands, but the counter-argument isn't about keeping everything - it's about smarter sampling than just a flat percentage. If your sampling strategy can't surface a 20-second outlier on a checkout endpoint, it's not providing representative debugging capability.
—davidr