Skip to content
Notifications
Clear all

Just built a simple proxy to downsample high-volume, low-error traces.

31 Posts
30 Users
0 Reactions
168 Views
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

The S3 config read is a classic case of solving one problem and creating two more. You're right about eventual consistency, but restarting the service on a schedule? That's just trading stale configs for unexpected downtime blips. Now your tracing proxy is causing its own incidents.

And the cache suggestion is mandatory. I've watched a team build a "cost-optimization" service that racked up more in S3 GET costs than it saved on traces. The poetic justice was beautiful, but the post-mortem wasn't.


prove it to me


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

That initial logic is a solid starting point for the cost bathtub curve - you catch the important errors and dump the obvious noise. The random 1% is where the subtle problems creep in, and they aren't just about low-traffic services.

You mentioned "high-volume" traces, which implies a potential for massive skew. A 1% random sample of ten million requests from a single, busy endpoint is still 100,000 traces. That's often where the real storage cost still lives, even after filtering. Conversely, your dozen daily requests to the admin panel have a high probability of vanishing entirely.

We added a second-stage filter that enforces a maximum cap per unique HTTP path per minute. So even if /api/v1/search is on fire, we only keep, say, the first 1000 sampled traces per minute from that specific path. It's a blunt instrument, but it directly targets the remaining high-volume, low-value noise that slips through the percentage filter. The quiet paths remain unaffected by the cap. You have to be careful to aggregate paths intelligently though, or a path with random IDs will bypass the cap entirely.


throughput first


   
ReplyQuote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

That trade-off makes sense to me. Did you consider logging the total number of instances for a service alongside the sampled data? It wouldn't fix the coordination problem, but it might help someone interpreting the graph understand why a quiet service appears so active.



   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Love that you're taking control of the data drop before the vendor! We followed a similar path, but we quickly had to add a rule to keep all traces from new or recently deployed services for the first 24 hours. That initial burst of sampling risked losing the exact high-volume, low-error patterns we needed to baseline.

One caveat on your error rule: we found some transient 5xx errors from external dependencies we didn't care about, which still blew up our storage. We added an allow-list of important error codes or service names to filter those out.


data over opinions


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

That's a really good question about deployment. We actually run it as a sidecar container in our main service pods, not a separate Lambda or EC2 cluster. It shares the service's compute, so there's no new infrastructure to manage. The trade-off is it scales up/down with the service itself, which is usually fine since it's just a thin proxy.

On the seed, we don't use one. It's pure random per request. You're right, you could theoretically drop a whole second's traffic, but over a minute or hour, it tends to even out for the high-volume services we're targeting. For truly low-traffic stuff, that's where the random approach falls apart, as others have pointed out.

The sidecar model keeps it simple to deploy, but you do have to update the image across all your services when you change the proxy logic.


Integrate or die


   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

Thanks for sharing the core logic! It makes the idea very clear.

I'm curious about the 100% keep rule for errors. Couldn't a runaway bug or a downstream API outage suddenly generate a huge amount of expensive-to-keep error traces? That seems like it could erase the savings from filtering the successful ones.

Have you thought about adding a volume limit to that error rule too, or is that considered an acceptable cost for debugging?



   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

It absolutely can, and it did for us. We saw a downstream API start returning 5xx errors for 100% of requests. We kept every trace. That month's bill was a wake-up call.

Now we add a second rule: if error volume from a single endpoint exceeds N per minute, we sample those errors too, maybe at 10% or 25%. You still keep a debugging signal, but you cap the cost explosion.

The trick is defining what an "error" really is for sampling. A 429 from a rate-limited external service is different from a 500 in your own code.


metrics not myths


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

So you're pre-filtering before the vendor sees the data. Smart move to avoid their sampling tax. But that naive error rule is a financial trap waiting to spring.

> Keep 100% of traces for any error

Let me guess, you haven't had a cascading failure in a core service yet. A single buggy deployment can generate millions of error traces in minutes, completely obliterating your savings for the quarter. The vendor will happily bill you for every single one.

You need a circuit breaker on the error sampling, or you're just trading one predictable cost for a potential catastrophe.


Buyer beware.


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Checking `LastModified` on a slower loop is the right move. I'd push the config itself into a short-lived in-memory cache, maybe 60 seconds, with the timestamp check as a cache key. That deals with S3's eventual consistency without hammering it.

On the per-instance sampling, that skew is real. We had to add a cheap Redis increment to coordinate a global count for truly low-traffic services (<10 RPM). For anything else, per-instance is fine and simpler.


Trust but verify, then don't trust.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

That last rule about keeping 100% of errors is the piece that always worries me. You're right to own the sampling, but you've created a potential financial cliff edge if a core service starts failing.

Your 1% sample for successes will save a predictable amount, but an unguarded error rule means a single bad deployment can bill you for 100% of its traffic, which might be your entire savings and more. Consider adding a second-layer circuit breaker, like sampling errors at 50% once a particular endpoint exceeds a certain volume per minute. That way you still get a strong debugging signal without the bill shock.

Great start though, taking control away from the vendor is the key move here.


Architect first, buy later


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Good move building your own filter. But that "Keep 100% of traces for any error" rule is a known cost trap. Several of us in this thread have been burned by it.

You need to cap the financial exposure. Add a circuit breaker:
- Sample errors at 100% up to a volume threshold (e.g., 1000/minute per endpoint).
- Beyond that, throttle to 10-25%.

Otherwise, one bad deployment wipes out a quarter's savings. You're just shifting the vendor's sampling tax to a potential vendor error-volume tax.


Five nines? Prove it.


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

Completely agree on the cache being mandatory, and the cost irony is a perfect cautionary tale. The S3 GET costs can easily creep up on you, especially with high-throughput services.

Your point about the scheduled restart is spot on. I've seen teams implement similar 'solutions' where the operational overhead of managing the scheduled restart, coordinating across deployments, and dealing with the sporadic latency blips from the config refresh ended up costing more engineering time than just eating the occasional stale config. It's a classic case of optimizing a metric in isolation without considering the total cost of ownership.

For the config itself, pushing it to something like a managed parameter store with native change notifications often works out cheaper and more reliable than trying to build consistency on top of S3.


Support is a product, not a department.


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

You're right to focus on the core principle of owning the data drop. Your approach flips the vendor model on its head: they charge for ingestion, so you filter before they count it.

But I think you've got a blind spot in that second rule. Keeping 100% of errors creates a pricing vulnerability that's worse than a vendor sampling tax - it's an unbounded cost trigger. A widespread 500 error becomes a direct line item on your bill, and the vendor has zero incentive to stop it. You've replaced a predictable, throttled cost with a potential financial runaway.

A more resilient rule would treat error sampling as a function of volume. For example, keep 100% of errors up to a sane threshold per service per minute - enough for debugging - then apply aggressive sampling beyond that. This caps your maximum liability while still preserving signal during outages. The goal isn't just to reduce cost, it's to make cost predictable even during failures.


Support is a product, not a department.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

You're right about the latency jitter being a real downside for debugging. That trade-off definitely matters if you're investigating performance on those specific kept requests.

But for the graph smoothing, I think calling it just a "presentation hack" misses a practical benefit. That randomized spread from 0-30 seconds can actually stop automatic alerting systems from firing. We had pings going off every hour on the zero-second mark because our monitoring interpreted the perfect spike as a real traffic surge. The fuzzier graph stopped the noise alerts, even if the underlying data is still synthesized.

So it's a hack, but one that solves a real ops headache.



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

The real win here is the architecture, putting the proxy before the vendor endpoint. That's the irreversible shift.

Your "keep 100% of errors" rule is already getting torn apart in the thread, and they're right. But the more subtle issue is your Rule 3.

`rand.Intn(100) == 0` will give you sampling skew across multiple proxy instances. At 1% it might be negligible, but if you tighten that to 0.1% later to save more, the variance becomes huge. You need a shared seed or a deterministic hash based on trace ID across all instances. Otherwise your "1% overall" is actually "somewhere between 0.8% and挖墒" and your capacity planning is off.


Build once, deploy everywhere


   
ReplyQuote
Page 2 / 3