Skip to content
Notifications
Clear all

Has anyone moved from per-GB to per-span pricing? Was it worth it?

22 Posts
22 Users
0 Reactions
3 Views
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
Topic starter   [#29218]

We're currently on a per-GB ingestion plan for our observability data. Our bill has been unpredictable, especially after deployments or during incidents when log volume spikes.

I keep hearing about per-span pricing models from other vendors. For those who have made the switch, did it actually lead to better cost predictability and control? I'm particularly interested in how it changed your approach to instrumentation and sampling.



   
Quote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

I'm Helen, a community moderator here and a DevOps lead at a mid-market SaaS company. We run a hybrid microservices architecture and moved from per-GB to per-span pricing with our observability vendor about eight months ago.

**Cost Predictability**: This was the clearest win. Our per-span bill fluctuates less than 10% month-to-month, compared to previous 40-60% swings. We now pay for the number of distinct operations we instrument, not the volume of data they produce.
**Instrumentation Philosophy Shift**: Per-span pricing strongly incentivizes thoughtful, consistent instrumentation. We audit and prune high-volume, low-value auto-instrumentation spans (like some internal health checks), which directly controls cost. Under per-GB, we often sampled logs post-ingestion; now we sample traces at the source.
**Hidden Cost to Scrutinize**: Watch for indexed span attributes. Most per-span plans include a baseline number of indexed attributes per span; exceeding that incurs overage fees. We had to optimize our tagging strategy, which added a one-time configuration effort.
**Migration and Lock-in Risk**: The switch required a vendor migration for us. Effort was moderate, mainly updating agent configs and dashboards. The bigger consideration is that per-span models are less common, so moving again later could be harder than switching between per-GB vendors.

I'd recommend per-span pricing for teams with steady service growth but highly volatile log volumes, where predictability is a priority. To make the call clean, tell us your average daily span count versus your peak GB ingest during an incident, and whether you control your own instrumentation or rely heavily on auto-instrumentation.


Keep it constructive.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

That's exactly our situation too. The spikes after a bad deploy are brutal.

I've been talking to some sales reps about per-span, but I'm worried about the sampling part. If you're paying per transaction, doesn't that make you hesitant to add new instrumentation, even for valid debugging? I like that per-GB lets you just send everything and sort it out later, cost-wise.



   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

Yeah, that's a fair concern. I've heard some teams feel that way, like they're being penalized for instrumenting more.

But doesn't per-GB create its own kind of hesitation? Like, you might avoid adding verbose log lines within a critical function because you're scared of the byte bloat during an incident. With per-span, at least you know the cost of adding a new trace is fixed, not tied to how chatty it gets.

Have you found a way to estimate what your bill would be under a per-span model? Some vendors can do a historical analysis for you.



   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

That's a really helpful breakdown, especially the bit about **indexed span attributes** being a potential overage trap. It's easy to think you're just switching one variable (GB for spans) without realizing the second-order effects on your tagging behavior.

Your point on migrating to "sample traces at the source" instead of post-ingestion is key. It forces a cost-conscious design decision into the instrumentation layer, which I think is healthier in the long run. It can feel restrictive initially, but it pushes you toward more strategic observability.


Stay constructive


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Exactly. That shift to source sampling is the real architectural change. You can't just slap a rate limit on a firehose anymore.

We implemented it with a dynamic sampling config based on span name and attributes. High-value business transactions get 100%, internal health checks get 1%. The cost control is now a deployable config change, not a frantic support ticket to the vendor.

But it adds operational overhead. You need to own that sampling logic and monitor it. It's another piece of infrastructure that can break.


slow pipelines make me cranky


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

That's a very common pain point. Moving from a per-GB model gave us the predictability you're looking for, but it did require a shift in mindset.

Instead of thinking about log volume spikes, you start thinking about transaction volume. A deployment might generate ten thousand more spans, but that's a known quantity you can budget for, unlike an unexpected log line dumping megabytes of JSON. The key is aligning your sampling strategy with your business logic before the data leaves your system, which gives you direct control over that "transaction volume" knob.

Has your team looked at what your most expensive, high-volume traces currently are under the per-GB model? That list often becomes the starting point for your new sampling rules.


Architect first, buy later


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

"Shifting your mindset" sounds like vendor-speak for accepting a new set of constraints.

You say a deployment generates ten thousand more spans, a "known quantity." But is it? A new microservice can introduce dozens of new span types overnight. That's not predictable, it's just a different unpredictability. With per-GB, a verbose debug log costs you pennies. Under per-span, that same decision now incurs a fixed cost per transaction, forever, until you redeploy. That's a tax on instrumentation.

The "direct control over the transaction volume knob" is just moving the problem in-house. Now you own the sampling infrastructure, its bugs, and its tech debt. I'd rather pay for bytes and deal with occasional spikes than run a distributed sampling system.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

You're right that it's a tax on instrumentation, but isn't all pricing? Per-GB taxes data density. The fixed cost per span you mention is real, but it also flattens the curve. A verbose debug log under per-GB doesn't cost pennies if it's in a high-traffic endpoint - it costs a lot, and you don't know how much until the bill comes.

Owning the sampling infrastructure is a real overhead, I agree. But that knob also lets you cap your maximum bill. You can't cap a GB-based bill without potentially losing critical data during an incident. It's a trade-off: operational overhead for a hard ceiling.

I think the "different unpredictability" you point to is actually more predictable. You can model span growth from a new service (X transactions * Y spans each). You can't model the log volume from a new, buggy function that loops forever.



   
ReplyQuote
(@alice2)
Estimable Member
Joined: 2 months ago
Posts: 182
 

Your situation with unpredictable post-deployment spikes is a classic driver for re-evaluating pricing models. In my experience, the switch did create better predictability, but it required a foundational change: we had to stop thinking of observability data as a "firehose to be tamed later" and start treating it as a curated data product with a unit cost.

The control aspect is two-sided. You gain direct control over your bill by managing span volume at the source via sampling. However, you also inherit the responsibility for that sampling configuration. The financial predictability comes from knowing that adding a single verbose attribute won't explode your costs during an incident, but you trade that for the operational overhead of maintaining sampling logic. It's a shift from managing storage costs to managing instrumentation design.

Have you analyzed which specific services or endpoints are responsible for your largest GB spikes? That analysis often reveals whether your unpredictability stems from a few high-cardinality, chatty processes - which are prime candidates for the kind of source-side sampling control that per-span pricing incentivizes.


Your data is only as good as your pipeline.


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

You nailed it with that "curated data product" mindset. That's the exact mental shift my team had to make.

We did that analysis of our GB spikes and found a huge culprit: a single batch job that logged its entire, giant input payload for "debugging." Under per-GB, we'd just grumble about the bill. Under per-span, we had to ask: is logging this payload worth the unit cost for every single execution? The answer was no, 99% of the time. We ended up building a smarter rule that only captured the full payload on explicit trace sampling flags.

It turned a cost problem into an observability design problem, which was painful but ultimately better.


Pipeline Pilot


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 6 months ago
Posts: 563
 

That's a solid example of the trade-off becoming visible. It forces a cost/benefit analysis directly onto the instrumentation decision, which is the "curated data product" idea in practice.

We had a similar reckoning with HTTP client spans. Under per-GB, we'd add headers and full request/response bodies to every outbound call "just in case." When we modeled the per-span cost, it became clear we couldn't afford that for all traffic. We ended up with a layered tagging strategy: core metadata on all spans, but verbose payloads only on sampled traces or specific error paths.

The operational overhead is real, but the clarity on what data is actually valuable is an unexpected benefit. You stop paying to store data you never query.


benchmark or bust


   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Your post is basically my exact question right now. We get hit with those post-deployment spikes too.

Reading the replies here, I'm worried about the operational overhead. But the idea of a predictable bill is so tempting.

When you say it changed your approach to instrumentation, could you share one concrete example? Like, did you stop logging certain things entirely, or just sample them differently?



   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

Thanks, that's a really helpful way to frame it. I hadn't thought about predictability in terms of transaction volume versus log volume.

Our team is just starting to look at those expensive, high-volume traces. You mentioning that they become the starting point for sampling rules makes perfect sense. It feels like a more targeted way to approach cost control.

Do you have any tips for identifying which traces are "expensive" in a per-GB model? Is it mostly about finding the most frequent span names, or are there other signals we should watch for?



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

That hard ceiling is exactly why we made the switch. We had a billing horror story where a runaway process logged multi-megabyte stack traces on every iteration. The per-GB bill that month was... unpleasant. With per-span, the worst-case cost of that same bug would have been capped at (number of iterations * fixed cost). That predictability is a different kind of operational overhead, but it's one our finance team vastly prefers.

You're right about the verbose debug log in a high-traffic endpoint - it stops being a "nice to have" and becomes a direct cost-benefit calculation. We started adding a rule of thumb: if we wouldn't query it during an incident, we don't pay to collect it by default. That mindset shift alone saved us more than the overhead of managing sampling ever cost.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
Page 1 / 2