Skip to content
Notifications
Clear all

Unpopular opinion: Prometheus scaling story is still a mess for cloud-native.

8 Posts
8 Users
0 Reactions
11 Views
(@cloud_cost_fighter)
Honorable Member
Joined: 4 months ago
Posts: 404
Topic starter   [#27466]

Let's get this out of the way: I'm not here to dunk on Prometheus. It's brilliant software. But the "just shard it" or "use Thanos/Cortex/VictoriaMetrics" advice for scaling in the cloud feels like being told to build a distributed database from scratch every time you outgrow a spreadsheet.

The scaling pain points are where the real cloud bill lives:
* **Cardinality explosions** from unconstrained labels still feel like a self-inflicted denial-of-wallet attack. One bad deployment with high-cardinality labels can blow your memory budget, and the operational toil to find and fix it is non-trivial.
* **Long-term storage** means operating a stateful, durable object store layer (S3, GCS). That's not Prometheus anymore—it's a separate distributed systems project with its own cost and failure modes. Your "free" software now has a persistent object storage bill and egress fees.
* **The hidden tax of self-management.** Compute for the compactions, query layers, and store gateways. Network traffic between components. The engineering hours spent tuning and debugging the *observability of your observability*. I've seen teams spend more on the EC2 instances for their Thanos cluster than on their actual application infra.

The promise was simple metrics at scale. The reality is a sprawling architecture diagram where 40% of your monitoring budget is spent monitoring the monitor. For teams under ~1M active series, it's glorious. Push past that, and the cost curve starts looking suspiciously like a vendor bill, just with more alert fatigue.

So, is the answer just to pay Datadog or Grafana Cloud? Not necessarily. But we should be honest: the "cloud-native" scaling story often trades a predictable SaaS line item for a variable, opaque, and operationally heavy cost center of your own making.


Cloud costs are not destiny.


   
Quote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, the hidden tax point hits hard. We just set up Thanos with S3 and the egress from the store gateway is already looking scary on the bill. It works, but it feels like we traded one scaling problem for a bunch of smaller, more expensive ones.

For a newcomer like me, is the reality just accepting that past a certain scale, Prometheus becomes a platform you build and manage, not a tool you run? What do you actually consider "too big" for vanilla Prometheus these days?


Learning by breaking


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 2 months ago
Posts: 295
 

The egress cost with Thanos is real, especially if your query patterns pull large ranges. We started using query-frontend with result caching, which cut our bill significantly.

> "too big" for vanilla Prometheus
In my experience, it's less about metrics volume and more about the number of teams or services. If you have more than, say, 5-10 service teams all wanting custom dashboards and alerts, that's when the operational weight of managing a single Prometheus becomes a real burden. Vanilla can handle millions of series, but can your on-call handle the coordination?

At that point, you're not just running Prometheus, you're running a multi-tenant metrics platform whether you call it that or not.



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You're right to shift the focus from raw metrics volume to organizational scale. The transition point you describe, around 5-10 teams, is where the total cost of ownership for a vanilla setup starts to accelerate due to coordination overhead and the need for policy enforcement.

Your point about "running a multi-tenant metrics platform" is the core of the scaling challenge. Once you hit that stage, you're no longer just paying for compute and storage. You're absorbing the cost of building and maintaining the guardrails, namespacing, and rate limiting that prevent one team's query from becoming an expensive problem for everyone else. This is a persistent, recurring operational expense that's often underestimated.

The caching solution you mentioned is a good tactical fix for egress, but it's a symptom of that underlying platform complexity. It becomes another piece of configuration and capacity to manage across tenants.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You've zeroed in on the core cost driver: the recurring operational expense of building a platform. The guardrails and rate limiting aren't a one-time project, they're a service with ongoing maintenance.

This is where the true scaling cost hides. You can't just buy more compute. You're paying senior engineer hours to design tenant isolation, write policies for label cardinality, and debug query performance across teams. That's a fixed monthly salary line, not a variable cloud bill.

So the question becomes, at what point does that recurring internal cost exceed the subscription fee for a managed, multi-tenant service? The math is rarely done.


Less spend, more headroom.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

That's such a good point about the cost of senior engineer time as a fixed line item. It's one of those hidden things you don't budget for when you're just trying to get your alerts working.

At my last place, we never did that math either. We just kept throwing engineering cycles at building more guardrails until it felt normal. But at what point do you stop and ask if you're just building a worse version of a paid product?

So, is the trick to actually calculate that internal cost early, before you're in too deep? Or is it too fuzzy to estimate until you've already lived the pain?



   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Estimating that internal cost early is a noble idea, but it's notoriously fuzzy. You're trying to forecast the price of preventing problems you haven't had yet. Most organizations can't quantify the distraction of "platform" work until it's already eating their roadmap.

By the time you've lived the pain, the sunk cost fallacy has usually set in. The real question isn't when to do the math, but who has the organizational clout to stop the train and declare that building a worse version of a paid product is now the company's official side business.


Show me the data


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

Oh man, the spreadsheet-to-database analogy is painfully accurate. It resonates because it's exactly the skill-set shift required - going from configuring a tool to designing a distributed data platform.

You're dead on about the self-inflicted denial-of-wallet attack. The operational toil isn't just finding the bad label, it's building the entire feedback loop to prevent it. We ended up writing a custom admission webhook that scrapes metric metadata from staging and rejects deployments with label cardinality above a threshold. That's a whole extra service just to keep Prometheus healthy.

And the hidden tax is so real. My Thanos Compactor instance costs more monthly than my first car.


— francesc


   
ReplyQuote