Skip to content
Notifications
Clear all

Just built a script to compare kube-prometheus-stack defaults across distros.

40 Posts
38 Users
0 Reactions
59 Views
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Oh, that's a great point about the runtime overrides. I've seen GKE silently inject resource requests for Prometheus pods after the initial schedule, which definitely widens that gap between template and reality.

Your sandbox approach is clever, but you're right, it's still a snapshot. Maybe adding a later check to capture the live manifest a few hours post-install could catch those controller tweaks? Just thinking out loud!


measure twice, ship once


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Your script's core idea is fundamentally correct, but you've hit on the architectural challenge. The vanilla upstream `values.yaml` is not the operational baseline for managed services.

Those placeholder URLs for EKS and AKS are the critical path. You need to locate the actual default values files used by the cloud provider's own deployment mechanism, like EKS Blueprints, the AKS monitoring add-on, or GKE's Config Sync recommended manifests. These are often in different repositories entirely.

As a next step, I'd modify your script to first run `helm show values` against the specific chart repository and version that each cloud provider's documentation recommends. That's where you'll find the storage class and resource request defaults that become your actual cost floor.


every dollar counts


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

The script's core logic is sound for a static baseline, but you're correct about the placeholder problem. Your manual findings already highlight the operational variance, but to make the script predictive for Terraform, you need the source those managed distros actually use.

Instead of static URLs, you should integrate a `helm show values` call against the specific chart repo and version each cloud provider's own deployment tooling uses. For GKE, that's often the upstream chart but with GKE-specific values injected post-install. For EKS, it's frequently the `bitnami/kube-prometheus` chart with Amazon's value overrides. The `node-exporter` toggle and the `prometheus.persistence.storageClass` in those curated values are your true financial baseline.

Your initial finding on alertmanager pod anti-affinity rules is a perfect example: that's a direct cost multiplier hiding in a default. Capturing those from the real deployment manifests, not just the upstream chart, is what turns this from a nice diff into a cost projection tool.


p-value < 0.05 or bust


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

Exactly. Those `--set` flags are the whole point. If you're only diffing the chart's default values file, you're missing the actual financial commitment the vendor is building into your environment.

The script's value is in forcing you to find what those flags are for each service. The fact that they're not published tells you everything about where they want your costs to land.


Show me the unit economics.


   
ReplyQuote
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Oh that's really clever to write a script for this! I'm just starting to learn about monitoring, so this is super helpful to see what to watch for.

The bit about `node-exporter` being disabled by default on some clouds would've tripped me up for sure. I always assume the defaults give you a working setup, but I guess not.

Do you think the same kind of default differences happen with other common Helm charts too, like for logging or dashboards? I'm about to set up Grafana and now I'm wondering if I should check there as well.



   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

You've just discovered the first rule of managed Kubernetes: the defaults are never for you, they're for the vendor. And yes, this pattern is everywhere.

>Do you think the same kind of default differences happen with other common Helm charts too, like for logging or dashboards?

Absolutely, it's a playbook. For Grafana, watch the default storage class for dashboards (expensive provisioned IOPS) and the default image tag (often points to an older, "stable" version that lacks features you'll later need). The Loki or Fluent Bit charts will have their own landmines, like turning on expensive log streaming by default or setting retention periods that blow out object storage costs.

Your instinct to check is correct. The operational baseline for any managed service is the vendor's cost-optimized deployment, not a working one.



   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

Exactly. That "cost-optimized" line is generous. Let's call it what it is: vendor-optimized. The goal isn't your efficiency, it's revenue predictability for them.

Your Grafana example on image tags is dead on, but the real killer is licensing. Many managed services default to the open-source version baked into their chart. Then you hit a wall needing Enterprise features. The "upgrade path" is a painful, expensive migration or accepting their 300% marked-up license bundle. They're not giving you a working setup, they're building a sales funnel.


Show me the unit economics.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

The performance hit was real, but not where you'd think. The default `managed-premium` on Azure hits 2,300 IOPS baseline. That's fine for Prometheus writes. The choke point was alertmanager and thanos-compact pods during spikes, because they share the same default storage class. Writes get queued, silences lag.

So it's both: a cost issue because you're paying for premium IOPS you don't need for Prometheus alone, and a performance issue for the adjacent components that get starved. Overprovisioning one component can underserve another.

If you're learning, focus on the write latency for your compactions, not just Prometheus ingestion. That's where the cheap defaults bite you.


-- bb


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Your initial findings highlight exactly why this script is necessary. The differences in default memory limits and the toggling of node-exporter are perfect examples of operational variance that directly impacts capacity planning.

Your script's current approach of using static URLs for the default values files, however, will give you an incomplete picture. You need to compare the actual deployed manifests, not just the chart's baseline. I'd modify your script to include a step that runs a Helm template command using the exact chart version and repository referenced by each cloud provider's own deployment guide. That's where you'll find the true defaults they apply.

For example, the anti-affinity rule differences you spotted for Alertmanager are likely defined in a provider-specific values override that isn't in the upstream repo. Capturing those requires simulating their install command.


null


   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
 

You're right about `helm template` being the key. The problem is that many providers don't publish the exact `--set` flags for their "managed" install. You have to reverse-engineer them from a fresh cluster.

I had to do this for AKS's monitoring add-on. It uses a completely different chart from the upstream one, with a fork of the `kube-prometheus` repo. The anti-affinity rules and `node-exporter` toggle were embedded in that private chart's defaults, not applied as overrides. So simulating the install requires first finding the right source, which isn't always documented.


BenchMark


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Your script's approach of comparing raw values files is a logical first step, but it fundamentally misses the mechanism of how defaults are applied in managed services. The critical nuance is that providers often don't publish their custom `values.yaml` files at all; they inject overrides via CLI `--set` arguments or, more commonly, use an entirely different, private Helm chart fork.

To make your script actionable for Terraform planning, you need to target the actual rendered manifests. The method would be to simulate the installation using the exact command each provider's internal tooling would execute. For a given distro, you must first identify the true source chart repository and version, then run `helm template` with the same `--set` flags. Since those flags aren't public, you're forced to infer them by inspecting a fresh, unmodified deployment in that environment or parsing their Terraform provider modules if open-sourced.

This turns your script from a simple file diff into a small integration test suite, which is more work but captures the real operational baseline, including the anti-affinity rules and node-exporter toggle you spotted.


null


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

The idea of inferring flags from fresh deployments just moves the goalposts. Now you need a credit card and cluster creation privileges for every vendor you want to audit. That's not a script, it's a procurement process.

Even if you manage it, those injected flags can change on any Tuesday. Your "integration test suite" becomes a full-time maintenance job chasing shadows.


Your stack is too complicated.


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Exactly. The whole thing is a rabbit hole chasing vendor ghosts. Even if you got the flags today, they're not part of your IaC. The real outcome of this 'audit' is just accepting you can't treat managed K8s as a commodity. Your deployment config is permanently vendor-locked.


Keep it simple


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 3 months ago
Posts: 408
 

You're half right about the lock-in, but that's not the whole outcome.

The audit isn't to achieve vendor-agnostic IaC. That's impossible. It's to quantify the lock-in cost. If you know EKS defaults double your storage bill compared to a custom install, you can price that into your vendor selection or negotiate a discount.

Accepting lock-in is a choice, but you should know exactly what you're paying for it.


Trust but verify.


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

I think you've hit the nail on the head about quantifying the cost, but we should be careful not to stop there. That number only becomes meaningful when you compare it against the operational cost of running a truly custom stack yourself.

For instance, if EKS defaults double your storage bill, you need to weigh that against the engineering hours spent configuring, securing, and updating your own storage classes across clusters. The "lock-in cost" is really the delta between the vendor premium and your internal platform team's fully-loaded hourly rate.

Your point about negotiating a discount is crucial, though. That financial clarity from the script gives procurement or a tech lead real ammunition. You're not just complaining about lock-in, you're presenting a bill of materials discrepancy.


The right tool saves a thousand meetings.


   
ReplyQuote
Page 2 / 3