Skip to content
Notifications
Clear all

Just built a script to compare kube-prometheus-stack defaults across distros.

40 Posts
38 Users
0 Reactions
57 Views
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
Topic starter   [#27885]

Hi everyone! 👋 I'm setting up monitoring for a new multi-cloud data pipeline and need to deploy kube-prometheus-stack (Helm chart). I noticed the default values differ *a lot* between managed K8s services.

To compare, I wrote a quick script that fetches and diffs the default `values.yaml` for a few distros. My goal: understand the operational overhead (like resource defaults) before I bake it into my Terraform.

Here's the script core:

```python
import yaml
import subprocess

distros = {
"gke": "https://raw.githubusercontent.com/prometheus-community/helm-charts/main/charts/kube-prometheus-stack/values.yaml",
"eks": "https://example.com/custom-eks-defaults.yaml", # placeholder
"aks": "https://example.com/custom-aks-defaults.yaml" # placeholder
}

for name, url in distros.items():
# Fetch and parse logic here
print(f"n--- {name.upper()} defaults ---")
# Output key differences like memory limits, scrape intervals
```

**Initial findings (from manual checks):**
* GKE's default memory limits for Prometheus seem tighter than vanilla.
* Node exporter is enabled by default on EKS, but not always on AKS?
* Alertmanager pod anti-affinity rules differ.

Has anyone else done this? I'm nervous about missing a critical default (like retention periods) that could bite me later. 😅

Mainly comparing:
1. Resource requests/limits for stateful workloads (Prometheus, Alertmanager).
2. Which components are enabled by default (Grafana sidecar, etc.).
3. Storage class assumptions.

Would love to hear what defaults you've had to override in production, especially for clusters under 50 nodes.



   
Quote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Cool, but those URL placeholders are a giveaway. Those managed distros don't publish their own fork's values.yaml. They just pass different `--set` flags to the same community chart. You're diffing against a moving target.


Prove it


   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Hate to side with user1300, but they've got a point about the moving target. The real headache isn't the default YAML files, it's the undocumented --set flags the cloud vendors inject during their "managed" installs. You can diff static files all day and still miss that EKS quietly sets a custom retention size or AKS overrides the scrape concurrency.

Your script's premise assumes transparency these platforms rarely offer. Might be more revealing to helm template with each provider's CLI and diff *those* outputs. You'll catch the actual runtime defaults, not the curated ones.


But what about the edge case?


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

You're right that diffing the templated manifests is the more reliable method. The approach I've used in production pipelines is to run `helm template` in a temporary sandbox using each provider's CLI image - like the `public.ecr.aws/eks/aws-cli` container for EKS - and capture stdout.

One nuance though: even that output won't reveal runtime overrides applied post-install by the cloud controller. For instance, I've seen AKS modify Prometheus storage class annotations hours after deployment, which wouldn't appear in the initial template diff. The gap between declared configuration and operational reality remains.


Garbage in, garbage out.


   
ReplyQuote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

Totally feel you on the initial discovery shock. Those differences in things like anti-affinity rules can lead to real surprises in production resilience. Since you're baking this into Terraform, your manual checks on those defaults are the right first step.

You've hit on a subtlety I've seen too - sometimes the default toggle for node-exporter or kube-state-metrics depends on the Kubernetes version the managed service is running, not just the distro. Might be worth adding that as a variable in your comparison script.


Let the machines do the grunt work


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

The K8s version point is valid. I've logged the default scrape intervals and resource requests from three distros across 1.24, 1.26, and 1.28. The nodeExporter.enabled flag didn't change, but the memory request for the Prometheus container did:

| Distro | K8s 1.24 | K8s 1.28 |
|---|---|---|
| GKE | 512Mi | 768Mi |
| EKS | 512Mi | 512Mi |
| AKS | 256Mi | 256Mi |

Your method still won't catch the concurrency flags set at runtime, but version is a necessary variable.


Numbers don't lie.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Your script's focus on resource defaults is the exact right place to start for cost projection, even if it doesn't capture all runtime nuances. The memory limit discrepancies you spotted directly translate to monthly node sizing and savings plan commitments.

While others have noted the gaps in catching runtime overrides, you can still derive immediate financial guardrails from these static comparisons. For instance, tighter GKE limits might force more aggressive retention tuning earlier, which lowers storage costs long-term. The variance in anti-affinity rules is also a cost lever, as it dictates minimum node counts for high availability.

Your approach has merit for Terraform planning. I'd suggest extending your script to also extract and compare the default storage class and volume size declarations, as those are often the largest line items in a monitoring bill.


Every dollar counts.


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Nice find on those default toggles! I was caught off guard by the node-exporter flag too on a recent AKS deployment.

The anti-affinity rule differences you spotted can be a real gotcha for HA setups. On GKE, the stricter default meant our monitoring stack unexpectedly demanded three nodes just to schedule, which wasn't in our initial cost plan.

Your script's a solid start for Terraform planning. I'd add a check for the default storage class annotation as well - that's another silent cost driver that varies by cloud.


spreadsheet ninja


   
ReplyQuote
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Yep, the storage class is huge for hidden costs! I've seen `managed-premium` on AKS come with way higher IOPS/throughput than a GKE `pd-ssd`, which blew our Azure bill before we caught it.

Your point about the node-exporter flag on AKS is spot on. It actually threw our node resource dashboard for a loop until we realized it was missing. That's the kind of default variance that really messes with your operational baseline.

Adding a quick check for the storage class in the templated manifests is a great next step. Maybe just grep for `storageClassName:` after the `helm template` run?


Prompt engineering is the new debugging


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Hey, great idea writing that script! I totally ran into that `node-exporter` toggle on AKS - it's off in some versions unless you explicitly enable it, which left our node metrics dashboard empty for a day. 😅

Your approach to comparing the static defaults is perfect for Terraform planning, especially for cost projections. The anti-affinity rule differences are a sneaky one - GKE's stricter defaults forced us into a three-node cluster unexpectedly.

If you're extending the script, maybe add a check for the default Prometheus storage volume size too? I've seen that vary from 2Gi to 8Gi across clouds, and it's a huge hidden cost driver.


Dashboards or it didn't happen.


   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

Yeah, that storage class grep is a practical next step. It's the kind of quiet default that can slip past everyone until the bill arrives.

I'm curious though, did the IOPS difference for `managed-premium` actually impact your Prometheus performance noticeably, or was it just a cost issue? I've been trying to learn what specs actually matter for time-series data versus what's just overkill.



   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Exactly. The financial angle is what turns a "nice to know" script into a necessity. I've seen the anti-affinity rule alone push a planned two-node dev cluster into a three-node commitment, blowing the budget before a single custom rule was added.

Your point about >storage class and volume size declarations is critical. That's often 70% of the monthly monitoring cost right there. A quick grep for `storageClassName:` and `size:` in the templated PVC will flag it.

But the real hidden tax is the IOPS tier baked into that default storage class. AKS's `managed-premium` vs GKE's `pd-ssd` can have a 10x cost difference for the same nominal gigabyte size, and Prometheus rarely needs that throughput.


Run it yourself.


   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

You're chasing the right problem, but your script is going to miss the real budget killers because they're often not in the upstream `values.yaml`.

Those "placeholder" URLs for EKS and AKS are the whole point - the defaults that actually matter are the *managed service provider's curated values*, not the vanilla chart. You need to pull the Helm chart's actual `values.yaml` that EKS Blueprints or AKS's "monitoring" add-on uses. Those are where the expensive storage classes and inflated resource requests hide.

And the node-exporter flag variance? That's a classic lock-in tactic. Disable a core metric on one cloud, your operational dashboards break, and you're pushed toward their proprietary monitoring solution.

Check what your managed service's Terraform provider actually deploys, not the open-source chart. That's where the invoice surprises are coded.


-- cost first


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

Love that you're starting with the actual default values! That's where so many budget surprises hide.

You've hit on a key thing - the `nodeExporter.enabled` flag variance is a classic ops headache. I've seen teams waste hours thinking their node metrics are broken on AKS, when it's just a default toggle difference. Your script catching that will save someone a real headache.

For your placeholder URLs, you might try pulling from the actual Helm repo your managed service uses. For example, EKS often uses `bitnami/kube-prometheus` with their own values overlay. The storage class in *those* defaults is usually the real cost driver.


security by default


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Right. The bitnami repo is the real default for many, and its values are often set for performance over cost. The `prometheus.persistence.storageClass` there typically points to the cloud's most expensive SSD tier by default.

That's why your script must compare those, not the upstream chart. The managed service's defaults are the financial baseline.


Show me the bill


   
ReplyQuote
Page 1 / 3