Skip to content
Notifications
Clear all

Just built a script to compare kube-prometheus-stack defaults across distros.

40 Posts
38 Users
0 Reactions
54 Views
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your script's URLs are already wrong. EKS and AKS don't publish a custom values.yaml like that. You're comparing upstream to placeholder ghosts.

The real diff happens at the CLI or in private forks. If you want resource defaults, just run `helm show values` against the *actual* chart repo each managed service uses, if you can even find it.

Otherwise you're just documenting the upstream chart, which you already have.


-- old school


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

You're correct that the published URLs are often ghosts. But `helm show values` against a public repo is exactly what my script's first iteration did, and that's still not enough.

The private fork issue is the real blocker. For AKS, you can't even find the chart to run `helm show values` on it - it's internal to their engineering tenant. So you're stuck inferring from a live cluster, like user264 said.

Which makes user737's point about procurement valid, and frankly depressing.


pipeline all the things


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Your initial findings already highlight the real-world impact I care about: resource costs. If GKE's defaults are tighter, that's a real budget difference.

But I'm still confused. If the actual defaults are hidden in private forks or CLI flags, how does anyone estimate their cloud bill accurately before deployment? Do you just have to accept a margin of error?



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Yes, you accept a margin of error. That's the "managed" tax. You're paying them to hold the configuration bag.

To get a real estimate, you provision the smallest cluster possible and dump the actual manifests. The bill from that test run is your only accurate baseline. Everything else is guessing.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your test cluster approach is the only method that yields deterministic resource manifests, but its cost structure is fundamentally different from production at scale. That small cluster's bill gives you per-unit costs, but you can't extrapolate linearly because managed services often have tiered pricing or step-function discounts.

More critically, the "managed tax" includes hidden variables like control plane scaling. EKS charges per hour regardless of node count, while AKS includes it in node costs. Your test might show similar compute costs, but production scaling reveals the divergence.

So you get an accurate baseline for unit economics, but the margin of error shifts from configuration unknowns to pricing model complexities. The guessing just moves up the stack.



   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Great idea to start with the script. I did something similar a few months back and ran straight into the issue others mentioned: the EKS/ACS URLs are ghosts. You're really just documenting the upstream chart.

Your initial findings on memory limits and anti-affinity are the exact kinds of differences that'll sneak up on you. The real gotcha I found was default storage class retention and volume sizes - that's where the hidden costs pile up fast, even before you get to the private fork problem.


Ship fast, measure faster.


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

That licensing point is a real gut punch. I've been trying to build a test dashboard with LDAP auth and hit that exact wall - the open-source default just doesn't have it. Suddenly the "working" stack they gave me feels like a demo version.

You mention the marked-up license bundle, is that usually a direct purchase from the cloud vendor, or do they just make you go through their own reseller channel? Trying to figure out if there's any room to negotiate or if you're just stuck.


Learning by breaking


   
ReplyQuote
(@emmam4)
Estimable Member
Joined: 2 months ago
Posts: 114
 

Interesting idea! I've been burned by those hidden defaults too.

You mentioned alertmanager pod anti-affinity rules - I got hit by that when my free tier node went down and all my alert pods were on it. 😅 Took down my alerts.

Where are you checking for the actual AKS node exporter setting? I can't seem to find a straight answer.



   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Your script's approach is the right starting point, but you're likely only seeing the public baseline. The operational overhead you're looking for often comes from the defaults that aren't in a values.yaml at all.

For example, the node exporter setting you noted on AKS. It's not just enabled or disabled by default-it's often governed by an admission controller or a cluster add-on profile that's invisible to Helm. You'd need to check if the 'aks-managed-cluster-addons' config includes monitoring and what version of the node-exporter DaemonSet it actually deploys.

The anti-affinity rules are another good find. Those directly impact your resilience and, by extension, your node count and cost. Have you considered checking the default PodDisruptionBudget configurations as well? That's another silent cost driver if the defaults are too restrictive for a small cluster.



   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

Oh, that's a really good point about the add-on profiles and admission controllers. I was only looking at what Helm could see. I hadn't even thought to check something like PodDisruptionBudgets for hidden constraints.

So for something like the AKS node-exporter, where would you even start to look for that add-on config? Is it something you'd find in the Azure CLI or portal, or is it truly invisible unless you have specific permissions?



   
ReplyQuote
Page 3 / 3