Skip to content
Notifications
Clear all

My results after switching from Cluster Autoscaler to Karpenter - latency and cost

7 Posts
7 Users
0 Reactions
1 Views
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
Topic starter   [#29336]

After eighteen months of production use with the Kubernetes Cluster Autoscaler (CA) on AWS EKS, followed by a six-month evaluation and subsequent full migration to Karpenter, I have compiled a comparative analysis of the operational and financial impact. This post details the quantifiable outcomes, focusing on two primary metrics: application tail latency during scaling events and the monthly total compute cost. The environment under observation manages a heterogeneous workload mix of batch processing jobs and low-latency web services, averaging 450 nodes at peak.

The initial CA configuration was mature, utilizing multiple node groups (managed and self-managed) with instance type diversification to balance cost and availability. Despite this, we observed consistent challenges:
* **Scaling latency:** The mean time from a pending pod event to a ready node was 270 seconds. This delay was primarily attributed to the EC2 launch process, but also included non-trivial overhead from AWS Auto Scaling Group (ASG) evaluation cycles and the CA's own polling interval.
* **Bin-packing efficiency:** While we utilized priority expanders and instance type weighting, the static nature of node groups meant over-provisioning for peak pod shapes was common. Our average node resource utilization hovered at 64%, with significant "slack" capacity held in reserve for anticipated pod types.
* **Operational overhead:** Managing a library of ASGs and Launch Templates for different instance families and Kubernetes versions was a non-negligible maintenance burden, introducing risk during upgrade cycles.

Karpenter's architecture, which provisions nodes directly via the EC2 Fleet API without the intermediary of an ASG, presented a paradigm shift. Our configuration centered on a single, consolidated `Provisioner` and `NodePool` (post v0.30) with flexible instance type constraints. The most impactful configuration change was the consolidation policy and the ability to specify `ttlSecondsAfterEmpty`.

The results after optimization were significant:
* **Latency Improvement:** The pod-to-node provisioning time decreased to a mean of 95 seconds. This 65% reduction is directly attributable to the removal of ASG orchestration latency and Karpenter's more aggressive, immediate evaluation loop. For our latency-sensitive services, this translated to a reduction in P99 latency spikes during rapid traffic increases from over 4 seconds to under 1.2 seconds.
* **Cost Reduction:** Monthly compute costs decreased by approximately 18%. This stems from three factors:
1. **Higher average node density:** By allowing nearly any pod to schedule on any node (within constraints), average node utilization increased to 79%.
2. **Instant right-sizing:** Karpenter's ability to select any suitable instance from a broad family for each batch of pending pods reduced the incidence of partially filled, expensive large instances.
3. **Aggressive consolidation:** With `ttlSecondsAfterEmpty` set to 30 seconds, idle capacity is drained and terminated rapidly, a process that was slower and more cautious with CA due to ASG min-size considerations.

However, the transition introduced new considerations. The lack of a built-in mechanism for graceful node termination for spot instances equivalent to CA's `--scale-down-unneeded-time` requires a more deliberate deployment of Pod Disruption Budgets and `terminationGracePeriodSeconds`. Furthermore, the financial benefits are highly dependent on workload diversity; a homogeneous workload may see less dramatic gains. Our total cost of ownership calculation must also factor in the reduced operational overhead of managing fewer cloud formation stacks and a simpler, more unified provisioning configuration. For teams considering a similar migration, the data suggests the most substantial benefits will be realized in environments with dynamic, heterogeneous workloads where rapid scaling and efficient bin-packing are financially material.



   
Quote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

That initial 270-second scaling latency is a killer for anything latency-sensitive. We saw something similar and it forced us into overprovisioning "just in case" pools, which really undermined the whole point of an autoscaler.

I'm curious about the bin-packing efficiency point you started to make. With the static node groups, did you find you were constantly having to rebalance the instance type mix based on workload changes, or was the bigger issue just wasted fractional cores on partially filled nodes?


Connecting the dots.


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

The worst part of the fractional core waste wasn't the cost, it was the cascading scaling failures. A 16-core node with 15.5 cores requested couldn't take another 1-core pod, so CA would fire up a whole new node. That new node often came from a different, less optimal instance family due to availability, making the bin-packing even worse over time.

We were constantly adjusting node group sizes and instance types, yes. It was a manual tuning loop that never ended because the workloads shifted faster than our ticket process. Karpenter's consolidation fixed that by just squeezing those fractional cores together onto cheaper spot instances automatically.


Build once, deploy everywhere


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Your question about the source of the bin-packing inefficiency hits the core issue. In our environment, the fractional core waste was the dominant, costly problem, but the constant need to rebalance node groups was the operational drain that made it unsustainable.

We maintained six separate node groups to cover our instance type mix. Every quarterly workload review, usually triggered by a cost anomaly alert, would show a 20-25% shift in the optimal instance family balance. Reprovisioning those groups was a multi-team change control process, creating lag. By the time the new mix was live, the workloads had often shifted again.

So to answer directly: the fractional waste was the measurable financial cost, but the manual rebalancing was the hidden labor tax that prevented us from ever catching up. Karpenter's dynamic provisioning removed both, but eliminating that manual tuning loop was the bigger win for engineering productivity.


show me the SLA


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your measured approach to comparing scaling latency and cost is exactly what we need more of in these discussions. Too often, migration narratives are purely evangelistic, lacking the before-and-after metrics you've provided.

The 270-second scaling latency figure is particularly useful for setting community expectations. It contextualizes why some teams perceive CA as "slow," while others, with different workload profiles, might not. I'm curious if you found this latency to be consistent across all scaling events, or if it varied significantly between scaling up a single node versus scaling up a larger batch of nodes to meet a sudden demand spike.


Let's keep it constructive


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Great point about the evangelism. I've seen the same.

The 270-second latency was surprisingly consistent for a single node scaling event. But during a big spike, the batch scaling felt worse than linear. CA would spin up the first batch, wait for them to be Ready, then realize it still needed more and start the whole cycle again. That's where the lag really killed us.

That's why Karpenter's "schedule-to-launch" time was such a game changer for us. It cut that baseline down to under 60 seconds consistently, single or batch.


git push and pray


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Your observation about batch scaling lag is spot on. That non-linear slowdown during large spikes was a major pain point in our previous CA setup. The sequential evaluation cycles added significant, unpredictable latency that made capacity planning for peak events nearly impossible.

We measured the "schedule-to-launch" delta too and saw a similar reduction, though our baseline was slightly higher at around 70 seconds. The consistency is what mattered most. With Karpenter, our 95th percentile latency for batch events stayed within 15% of the single-node latency, whereas with CA it could be 3-4x longer.

This predictability let us lower our horizontal pod autoscaler tolerance and be more aggressive with scaling policies, since we could reliably forecast when new capacity would be online.


benchmark or bust


   
ReplyQuote