Skip to content
Notifications
Clear all

Breaking: Karpenter adds native support for Azure Spot VMs in latest commit

4 Posts
4 Users
0 Reactions
9 Views
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
Topic starter   [#26959]

Just spotted this merge in the main branch and thought it deserved a signal boost here. For those running on Azure Kubernetes Service (AKS) or self-managed clusters on Azure, this is a significant step toward parity with AWS support.

The commit introduces a new `Azure` provider type and, crucially, a `Spot` field within the `AzureNodeClass` spec. This allows Karpenter to directly provision spot VMs via the Azure Instance Metadata Service (IMDS), rather than relying on workarounds like manually tainting nodes or using separate nodepools. The key benefit is that Karpenter can now handle spot interruption termination natively, potentially cordoning and replacing nodes proactively.

I'm particularly interested in how the community finds the configuration experience compared to the AWS implementation. The initial spec suggests it's straightforward, but real-world testing often reveals nuances. Has anyone started testing this in a staging environment? I'm curious about the actual interruption handling latency and if there are any gotchas with specific Azure VM series.

This feels like a move that will really bolster Karpenter's position as a multi-cloud tool. It also raises a broader question for this subforum: as these tools expand their cloud support, does it change your evaluation criteria for cluster autoscalers? Do you prioritize a single, robust cloud integration or broader, shallower multi-cloud support?

—HR


—HR


   
Quote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That's great to see Azure getting more love. The native spot handling is a game changer for cost savings on AKS. I'm still learning Karpenter, so maybe this is obvious, but does the new `Spot` field work with the same `ttlSecondsAfterEmpty` setting we use for on-demand nodes? I'm wondering if spot nodes get recycled faster.



   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

Great question! Yes, `ttlSecondsAfterEmpty` applies uniformly, regardless of whether the node is spot or on-demand. The behavior is the same.

One important caveat, though: with spot nodes, you might find that Azure's own interruption notice comes before your `ttlSecondsAfterEmpty` timer expires. In practice, this means spot nodes could be reclaimed by Azure for capacity reasons *before* they're considered "empty" by your own settings, which adds another layer of churn to manage.


Keep it real, keep it kind.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

> configuration experience compared to the AWS implementation

Having used both in early testing, the configuration feels more aligned now, which is good for multi-cloud teams. The AzureNodeClass spec mirrors the AWSNodeTemplate's structure for spot configuration. The main nuance I've observed is that Azure's IMDS for scheduled events can have a slightly more variable propagation delay compared to AWS's more predictable two-minute warning. This might affect how quickly Karpenter's native interruption handling can react.

On VM series, I'd suggest caution with the Lsv2 and Lasv3 series for spot. While cost-effective, they have a higher observed eviction rate in some regions. Testing should include a mix of series to see how your workloads tolerate the different interruption profiles.


Data is the only truth.


   
ReplyQuote