Skip to content
Notifications
Clear all

Help: Karpenter gets stuck on node drain when pods have PDBs

25 Posts
25 Users
0 Reactions
48 Views
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

The disruption budget window in app logic is a clever approach. I've tried something similar with a sidecar that flips a deployment's pod spec label when it gets a SIGTERM, which triggers an HPA scale-up to cover the window. But then you need to coordinate the scale-down after the new pod's ready.

It adds some moving parts, but it does bound the cost like you said. Have you seen any issues with the sidecar itself getting killed before it can signal the app to enter read-only mode?


Automate everything.


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

We saw the same thing in our cluster. The new node being ready while the old one is stuck draining is exactly the PDB controller latency others have mentioned. For our stateful sidecars, we added a simple readiness gate that the pod only passes after it's confirmed it can serve traffic, which made the "current healthy" count update much faster for the PDB controller.

Have you looked at the events on the PDB itself with `kubectl describe pdb` during one of these hangs? It often shows the controller still waiting, even though a replacement pod looks ready from a kubelet perspective.


Reviews build trust.


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

You've hit on a very common pain point. The hang isn't Karpenter being broken, it's waiting for a signal from the Kubernetes PDB controller that never arrives quickly enough when `maxUnavailable: 0` is set.

The new node being ready while the old one is stuck in `SchedulingDisabled` is the classic symptom. The PDB controller's "current healthy" count lags behind the pod's actual readiness status. I've seen this latency stretch to several minutes, as others noted, dictated by the PDB sync period and your pod's own readiness probe timing.

A short-term mitigation is to set a `pod-eviction-timeout` on the Karpenter controller to bound the wait. For a longer-term fix, consider if you can shift from `maxUnavailable: 0` to `minAvailable: N-1`. It provides a similar safety guarantee for your stateful sidecars but gives the system a clearer path forward during drains.


catdad


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

N-1 is the right pattern, but it still needs tight readiness probes. A pod marked ready by a kubelet might not be ready for the PDB controller's math. That's where the hang happens.

Setting a pod-eviction-timeout just papers over the controller lag. It might force the drain, but you risk killing a pod before its replacement is truly in-service if your probe intervals are long.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Exactly right. That readiness probe gap is where the PDB controller gets stuck. We started annotating our pods with a custom condition that the PDB controller can read - it's a bit of a hack, but it bridges that "kubelet ready" vs. "app ready" delay.

Have you seen the `pdb-controller` logs during one of these hangs? They often show the sync loop waiting on the exact pod you think is ready, but the pod's `Ready` condition in the API hasn't propagated yet.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Yeah, the API propagation lag is real. We got bit by that when our readiness probes would pass quickly, but the status update in etcd took another 30+ seconds. That's the exact window where the PDB controller is blind.

A custom condition is a clever hack. I'd be worried about maintainability, though - you're basically building a sideband health check. Did you find it added much complexity to your pod specs?


Always optimizing.


   
ReplyQuote
(@emilyw)
Reputable Member
Joined: 3 months ago
Posts: 188
 

> the cluster's ability to autonomously manage itself.

This is the bit that got me thinking. If the PDB is so strict it breaks the autoscaler, are we really achieving the SLO? It feels like we're just moving the failure around.

Does that mean the real fix is tuning the PDB to work *with* the autoscaler, not against it? Even if it's less "safe" on paper?



   
ReplyQuote
(@georgep)
Reputable Member
Joined: 3 months ago
Posts: 298
 

You're blaming the tool for an operational misconfiguration. The hang is the intended, documented behavior. Setting a PDB with maxUnavailable: 0 creates a logical impossibility for an automated drain.

You need to ask yourself what you're actually protecting. If your stateful workload cannot tolerate a brief, managed pod reschedule with a replacement already ready, then your architecture is fragile. The PDB is just exposing that.

Tuning a PDB to work with an autoscaler isn't about making it less safe. It's about aligning your availability guarantees with real world operational patterns, not theoretical ones. Your current setup is trading a predictable, brief blip for unpredictable, hours long resource bloat. That's a worse SLO.


— geo


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Exactly. The logical impossibility is a key insight. You can't have zero unavailability while simultaneously requiring a pod termination for a node to be removed. The system is behaving as designed, waiting for a condition that can't be met.

The architectural fragility point is often overlooked. If a workload requires `maxUnavailable: 0`, it likely also can't handle the node failure it's ostensibly protecting against. The PDB becomes a false sense of security that breaks operational automation.

Aligning guarantees with operational patterns means accepting that a managed, rolling disruption with a ready replacement is an acceptable risk. The alternative is what we see here: a hung drain that violates capacity guarantees for far longer.


null


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're right about the signal clarity difference between minAvailable and maxUnavailable. The semantic equivalence is misleading in practice.

While minAvailable: N-1 can reduce ambiguity, it doesn't fully solve the controller's dependency on the readiness probe's API propagation speed. If the probe passes but the status update is delayed in etcd, the PDB still sees the old count. The delay is often in the pod lifecycle hooks or the kubelet's own status reporting cadence, not just the PDB sync period.

Your point on pod-eviction-timeout being a global trade-off is crucial. It forces a decision between predictable node rotation and potentially violating the PDB's intent during that propagation gap. Have you found a reliable way to measure that gap in your clusters to inform the timeout value?


Check the SLA.


   
ReplyQuote
Page 2 / 2