Skip to content
Notifications
Clear all

Anyone else having issues with Calico and Windows nodes on AKS?

2 Posts
2 Users
0 Reactions
28 Views
(@julieh4)
Trusted Member
Joined: 3 months ago
Posts: 53
Topic starter   [#12349]

Hey everyone, been wrestling with something in our hybrid AKS cluster and wondering if it's a common pain point.

We're running a mixed node pool cluster (Linux and Windows Server 2019 nodes) on AKS, using the Azure CNI with Calico for network policy enforcement. Our Linux workloads are humming along fine, but we keep hitting intermittent network connectivity drops on the Windows nodes. Pods on Windows nodes will just lose connection to services on Linux nodes (and vice versa) for short periods, and the Calico Felix logs on the Windows nodes are spitting out a lot of timeout errors related to the datastore.

What we've tried so far:
* Verified we're using the AKS-recommended Calico manifest for Windows.
* Increased resource limits for the `calico-node-windows` DaemonSet.
* Tried tweaking the `datastore.type` parameter (switched from `kubernetes` to `kdd` which is default for AKS, but no change).

It *feels* like a stability issue with the Calico Windows agent under load. We're not at a huge scaleβ€”maybe 15 Windows nodes.

**My questions:**
1. Is this just a known quirk of the AKS/Calico/Windows combo right now? Should we consider moving these workloads to Linux (a longer-term project for us)?
2. Has anyone found a reliable config tweak or a specific Calico version that improves stability?
3. Or, is the better path to ditch Calico network policies for Windows nodes and rely solely on Windows-native ACLs/host firewall rules (even though that means managing two policy models)?

Would love to hear if you've been down this road and what your operational fix looked like. The inconsistency is giving our monitoring alerts a serious workout 😅

– Julie


Data-driven decisions.


   
Quote
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Oh that sounds frustrating. I'm just starting to explore AKS for our team, and hearing about issues like this is a bit daunting. I'm curious, have you noticed if these drops happen more during specific times, like when there's a lot of pod creation or deletion happening?

We're considering a Windows node pool too for some legacy apps, so this is really helpful to see. If it's a known quirk, maybe we'll look into that Linux migration sooner. Thanks for sharing the details!



   
ReplyQuote