We're currently evaluating a move from the default Canal CNI to Cilium across our RKE2 clusters. The promise of Hubble for observability and potentially better network performance is really appealing, especially as we scale beyond 50 nodes per cluster.
I'm looking for real-world feedback from teams who've made this switch. Specifically:
* **Upgrade stability:** Did you install Cilium fresh or migrate an existing cluster? Were there any hiccups during RKE2 version upgrades afterward?
* **Operational overhead:** How has day-to-day management compared to Canal? Any surprises with resource usage or troubleshooting?
* **Networking & features:** Are you using any of Cilium's advanced features (like Kubernetes Network Policies, service mesh, or cluster mesh)? Did the default configurations work well with RKE2's architecture?
Our primary driver is gaining better visibility into east-west traffic. If you've run both, what was the tangible difference in operational insight? Any gotchas in the Helm chart values for a stable RKE2 integration would be invaluable.
I've seen the official docs, but practical benchmarks and team adoption stories are harder to find. Thanks in advance for sharing your experiences