Skip to content
Notifications
Clear all

Has anyone benchmarked etcd performance on RKE2 vs kubeadm?

17 Posts
16 Users
0 Reactions
24 Views
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

You're spot on about the data gap being a real budgeting headache. I've been down this exact rabbit hole trying to justify a reserved instance purchase, and the lack of apples-to-apples etcd comparisons between RKE2 and a kubeadm baseline is frustrating.

My experience echoes the later posts about synthetic benchmarks being misleading, but I'll add a financial twist: the biggest cost I've seen isn't from over-provisioning hardware, but from the engineering time spent trying to tune a kubeadm-managed etcd cluster to chase theoretical gains. You're paying for that toil, either directly on salary or indirectly through delayed projects. RKE2's bundled etcd might have a fixed performance ceiling, but its operational cost is near zero, which for many orgs is the more expensive variable.

So my methodology shifted from chasing raw throughput numbers to building a simple simulator that replicates our worst-case "configuration storm" - think a Terraform apply that hits fifty CRDs at once. The output I cared about wasn't max ops/sec, but the point where request latency started to impact our developers' workflows. That's the real threshold for right-sizing.


hugo


   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Oh, the quota-backend-bytes default! We hit the exact same thing during a major Helm history prune. It's that kind of hidden limit that makes RKE2's simplicity a double-edged sword.

Your point about needing deeper etcd expertise for kubeadm tuning is the real hidden cost. It's not just knowing the flags, it's knowing *when* to change them without causing instability. We learned that the hard way after tweaking snapshot-count and causing a spike in memory usage during compaction.

Curious, for your baseline throughput testing on c5.large, did you hit any network-related throttling between the control plane nodes that became the actual bottleneck, rather than the etcd configuration itself?


Happy customers, happy life.


   
ReplyQuote
Page 2 / 2