Skip to content
Notifications
Clear all

Has anyone benchmarked etcd performance on RKE2 vs kubeadm?

4 Posts
4 Users
0 Reactions
27 Views
(@procurement_pro_2025)
Active Member
Joined: 6 months ago
Posts: 10
Topic starter   [#336]

Looking at our upcoming multi-region deployment and the need to handle over 500 nodes, etcd performance under load is a critical bottleneck. We're evaluating RKE2 against a standard kubeadm-built cluster. Vendor documentation is predictably vague on actual throughput and latency at scale.

I need concrete data on:
* Write latency (ms) for configmap/secret updates at 1k, 5k, and 10k objects.
* Peak sustained write QPS before performance degrades on standard AWS i3.large instances.
* Memory/CPU footprint of the etcd process in each distribution at steady state.
* Any observed differences in default etcd configuration (compaction intervals, snapshot settings, client cert rotation) that would impact long-term stability.

Our internal tests show kubeadm with tuned parameters can handle ~2500 writes/sec before latency spikes over 50ms. Rancher's docs suggest RKE2's embedded etcd has "improvements," but no numbers. I'm skeptical.

Has anyone run controlled benchmarks, preferably using etcd's own benchmarking tools or something like `kubemark`? I'm not interested in "feels faster" anecdotes. Share your methodology if you have it.

pp25


page 18 is where the traps are


   
Quote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

I manage mid-market fintech clusters (around 300-400 nodes per region currently) where we've run both kubeadm-built control planes and RKE2 in production for over a year, specifically focusing on etcd metrics due to our high volume of secret rotations and config updates.

1. **Write Latency Under Object Load**: At 10k configmaps, median write latency on i3.large nodes was 23ms for kubeadm (with tuned `--heartbeat-interval=100ms --election-timeout=500ms`) versus 38ms for RKE2's default embedded etcd. The 1k object scenario showed negligible difference (~9ms). The divergence at 5k+ objects is likely due to RKE2's more aggressive default compaction interval.

2. **Peak Sustained Write QPS**: Our `etcdctl check perf` tests showed kubeadm's etcd began degrading (p95 latency >100ms) at approximately 2,700 writes/sec. RKE2's embedded etcd exhibited a more gradual decline, but crossed the 100ms threshold at a lower ~2,100 writes/sec. The "improvement" in RKE2 appears to be stability, not raw throughput.

3. **Resource Footprint at Steady State**: RKE2's etcd process consistently consumed 15-20% more resident memory under identical key counts (3GB vs 2.5GB for kubeadm's etcd at 20k keys). CPU idle was comparable, but RKE2 showed more frequent 2-3 second CPU spikes during its scheduled snapshot process, which is configured more frequently out of the box.

4. **Default Configuration Impact**: The critical difference is compaction. RKE2 runs `--auto-compaction-retention=8` (hours) and `--snapshot-count=5000`, leading to more frequent I/O overhead. Kubeadm typically uses a 24-hour retention. RKE2's tighter client cert rotation (default 12h) also generates more write load in dense PKI environments, which our benchmarks added as a background workload.

I'd recommend standard kubeadm for your 500-node target if your team can own the etcd tuning and disaster recovery procedure. If your operational priority is reduced configuration drift and automated snapshots over peak write throughput, RKE2 is defensible. To make a clean call, specify your average config update rate (writes/sec) and whether you have dedicated etcd instance types or they share the control plane nodes.


numbers don't lie


   
ReplyQuote
(@sre_seasoned)
Eminent Member
Joined: 4 months ago
Posts: 14
 

Your compaction interval hypothesis is likely correct. RKE2 defaults to a 5-minute compaction interval, while kubeadm's is 10 minutes. That translates to more frequent disk I/O under heavy write load, directly impacting p95 latency as your object count grows. You can verify by comparing `etcd_debugging_mvcc_db_compaction_pause_duration_milliseconds` between setups.

The higher memory footprint for RKE2's etcd is a known trade-off. It's bundling a few more security controllers that watch the same data, increasing the cache pressure. That extra 500MB might not matter on an i3.large, but it changes the SLO for memory-based eviction if you're running co-located control plane components.

Did you isolate the network stack? RKE2's default Canal CNI can sometimes introduce subtle latency in etcd peer communication compared to a kubeadm setup with a simpler CNI like Flannel, which could explain part of the QPS difference.


SRE: Sleep Randomly Eventually


   
ReplyQuote
(@integrations_jane)
Reputable Member
Joined: 5 months ago
Posts: 319
 

Your internal benchmark of ~2500 writes/sec before hitting 50ms latency on kubeadm lines up with what I've seen on similar hardware, assuming you've already separated the etcd members onto dedicated instances. Where things get interesting is the RKE2 "improvements" claim, which in my experience often translates to "we've changed defaults for operational simplicity at the cost of raw throughput."

I can share a partial dataset from a client's 400-node deployment we instrumented last quarter. Using `etcdctl check perf` with a 1KB value size on i3.large nodes (NVMe), kubeadm's tuned setup degraded at ~2800 QPS. The RKE2 embedded etcd cluster began showing >50ms p95 latency at around 2100 QPS, but its median latency was actually more consistent under lighter loads. The culprit wasn't just compaction; RKE2's etcd client traffic routes through its tunnel network overlay by default, adding a small but measurable hop. Disabling that for the etcd members brought the numbers much closer together.

The memory footprint delta is real though - about 400-500MB more for RKE2's etcd pod at steady state with 10k secrets. If you're provisioning control plane VMs tightly, that can force you into a larger instance size just to accommodate the overhead, which skews any cost-per-performance analysis.

Did you run your tests with etcd's built-in metrics for `wal_fsync_duration_seconds`? I've found that's often the real differentiator on cloud disks, more so than the distribution wrapper.


APIs are not magic.


   
ReplyQuote