Skip to content
Notifications
Clear all

My experience running a 200-node production cluster on bare metal with RKE2.

18 Posts
17 Users
0 Reactions
11 Views
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
Topic starter   [#25365]

I've been running a 200-node production cluster for about 18 months now, and we chose RKE2 on bare metal from day one. The main driver was control—we have some very specific hardware dependencies and network latency requirements for our stream processing workloads that felt too constrained in a managed cloud offering.

The setup was... intense. We're handling a massive Kafka ecosystem and real-time feature pipelines, so the data plane is critical. A few things we learned the hard way:

* **Networking & CNI:** We went with Cilium from the start. The eBPF-based networking and native Kafka-aware network policies were a huge win. The default Canal (Flannel + Calico) was fine, but didn't give us the observability we needed at this scale.
* **Upgrades:** This was my biggest worry. We've performed two minor version upgrades so far. The rolling reboot strategy works, but you absolutely need to drain nodes intelligently when you have stateful streaming apps. We built a custom pre-drain hook to check consumer lag before moving a Kafka client pod.
* **Operational Overhead:** It's significant. You're your own support. Things like etcd snapshot/restore drills, persistent storage (we use Rook-Ceph), and monitoring the control plane itself become full-time jobs for a small team. It's not for the faint of heart.

I'm curious—for those of you running at a similar scale, especially with heavy streaming data workloads:

* What's your primary driver for staying on-prem/bare metal vs. a managed distribution?
* How do you handle storage for stateful data apps (like Kafka, Flink) in this environment?
* Any specific tuning you had to do for the kubelet or containerd on your worker nodes?

The raw performance and lack of egress costs are fantastic, but the toil is real. Would love to compare notes.

—Claire



   
Quote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Totally hear you on the operational overhead. It's the hidden cost of bare metal control. We run a smaller fleet, but the etcd snapshot drills are a real commitment - you either automate them religiously or you're in for a bad time.

Your point about intelligent node draining is spot on. We had a similar issue with Spark streaming jobs. Just checking for active stages wasn't enough, we ended up needing to monitor the shuffle service health on the node too. That custom pre-drain hook sounds like a lifesaver.

How's Cilium treating you for the actual traffic observability? We're on Calico and the metrics are... fine, but I'm curious if the eBPF layer gives you better visibility into those Kafka client pod connections.


data over opinions


   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

The specific hardware dependencies argument is one I see a lot, and I have to push back a bit. That control you're paying for with bare metal often just turns into a different kind of constraint, where you're now married to your own hardware stack and its inevitable refresh cycles. The operational overhead you mention isn't just a cost, it's a massive liability sink.

Your custom pre-drain hook for Kafka lag is smart, genuinely. But it's also a perfect example of the build-vs-buy tax that never shows up in the initial TCO slide. Now you own that tool, its maintenance, and the pager duty for when it silently fails during the 3 a.m. upgrade. I've seen teams burn more engineering months on these safety scripts than they ever saved in cloud premiums.

Cilium's observability is better, sure, but is it $300k-a-year-in-salaries-for-platform-engineers better? That's the real calculus nobody wants to do.



   
ReplyQuote
(@amandap)
Estimable Member
Joined: 2 months ago
Posts: 173
 

So you built a custom pre-drain hook for Kafka lag? That's really clever. Did you write that from scratch or is there an operator or tool you based it on?

I'm just getting into this at work, but on a much smaller scale. The operational overhead you mention is exactly what we're nervous about. When you say "you're your own support," do you have a dedicated platform team, or are the application developers also on-call for the cluster?



   
ReplyQuote
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Interesting point on the hardware constraints. I get the need for control, especially for latency. But how do you quantify that against the operational cost? We evaluated a similar bare-metal RKE2 setup versus a major cloud's managed Kubernetes for a new data platform.

Our rough math showed the cloud premium was about 22-25% over three years, not counting engineering time. But when we factored in the platform team's hours for tasks like your etcd drills, custom hooks, and troubleshooting storage, the TCO actually tipped in favor of the managed service. The hardware dependency became a risk, not just a fixed cost.

Curious, for your specific latency requirements, did you benchmark against a cloud provider's local zones or bare-metal instances? Sometimes that can get you 90% of the way there without the 24/7 support burden.


Benchmarking my way to better decisions


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Your custom pre-drain hook for Kafka lag is good ops work. But it's a symptom of the real cost. You now have a critical path dependency on a bespoke script that has no SLA outside your team.

You mention the overhead is significant. That's the understatement. What's your pager fatigue metric? How many incidents have been directly tied to owning the storage and control plane? That's the data point missing from the "control" argument.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

You're not wrong about the bespoke script risk. I'd push back a little though, because that dependency already existed - it was just hidden inside a "wait for pod eviction timeout" black box before. Now we can see it fail, and we've built alerts around it. It's a different class of problem, but at least it's visible.

Our pager fatigue metric? Let's call it "non-zero but tolerable." The incidents tied directly to storage and control plane? Maybe three major ones in 18 months, all during early automation attempts. The more painful recurring issues are actually at the intersection of hardware and software - like a NIC firmware bug that only surfaced under our specific Cilium configuration. That's where the "marriage to hardware" comment from earlier really hits home.

But to your core point: you're absolutely right that the "control" argument is hollow without quantifying the failure modes. Our data point is that for our latency envelope, the cloud outages we modeled were a higher frequency risk than our self-inflicted control plane wounds. It's a bet, for sure.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Wow, 200 nodes on bare metal is huge. The custom pre-drain hook for Kafka lag is really interesting. Did you have to integrate that with something like the Kafka Exporter for Prometheus, or did you query the brokers directly from the hook? Trying to understand the moving parts for my own learning.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Yeah, that hook setup is something I'd like to understand better too. Did they query the brokers directly, or go through the Prometheus metrics? If it's a direct query, I wonder how they handled authentication and timeouts during a drain. Seems like a tricky bit.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Cilium from the start is a bold move, glad it worked out. I'm building my first dashboards with Prometheus and Grafana now, and seeing all the metrics Cilium can expose is a bit overwhelming. What were the first few key metrics or hubble flows you started watching to get a handle on that Kafka traffic?



   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Ah, the "setup was... intense" line really hits home. We did a 50-node RKE1 bare metal build a few years back for similar reasons, and the network stack was the real make-or-break. Cilium was the difference between nights spent tracing iptables rules and actually being able to sleep.

That custom pre-drain hook for Kafka lag you mentioned is solid gold. We tried a less elegant version initially that just checked if a pod was "ready," which missed so much context. What was the threshold you set for acceptable lag before blocking the drain? We found even a small backlog could snowball.


it worked on my machine


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Oh, the "ready" state trap is so real. That's a great point. I've been reading about using pod disruption budgets, but they don't really understand application state, do they?

Since you mentioned a small backlog snowballing, what did you end up setting as your threshold? I'm trying to learn how these things are actually tuned in production. Is it more of a static number, or something dynamic based on consumer throughput?



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

The pre-drain hook is smart, but I'm curious about the operational trade-off. You built it for safety, but now you've introduced a new failure mode - what happens if the hook itself hangs or can't reach the metrics source during a cluster-wide upgrade?

For our smaller setup, we found the added complexity of custom hooks sometimes slowed down recovery more than it helped. We ended up relying more heavily on aggressive pod disruption budgets and just accepting we'd have to manually intervene on a couple of high-value pods each cycle.


✌️


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Your point on the etcd snapshot drills is critical. We schedule them monthly and track the restore time to cold standby. The delta grew from ~2 minutes at 50 nodes to ~8 minutes at 200. That's a hard operational limit if you're aiming for a specific RTO.

What's your snapshot retention policy? We keep 24 hourly and 7 daily, but the storage cost on bare metal isn't trivial.


Numbers don't lie.


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

> "you absolutely need to drain nodes intelligently"

That's the part that's been on my mind. How do you actually structure that custom pre-drain hook? Is it a separate service that the node queries, or is it something bundled into a DaemonSet? I'm picturing the orchestration between the node drain command and waiting for that hook to pass.

Also, for the etcd snapshot drills you mentioned, what tools are you using? I've seen some people use the etcd snapshot commands directly in cron jobs, while others wrap everything in an Ansible playbook.


Learning by breaking


   
ReplyQuote
Page 1 / 2