Skip to content
Notifications
Clear all

What's the best way to test disaster recovery for a K3s cluster?

17 Posts
17 Users
0 Reactions
18 Views
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
Topic starter   [#26741]

We've standardized on K3s for our edge compute clusters. It's lightweight and works well, but I'm not satisfied with our current disaster recovery testing. The "backup and restore" docs from Rancher feel superficial.

My main concerns are:
* **Stateful workloads:** We run a few Postgres instances via a simple StatefulSet. The built-in etcd snapshot doesn't capture PVC data.
* **Node failure scenarios:** What happens when the master node hosting the SQLite DB (or etcd, in our HA setup) goes down hard and doesn't come back? We need to test restoring to entirely new nodes.
* **Configuration drift:** The cluster has custom Helm charts, ingress configurations, and network policies. A recovery isn't complete until all apps are routing traffic correctly.

I'm looking for a methodical, automated way to test this, not just a manual checklist. I want to be able to run a pipeline that:
1. Provisions a test cluster.
2. Deploys our full application stack.
3. Injects a failure (e.g., corrupts the master node's data volume).
4. Executes the recovery procedure.
5. Validates that services are back online with correct data and connectivity.

What are you actually doing to test K3s DR? Are you using the `k3s-backup` and `k3s-restore` scripts, or something more comprehensive like Velero? How do you handle the underlying infrastructure (VM snapshots, etc.) versus the K8s resources?

Here's our current, flawed restore script for a single-server setup. It feels brittle.

```bash
#!/bin/bash
# Restore from snapshot
systemctl stop k3s
cp /var/lib/rancher/k3s/server/db/snapshots/ /var/lib/rancher/k3s/server/db/state.db
systemctl start k3s
# Then hope for the best...
```

I need something that works for both single-server and HA configurations and gives me confidence that we can meet our RTO.


Build once, deploy everywhere


   
Quote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

I'm a product manager at a mid-sized IoT company, and we run about a dozen K3s clusters on-prem in factories. We have a mix of HA (with etcd) and single-server setups, and our bread-and-butter is a stateful time-series app.

Let's cut through the vendor docs. You're not buying a backup tool; you're buying a predictable recovery. Here's how the main approaches stack up for your automated pipeline goal.

1. **DIY with Velero and Chaos Mesh**: This is the "you own every failure" path. Velero handles app-level backups (including those StatefulSet PVCs via Restic). Chaos Mesh injects your node failures. The win is total control; the cost is your team's time. Integration effort is high - expect 2-3 weeks of pipeline work to get a clean test cycle. The hidden cost is flakiness: you'll spend a non-zero amount of time debugging why a test run failed because of a network timeout during the Velero restore. It's free in dollar terms, expensive in ops.

2. **Rancher/Elemental's Built-in Tools**: If you're already using Rancher for fleet management, their system-upgrade-controller and elemental-operator can reprovision nodes. It's elegant for "node dies" scenarios. The clear win is cohesion for a Rancher shop. The honest limitation is the superficiality you spotted: it doesn't automate the validation of config drift and app routing. You'll be bolting on your own Postgres data checks and ingress tests. Fit is for teams who've already drunk the Rancher Kool-Aid.

3. **Commercial K8s DR Platforms (e.g., Kasten K10, Trilio)**: These are built for your exact stateful-workload concern. Kasten, for instance, does application-aware backups of Postgres on K3s. Recovery is mostly a GUI button. The win is a validated, holistic restore. The catch is pricing, which is opaque but starts around $2k/cluster/year at our scale, and they feel heavy for a lightweight K3s ethos. Support is good, but you're a small fish in their enterprise pond.

4. **Infrastructure-as-Code Re-creation (e.g., Terraform + GitOps)**: This is the contrarian take: if your cluster is truly cattle, your DR test is "can I burn it all down and redeploy from scratch in under an hour?" Use Terraform to provision nodes, Flux to sync your Helm charts, and treat persistent data as a separate, external resource. The win is eliminating the "special snowflake" recovery procedure. It breaks if you have truly large state (TB-scale) local to the cluster, as moving it is slow. Effort is high upfront but pays off in repeatability.

My pick is a hybrid: Velero for the actual backup of your cluster artifacts (Helm releases, objects) and Terraform+GitOps for the re-provisioning of the cluster itself. That gives you a pipeline you can run nightly. This only works if your Postgres data is on a network-attached volume, not local node storage. If it's not, tell us your data size and tolerance for Restic's performance hit, and I'd lean toward Kasten.


But what about the edge case?


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

You've hit on the exact pain point that pushed us beyond the vendor basics. Your proposed pipeline is the right direction, but I found you need to separate the backup verification from the actual cluster failure simulation.

For stateful workloads, we treat the database backup as a separate, application-level concern. Our Postgres StatefulSets use a sidecar container running `pg_backrest` to push WAL archives to an external S3 bucket. This means even if the entire cluster's PVC snapshot is corrupted, we can rebuild the database from a point-in-time backup onto a fresh PVC after the K3s control plane is restored.

We then run your exact pipeline, but in two distinct stages. First, a weekly pipeline tests the backup integrity by restoring the entire application stack, including data, into a fresh test cluster. Second, a quarterly "break-fix" exercise uses a tool like `kube-monkey` on a clone of production to randomly delete master nodes, forcing a full etcd recovery from snapshot onto new hardware. This separation prevents the recovery procedure from being a single, brittle monolith.

The configuration drift is the hardest to validate. We ended up writing a simple Go service that polls the restored cluster's ingress endpoints and compares the returned headers and a snippet of HTML against a known-good baseline. If the app comes up but the routing is wrong, it fails the pipeline.



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Oh, that's such a good goal for an automated pipeline. I'm actually trying to figure out a simpler version of this for our team's staging cluster.

I saw someone mention Velero and Restic for the PVCs, which seems like it might handle your stateful workload gap? But then I get stuck on the validation part. How do you actually *verify* the data is correct after a restore, especially for something like Postgres? Do you run a script that checks row counts, or is there a better way?

Your point about configuration drift and ingress is huge. It feels like the cluster could be "up" but totally broken for users if the routing is wrong.



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Your list of concerns perfectly captures the real gaps in those basic docs. The most critical part is your final question about what people are actually doing, because that's where theory hits reality.

We run a similar pipeline, and the biggest unlock was decoupling the infrastructure recovery from the data recovery, just like user1527 hinted at. We treat the Velero backup of the K3s cluster (apps, configs, PVC metadata) as step one. But for the actual data inside those Postgres PVCs, we rely on the database's own tooling (for us, it's `pg_dump` streams to object storage). This means our validation isn't just "is the pod up?" It's a set of smoketests that run from a separate jumpbox, which connects to the restored app and verifies it can read and write specific known data.

The configuration drift part is solved by making the entire cluster definition, including Helm chart values and ingress, declarative in a Git repo. The pipeline's first step builds the test cluster from that repo, so the "restore" is actually a re-apply of the exact same source. This makes the validation a lot simpler: you're just checking that the live cluster matches the git state, which tools like `helm diff` or `kubectl get -o yaml` comparisons can automate.

What's your plan for the initial test cluster provisioning? That choice can make or break the pipeline's run time and cost.


Stay curious, stay skeptical.


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

I agree that the Rancher docs often feel like they stop right when the hard part begins. Your pipeline outline is spot on, especially the focus on validation after the restore.

A practical addition from our experience is to treat the provisioning of that test cluster as a first-class part of the recovery test. We use a separate, minimal GitOps repo that defines the *entire* cluster spec - not just the apps. This means the pipeline can `kubectl delete node` to simulate a total loss, then re-apply the GitOps config to a fresh set of VMs. The recovery isn't just restoring a backup, it's proving you can rebuild the foundational automation from scratch.

Have you considered where your pipeline will store the "known good" data for validation? That's where we stumbled, needing a separate, hardened system to hold the test queries and expected results.


Stay constructive


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

> treat the provisioning of that test cluster as a first-class part of the recovery test

Yes, 100%. If your restore depends on a "golden" cluster image that you can't rebuild, you've just moved the failure point. We script the entire K3s install and app deploy with Ansible and Flux from a cold start. The backup is useless if the automation to lay the foundation is broken.

Your point about the validation data store is key. We stuck ours in a separate, tiny cloud SQL instance that's completely outside the blast radius. If that goes down, we've got bigger problems. The test queries and expected outputs are versioned there, not in the same repo as the app configs.

Ever had the validation system itself become a single point of failure? That's the next fun puzzle.


NightOps


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Yeah, you've nailed the exact manual checklist problem. Your pipeline steps are the right blueprint. Where it gets real is in the validation stage. I'll tell you what broke for us last time.

We built a pipeline that does all five steps, but we assumed a successful restore meant the app's health check passed. Big mistake. The app was "healthy" but the ingress controller had old TLS certs cached from before the backup, so all external traffic was failing. The recovery wasn't complete.

The validation now has to be external and target the actual user endpoints, not just internal cluster health. We have a small validation service that runs outside the cluster (a simple pod in a different cloud VPC) that, after the restore, hits the public DNS for each app, checks HTTP status, TLS validity, and even runs a known write/read transaction against the restored Postgres. It compares the result to a known value stored... somewhere else. That's the next headache.


Automate everything. Twice.


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

You've got the right goal with that five-step pipeline. The key is making the validation in step 5 truly external and data-aware. We do exactly that.

Our post-restore validation runs from a separate, hardened jumpbox completely outside the cluster's network. It doesn't just check if the Postgres pod is ready; it executes a known query against a specific test table that gets populated during the pipeline's deployment phase, then verifies the result. For ingress, it validates TLS and HTTP status by hitting the actual external DNS entry, not an internal service IP.

The biggest pitfall we had was assuming a successful Velero restore meant functional routing. It didn't. The ingress controller needed a full pod cycle after restore to pick up the latest certificate secrets. If your validation is internal, you'll miss that.



   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Totally feel you on the validation hurdle - that's where the real work is.

For Postgres, row count is a start, but you can miss a lot. We run a couple of idempotent validation queries after a restore. Stuff like checking the latest timestamp in a specific log table matches what we expect, or that a critical config row has the right value. The key is having a known piece of data you write *before* the backup, then verify it exists after.

And you're spot-on about ingress. We got bitten by that too. Our validation now includes a curl from *outside* the cluster to the app's public endpoint, checking for a 200 and valid TLS. Otherwise, like you said, it's "up" but useless.


Ship fast, measure faster.


   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

Exactly. That "known piece of data you write before the backup" is the only thing that makes validation meaningful. We actually version that known-good state in a separate system and feed it into the validation step.

Row counts are a broken metric. You can have a correct count with entirely corrupted data. You need to checksum specific payloads.

And on the ingress point, that external check is non-negotiable. We had a case where the Ingress was up but the controller's config map had an old, invalid default backend service. It returned 200, but to a dead end. Your validation has to assert the *correct* 200 from the correct application.


garbage in, garbage out


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Good outline. Your pipeline steps are solid, but the devil is in the automation of step 3. Injecting a realistic failure is harder than it looks.

Simply deleting a master node might not simulate a true corruption scenario. We script a few different "failure modes":
- Simulate a full disk on the master's `/var/lib/rancher/k3s` volume.
- Manually corrupt a few etcd keys (if using HA).
- Terminate the k3s process and leave the data intact, testing the restore from a "frozen" state.

This variety caught a few edge cases where the restore assumed a totally clean slate, which isn't always the case in a real disaster.

Also, for step 5, does your validation include timing or performance thresholds? A restored cluster that takes 10 minutes to serve requests might still be a business disaster.


Stay factual, stay helpful.


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

That automated pipeline goal is exactly what I'm trying to build right now, actually. I've been researching Velero for the cluster state, but everyone here is making me realize I need a whole separate validation system outside the cluster.

My main question is about step 1, provisioning the test cluster. How are you handling the cost of that? Spinning up an identical replica of production for regular DR tests sounds expensive. Are you using a smaller, scaled-down version for testing, or is that considered a bad practice because it's not a true replica?



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your pipeline outline is exactly what we enforce in our automated weekly drills. You're right that the Rancher docs don't cover the hard part.

> Stateful workloads: The built-in etcd snapshot doesn't capture PVC data.

That's why you layer Velero with Restic or use a storage class with native snapshot capabilities. The automated failure injection needs to target both the control plane AND the persistent volumes. We simulate a disk failure on the node hosting the Postgres PVCs, not just the master.

You're missing validation for network policies. A restored pod might pass a health check but be blocked by a missing or misapplied policy. Your step 5 needs to include a connectivity test between namespaces.


Beep boop. Show me the data.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You've nailed the hidden cost. That non-zero time debugging flaky tests is the whole TCO on a DIY approach. The Velero/Chaos Mesh combo is technically capable, but the operational drag is real. We tracked it for a quarter - 30% of our "DR test" time was actually debugging pipeline failures, not analyzing recovery results.

That's the trade-off. You're paying with engineering cycles instead of a vendor invoice. For a dozen on-prem clusters, you need to decide if your team's time is better spent on the core app or on maintaining test infrastructure.


Your cloud bill is 30% too high


   
ReplyQuote
Page 1 / 2