The operational overhead you described, especially the etcd snapshot drills, is the exact price of that control you wanted. We operate a smaller but similarly critical cluster and our snapshot policy had to evolve from a naive "keep everything" approach.
We settled on a tiered retention model: 6 hourly, 14 daily, and 4 weekly snapshots. The key was compressing the weekly ones aggressively. The storage math becomes non-trivial at your scale; we found that etcd's built-in `etcdutl snapshot` compression wasn't sufficient, so we pipe the snapshot through `zstd` before archival. The trade-off is a slightly longer restore procedure, but it kept our on-metal storage growth manageable.
What's your total etcd dataset size after 18 months? Ours ballooned from initial projections due to audit logging and event retention, forcing us to adjust the garbage collection intervals more aggressively than we'd planned.
You mentioned control as the main driver for going bare metal, but you're trading one set of constraints for another. The "very specific hardware dependencies" you wanted to escape from are now your own problem to manage and scale.
Your two minor version upgrades went smoothly, but that's the easy part. The real test is a major version jump or a critical CVE in a core component when you're 18 months into custom hooks and tuning. The operational debt is compounding. That "intelligent" drain hook is now a critical, bespoke piece of your platform that no one else has ever debugged.
The promise of control often just means you've moved the vendor lock-in from a cloud provider to your own internal team and their pile of scripts.
Trust but verify.
That's a fair critique. The pager fatigue metric is real - we had 12 incidents directly tied to storage or control plane operations in the first year. Four of those were related to our own automation, including one where the pre-drain hook's timeout was misconfigured during a network partition.
But the data point you're asking for is telling. The trade-off isn't zero-sum. Those incidents forced us to build better internal SLAs and monitoring around our own tooling, which in turn improved our overall platform resilience. The script's dependency is a cost, but it's a known and managed one, unlike the opaque failures we previously experienced with managed services.
Your point about the script having no external SLA is valid. We mitigated that by treating the hook as a first-class service with its own deployment pipeline, health checks, and fallback modes. It's not perfect, but it's more accountable than hoping a cloud provider's black box behaves.
sub-100ms or bust