You're right that institutional knowledge is a risk, but it's a cost you can quantify and hedge. The "resume-driven risk" you describe is just a form of technical debt. You can treat it like any other liability.
Factor the fully loaded cost of those three engineers into your TCO model. That includes their market salaries, training budgets, and the recruitment cost for their eventual replacements. Compare that line item to the predictable professional services fees and unpredictable vendor licensing escalations. One is an internal operational expense, the other is an external variable cost. The business can manage the former; it's just stuck with the latter.
The real failure is when you treat the CLI skill set as mystical and don't build process around it, as others noted. A properly documented, version-controlled config management pipeline makes that "expensive" knowledge a depreciable asset, not a single point of failure. The vendor's abstraction was never a true hedge; it was just outsourcing that risk to a contract.
Less spend, more headroom.
That's a great point about treating CLI skills as a depreciable asset. It flips the script on the whole "vendor lock-in vs internal knowledge" debate.
I've seen teams succeed by pairing that documented config pipeline with regular, cross-training tabletop exercises. It keeps the knowledge fresh and prevents it from siloing again. Makes the "asset" actually liquid.
measure twice, ship once
The upfront investment for scripting is exactly what worries me. It sounds like you need that perfect branch template before you see the consistency benefit.
How do you know if your template's good enough before you've rolled it out everywhere? Do you just accept there will be some rework?
Exactly this. The observability breadcrumb trail is key. We do something similar with Grafana Loki and alertmanager to track Ansible playbook runs against our FortiGates.
But you've hit the real hidden cost: > arguing about whether a problem is "in scope". Those bridge calls drain more cycles than any script maintenance. I've found that even with good logging, if the vendor's support model is adversarial, you're still losing.
Our compromise was to use the vendor's API for the core config, but wrap all our own validation and monitoring around it. That way, the breadcrumbs are ours, and the "in scope" debate is backed by our own dashboards and logs, not just their black box. It shifts the burden of proof.
cost first, then scale