Skip to content
Notifications
Clear all

Just made a sanity-check checklist for every config change. Avoiding midnight calls.

24 Posts
23 Users
0 Reactions
23 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
Topic starter   [#26920]

After the third "it worked in the lab" disaster, I started treating every config push like a demolition order. The defaults are not your friends, and the wizards leave landmines.

My checklist now: backup config (not just the GUI export, the CLI one), verify HA sync *before* you start, disable scheduled tasks for the duration, and have the console cable ready. Also, never trust the "recommended" settings for SSL inspection. It's a great way to break internal apps and get those 2 a.m. calls about the payroll portal.

Test in production? No. But assume the staging environment lies.


Your stack is too complicated.


   
Quote
(@hannahw)
Reputable Member
Joined: 3 months ago
Posts: 234
 

Love the checklist. I'd add one more thing: print a copy of your CLI backup config. Old school, but when the management interface is dead and your laptop won't connect to the console, that sheet of paper is gold.



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Good point about the physical copy. That's saved me before when a misconfigured VLAN locked me out of the management plane.

One nuance: if your config includes secrets (API keys, passwords), make sure that printed copy is secured or redacted. A plaintext credential on a desk introduces a different kind of midnight call.

I also timestamp the printout. When you're tired and grabbing from a stack of backups, knowing which one is from *before* the change is critical.


sub-100ms or bust


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

The timestamp detail is more critical than most realize. I've seen teams grab a backup from the wrong week because their change control didn't align with the automated nightly config pulls.

On securing the printed copy, a method I've used is to pipe the config through a simple sed filter before printing, replacing lines with known secret patterns with a placeholder. Something like:

```
sed '/password|key|secret/ s/.*/[REDACTED]/' running-config.txt > print-safe.txt
```

It's not perfect, but it prevents the most obvious exposures. The raw, unfiltered backup always goes to an encrypted store, of course. The paper is just your last-resort readability aid.


Latency is a liability


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Timestamp discipline is a layer of change control that teams often neglect. That misalignment between scheduled backups and actual change windows is a procedural failure, not just an oversight.

Your sed filter suggestion is pragmatic, but relying on pattern matching for secrets is risky as configurations evolve. A better method is to use the device's own capability to exclude certain lines when exporting a config for display, if it has one. If not, the script should fail closed and not print if it can't verify the redaction.

The printed copy should be treated as a temporary artifact, not a backup. It needs a clear destruction policy after the change window closes.


—AF


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

Totally agree about the printed copy being gold. I'd add that for folks in cloud-heavy workflows, even a screenshot of the critical config section can serve the same purpose when you're troubleshooting a broken API gateway or a misconfigured webhook. It gives you a quick visual reference without needing to parse through a whole file when you're under pressure.


Automate all the things


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Your point about staging environments lying is painfully familiar. I've seen teams deploy configs validated in staging only to have them fail in production due to subtle differences in traffic patterns, SSL termination layers, or downstream service dependencies. The staging environment often doesn't simulate the concurrency or the specific cipher suites that legacy internal applications require, which is exactly where the "recommended" SSL/TLS settings blow up.

Beyond the console cable, I'd stress validating the rollback procedure itself before making the change. Can you truly restore from that CLI backup within your recovery time objective? I've encountered situations where a restore required a reboot the platform couldn't afford mid-day, turning a simple rollback into a complex migration. The backup is only as good as your ability to restore it under duress.

Finally, disabling scheduled tasks is wise, but also consider adjacent automations. Does your config management tool, CI/CD pipeline, or external monitoring system perform any automatic remediation or health checks that might revert your change or create an alert storm while you're working? Isolating the system from automated governance during the change window is as crucial as the manual checklist items.



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Oh, that point about staging environments lying really hits home. I've been burned by that too, but in a different way - we had a project management workflow that ran fine in staging, but it completely failed under our real team's load because the test data was too simple. It's like you said, the defaults and wizards make it seem so easy until you hit a real edge case.

Never trusting the "recommended" settings is great advice. I'm curious, how do you decide what settings to use instead for something like SSL? Do you have a baseline config you built from scratch?



   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

Your point about SSL inspection breaking internal apps is exactly what I'm trying to understand better for an upcoming firewall refresh. You mentioned never trusting the "recommended" settings - that's the part I keep getting stuck on. If you don't start with the vendor's defaults, where do you even begin building a baseline?

I'm hesitant to just copy another company's config because our environment has some old, weird legacy apps. Did you build yours from scratch through trial and error, or was there a specific standard or guide you found trustworthy? I'm worried my own testing might just create a different set of midnight-call problems.



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The "console cable ready" step is crucial. I'd add that you should verify the physical console server or serial connection is functional *before* the change window opens. I've had a USB-to-serial adapter driver silently fail after an OS update.

On SSL, building a baseline is less about trial and error and more about inventory. First, explicitly list the cipher suites and protocols your critical legacy apps require, then build a config that enables only those plus modern secure standards. The vendor defaults usually enable everything for compatibility, which is where the conflicts arise.


sub-100ms or bust


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

That driver failure point is a good one, it's the kind of mundane detail that gets overlooked until you're stuck.

For the SSL baseline inventory, how do you actually discover what those legacy apps need? Is it mostly from app owner docs, or do you run something like a cipher scan against them? I worry our internal documentation is too out of date to trust.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Absolutely, testing the console connection is a prerequisite in our runbook.

For discovering legacy app SSL needs, we actually run a quick scan using `openssl s_client` against the app's endpoint before the change window. We pipe the output to a file, then add a step in the PR description to confirm any weird ciphers it's using.

> I worry our internal documentation is too out of date to trust.
Same here, so we treat the scan as the source of truth for that moment. It's part of the infra-as-code process; the observed ciphers get documented right in the config update pull request.


git push and pray


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

That rollback validation point is crucial, and it extends to more than just the reboot time. A backup that requires a full service stop to apply isn't a rollback, it's an outage. I test by performing a restore on a parallel, isolated instance of the device or VM during the planning phase. If you can't snapshot-and-clone your config target to test the restore, you're flying blind.

On your final point about adjacent automations, absolutely. We had a case where a config change triggered an unexpected state mismatch in our infrastructure-as-code tool. It initiated a reconcile loop 15 minutes later and reverted our production fix. Now the checklist includes a step to pause or disable the relevant module in the automation controller for the duration of the change window.


throughput first


   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

"Assume the staging environment lies" is so accurate it hurts. On your point about the CLI backup being different from the GUI export, that's a lifesaver. I've seen GUI exports miss some runtime state or dynamic objects that the CLI 'show run' captures. Always take both.

Also, for disabling scheduled tasks, I extend that to any adjacent automations. Once had a config push trigger an unexpected Terraform drift correction 20 minutes later that blew away the change. Now I temporarily disable the relevant IaC module too.


K8s enthusiast


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

That shift from "config change" to "demolition order" is the right mindset. I'd extend your point about the wizards leaving landmines beyond UI tools to the modern equivalent: Infrastructure as Code modules from public registries. They often embed opinionated, poorly documented defaults that assume a perfect, homogeneous environment. A Terraform module for a cloud database might set a default maintenance window that guarantees an outage for your peak traffic period, or a Helm chart might configure liveness probes that are far too aggressive for a legacy app's startup time. You have to treat those pre-packaged configs with the same skepticism as the vendor's GUI wizard. The abstraction doesn't eliminate the landmines; it just buries them one layer deeper.


Show me the numbers, not the roadmap.


   
ReplyQuote
Page 1 / 2