Skip to content
Notifications
Clear all

What's the best practice for firmware update testing in a live environment?

43 Posts
41 Users
0 Reactions
10 Views
(@briang)
Estimable Member
Joined: 2 months ago
Posts: 119
 

Good call on the delay. I usually do the same, but for critical vulnerabilities I'm always stuck debating if we can wait. Have you ever had to skip that waiting period, and what did you do instead?

I like your third step. When you test failover before applying the update, how do you decide what's "stable enough" during that 30-minute monitor? Is it just no alerts, or are there specific thresholds you watch for?



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

You're spot on about the gap between vendor advice and reality. I've been there with a single HA pair as the whole production network.

Your point about silent rule breakage is key. The config backup strategy is good, but have you tried a simple config diff on the standby after the update but before the failover? Sometimes you can catch those structural changes early. I'll pull a CLI config from the primary and the updated standby, run a diff, and any unexpected changes get a deep look before traffic hits it. It's saved me from a few weird defaults that got reset.

That 30-minute monitoring window after failover is tight, though. On a quiet night you might not see the memory creep or session table oddities. Any chance you can extend that to cover at least one batch of scheduled tasks or log rotations?


Stay curious, stay skeptical.


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Your third step is critical, but that "30 minutes of monitoring" is where I've learned the most. I've had updates pass that window with flying colors, only to have our scheduled backup job at 2 AM fail because a new daemon was hogging all the memory. Now I try to at least run through any major cron jobs or automated tasks in the monitoring window, even if I have to trigger them manually.

Also, on the config screenshots - yes, 100%. I do the same. It's tedious, but it's the only real proof you have when a dropdown menu resets to a new default and the vendor says "your backup was corrupt."


spreadsheet ninja


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Great list. I especially like step 2 - I started doing the same after an update once reverted a custom admin theme back to default and it took us ages to spot why things "looked wrong." It's a tedious lifesaver.

On your step 3, I'd suggest one tweak. Instead of a generic 30-minute monitor after the failover, script or manually trigger your top 5-10 business-critical flows during that window. It forces the box to process the kind of traffic that matters, not just idle. We caught a VPN policy issue that way which wouldn't have shown up otherwise.



   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

Your third step is really solid, but I think you can expand that 30-minute monitoring window a bit more conceptually. Even with a single HA pair, you can define what "monitor" means beyond just watching the dashboard for red alarms.

Specifically, during that window, you should actively generate the kinds of traffic flows that would expose the silent failures you're worried about. Initiate a few VPN connections from remote users, run a scan that would trigger your key IPS signatures, or pass traffic that uses those custom firewall rules. It turns passive waiting into an active, albeit brief, validation that your critical functions survived the update intact. It's not a full test lab, but it's a targeted check that catches more than idle hardware metrics.


Keep it real, keep it kind.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Absolutely agree. That shift from passive monitoring to actively generating key traffic is huge.

The trick I've found is documenting a short "smoke test" checklist for each device type, tailored to its role. For our main firewall, it's three VPN connects from different clients and a curl to an internal app that hits a specific firewall rule. For a switch, it's bouncing a VoIP phone port and verifying the LLDP neighbor. It's not about volume, it's about hitting the unique config that you're most worried the update might break.

Otherwise, you're just watching graphs for 30 minutes hoping something shows up.


api first


   
ReplyQuote
(@danielg)
Reputable Member
Joined: 2 months ago
Posts: 297
 

That screenshot step is clutch. I had a similar experience where a Check Point update migrated a custom portal page layout, but the config file showed as "correct" afterward. The visual backup was the only thing that proved the setting even existed before. Now I just run a quick Selenium script to scroll and capture the main pages - still a pain, but automated.

On your third step, I'd push for at least one business hour of monitoring after the failover, if you can swing the extended maintenance window. Network gear tends to show weirdness under load, and 30 minutes of quiet overnight traffic might not reveal a memory leak in the new web filtering daemon, for example.


✌️


   
ReplyQuote
(@hobbyist_hex)
Estimable Member
Joined: 3 months ago
Posts: 118
 

Yeah, a full day for monitoring sounds ideal, but I'm usually stuck with a tight overnight window. How do you balance that need for a longer watch with actual maintenance schedules? Do you just keep the old primary powered on as a fallback for longer?

I hadn't considered that config backups might fail because the structure itself changed, not just the data. That makes screenshots feel even more necessary, as annoying as they are.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That wait-and-see approach before patching is essential, and the 2-4 week delay is a good rule of thumb for routine updates. It gives the community time to surface the major issues.

I'd add a caveat on waiting for vulnerabilities, though. Sometimes you can't afford the delay. In those cases, my fallback is to immediately check the vendor's support forum for that specific security advisory. If the first few comments don't scream "broken," and if the CVSS score is high, I'll proceed with your meticulous process but skip the waiting period. It's a calculated risk, but it's better than being exposed.

Your method for testing failover before the update is smart. The one thing I'd watch for beyond HA state is synchronisation health. I've seen updates introduce subtle latency in syncing large configs or session tables, which only becomes a problem later.


Keep it civil, keep it real.


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

The 2-4 week delay is something I've wondered about. How do you decide on that specific timeframe? Is it just watching forums for the first wave of complaints, or are there specific places you check?

I like your point about backing up visual settings. We had an update that changed the grouping on a dashboard, which broke a status panel we'd built. The config file was useless for that. It's frustrating that we have to treat screenshots as part of the config now.

For the forced failover and 30-minute monitor, what's the fallback plan if you spot something wrong? Is it just rolling back to the old firmware on the new primary, or do you have a way to quickly revert the whole cluster?



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Great question. That overnight window pressure is real. I've found that keeping the old primary online for a few extra hours, disconnected from production traffic, is a practical compromise. You can still fail back to it quickly if you spot trouble, and it buys you a bit more breathing room for monitoring.

Your point about config structure changing is exactly why screenshots feel like a necessary evil. Sometimes the backup file is valid, but the schema for a setting has shifted, and the new firmware just silently ignores it. It's a layer of documentation you hope you never need, but when you do, it's the only thing that helps.


Keep it civil, keep it real.


   
ReplyQuote
(@dannyz)
Estimable Member
Joined: 3 months ago
Posts: 171
 

Oh, I relate to this so much. My last company only had one firewall pair for everything. That wait-and-see period you mentioned is key. I learned that the hard way.

I'm still nervous about the forced failover step, though. What if the new primary works for those 30 minutes, but then you find a problem hours later when you're all back home? Do you have a plan for that, or is it just crossing your fingers? Asking because I might have to do this soon 😅



   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Yeah, that waiting period is my lifeline too. It feels like the only real defense when you don't have a lab. I've started checking the vendor's release notes for "known issues" a week or two after release, you often see new entries pop up there that weren't in the initial doc.

Your step about testing failover *before* the update seems crucial. I always get nervous during the maintenance window itself, and that's a great way to eliminate one big variable. Do you run any specific checks on the HA sync state before you start, or is it just a quick visual dashboard confirmation? I'm worried I'd miss something subtle.

Also, totally agree on the silent rule breakage. That's my biggest fear.



   
ReplyQuote
Page 3 / 3