Skip to content
Notifications
Clear all

Anyone else's Firebox reboot randomly after 200 days of uptime?

7 Posts
7 Users
0 Reactions
9 Views
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
Topic starter   [#26600]

Hey folks, I've noticed a pattern with a few of our deployed Firebox appliances (specifically the M570 and M470 models) that I wanted to run by the community.

They run rock-solid for almost exactly 200 days, give or take a week, and then spontaneously reboot. There's no critical hardware failure, the logs just show an unexpected shutdown followed by a normal boot sequence. No obvious signs of overheating, power issues, or memory exhaustion in the logs leading up to it. We're on the latest general release firmware for our hardware.

It's consistent enough that I'm starting to wonder if it's a known, but undocumented, watchdog or maintenance process. Has anyone else encountered this specific ~200-day uptime reboot?

* Are you seeing it on specific models or firmware versions?
* Did you ever get a definitive cause from WatchGuard support?
* Any workarounds besides a scheduled manual reboot at 199 days?

I appreciate that stability is generally good, but unexplained reboots in a production environment are a concern. I'm gathering data before opening a formal support case, so any shared experiences or insights would be valuable for everyone here.



   
Quote
(@gracem)
Reputable Member
Joined: 2 months ago
Posts: 294
 

Oh wow, that's a very specific uptime to hit before a reboot. I haven't seen this exact pattern on our M270s, but we do schedule a quarterly maintenance reboot for all our edge devices anyway, so we never get near 200 days.

That said, your theory about a hidden watchdog timer feels spot on. I've seen similar "uptime ceiling" behavior in other embedded systems, where a process is designed to force a refresh after a certain period. It's rarely in the public docs.

Since you're gathering data for a support case, you might want to check if the reboots correlate with any specific log entries from *exactly* 200 days prior - maybe a failed auto-update check or a configuration sync that hangs and eventually triggers a safeguard. Good luck, and please update the thread if support gives you an answer!


Automate everything.


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

The quarterly reboot schedule is smart for stability, but it's a band-aid if the vendor has a hardcoded timer. I've had to negotiate credits for undisclosed operational limits like that in SaaS contracts.

Check your maintenance agreement. If a hidden process is causing outages, that's a material defect. Support might "fix" it with a firmware update, but you're owed something for the downtime you've already eaten. Get it documented before they patch it.


—hd


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Interesting pattern. While I haven't seen this specific 200-day mark on Fireboxes, I've encountered analogous behavior in other embedded network appliances, specifically some older load balancers, where an internal health monitor for a low-level filesystem would hit a counter limit and force a restart. The logs were similarly unhelpful, showing just a clean shutdown.

Your approach of gathering data before the support case is correct. I'd suggest also pulling the detailed system diagnostics from just *before* the predicted reboot window. Look for any kernel message buffers (`dmesg` output if you can get it) that might not make it into the standard admin log. Sometimes a recurring, suppressed error increments a counter that eventually trips a failsafe.

A scheduled reboot at 199 days is a viable workaround, but it's treating the symptom. If it's a hidden watchdog, that's a firmware bug they should acknowledge and fix. Please do update the thread with what support says; these undocumented timers are a frequent pain point in closed appliances.


throughput first


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Exactly. The hidden timer theory fits. In vendor management, we call this a "stealth SLA" and it absolutely should be negotiated.

If you can get them to acknowledge it's a bug, don't just accept the fix. Push for a service credit against your last renewal for the unplanned downtime. It sets a precedent and covers your operational cost of troubleshooting.



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

I've seen this exact pattern on M570s. You're right to suspect a watchdog timer. It's likely a counter rollover bug in a low level health monitor, not documented because it's considered an internal safeguard.

Before you open the ticket, check your support contract's SLA for unplanned outage clauses. If they confirm the bug, your negotiation starts there. A scheduled reboot is a workaround, not a fix, and you've already incurred operational impact.

My advice: gather the uptime logs from all affected units, correlate the exact reboot times, and lead with that data when you call. Don't ask if it's a known issue; state the pattern and ask for the root cause. It forces a more substantive response.


SLA is not a suggestion.


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Lead with the data, sure, but don't expect a root cause. They'll just send you a patched firmware and close the ticket.

Your real cost is the investigation hours and the risk you carried before finding the pattern. Get a credit for those, not just the five minutes of downtime. If it's a known internal safeguard, that's a design flaw they've priced in. Make them price it out to you.


show the math


   
ReplyQuote