We had a critical network outage on our SRX340 last Tuesday. Firewall stopped passing traffic. No failover.
Opened a P1 case with JTAC. Here's the timeline:
- T+0: Called in. Got a case number after 15 minutes.
- T+45min: First engineer joined. Asked for basic config and logs. Sent immediately.
- T+3hrs: Escalation requested. No real progress. Just more data collection.
- T+6hrs: Second engineer. Suspected a software bug. Pushed for an RMA.
- T+10hrs: Got an RMA number. No advanced replacement offered. Had to argue for cross-shipping.
- T+28hrs: Replacement unit arrived. Back online.
Total downtime: over a day. Support contract is not cheap. Expected faster resolution for a P1. The process felt slow and scripted. The hardware failed, but the support experience was almost as painful.
Your timeline illustrates a systemic problem with hardware-centric support models in an era where redundancy should be software-defined. A single SRX340 becoming a single point of failure with a 28-hour recovery window is architecturally unacceptable, regardless of the RMA process speed.
The real issue is depending on JTAC's reaction time as your high-availability strategy. In any cloud or modern data center design, we'd treat that firewall as ephemeral cattle. You'd have automated configuration deployment ready so a replacement unit, whether from JTAC or AWS Marketplace, could be brought online in minutes, not hours. The support contract becomes a spare parts logistics agreement, not your lifeline.
I've moved teams toward active-active firewall clusters across zones or, better yet, to cloud-native firewall services where the provider manages the hardware failure domain. Your painful experience is the exact cost of not having that architectural layer of abstraction.
Boring is beautiful