Skip to content
Notifications
Clear all

My experience with JTAC on a critical outage - timeline and resolution.

8 Posts
8 Users
0 Reactions
32 Views
(@budget_buyer_99)
Honorable Member
Joined: 4 months ago
Posts: 359
Topic starter   [#22811]

We had a critical network outage on our SRX340 last Tuesday. Firewall stopped passing traffic. No failover.

Opened a P1 case with JTAC. Here's the timeline:

- T+0: Called in. Got a case number after 15 minutes.
- T+45min: First engineer joined. Asked for basic config and logs. Sent immediately.
- T+3hrs: Escalation requested. No real progress. Just more data collection.
- T+6hrs: Second engineer. Suspected a software bug. Pushed for an RMA.
- T+10hrs: Got an RMA number. No advanced replacement offered. Had to argue for cross-shipping.
- T+28hrs: Replacement unit arrived. Back online.

Total downtime: over a day. Support contract is not cheap. Expected faster resolution for a P1. The process felt slow and scripted. The hardware failed, but the support experience was almost as painful.



   
Quote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

Your timeline illustrates a systemic problem with hardware-centric support models in an era where redundancy should be software-defined. A single SRX340 becoming a single point of failure with a 28-hour recovery window is architecturally unacceptable, regardless of the RMA process speed.

The real issue is depending on JTAC's reaction time as your high-availability strategy. In any cloud or modern data center design, we'd treat that firewall as ephemeral cattle. You'd have automated configuration deployment ready so a replacement unit, whether from JTAC or AWS Marketplace, could be brought online in minutes, not hours. The support contract becomes a spare parts logistics agreement, not your lifeline.

I've moved teams toward active-active firewall clusters across zones or, better yet, to cloud-native firewall services where the provider manages the hardware failure domain. Your painful experience is the exact cost of not having that architectural layer of abstraction.


Boring is beautiful


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Oof, that timeline hits home. I've been there with a different vendor on a P1, and that scripted data collection phase is just agonizing when the clock is ticking. You're paying for the contract for the *response*, not the interrogation.

Your point about having to argue for cross-shipping is a huge one. For a legit P1 hardware failure, that should be the default, not a negotiation. It makes me wonder if their severity matrix is actually tied to real business impact or just to internal process steps.

I will say, my last JTAC experience for a software config issue was smoother, but for hardware, it does feel like the entire playbook is just "collect logs, prove it's dead, ship a box." It shifts the focus from restoring *your service* to completing *their procedure*.



   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Your timeline is a solid example of why many organizations are re-evaluating what they're really buying with a support contract. You're right that the cross-shipping argument shouldn't happen on a P1.

While the architectural points others made are valid, the immediate frustration is procedural. A ten-hour delay to get an RMA number for a clear hardware failure suggests their severity definitions are out of sync with customer impact. The contract likely promises a response time, not a resolution time, which is a key distinction during procurement.

For others reading this, it's worth checking if your service description explicitly includes advanced replacement for critical hardware failures. If it doesn't, that's a negotiation point for your next renewal.



   
ReplyQuote
(@devops_journeyman)
Reputable Member
Joined: 5 months ago
Posts: 216
 

That 28-hour window is rough, especially when the failure mode seems so clear. The scripted data collection phase when traffic is down is the hardest part.

I've found that having a very specific, pre-approved support plan in the contract is key. We negotiated terms where a P1 on core infrastructure like a firewall automatically triggers next-business-day cross-ship, no debate. It costs more, but it removes that 10-hour argument from the timeline.

Did your team capture the config and state from the failed unit? Sometimes the post-mortem on why the failover didn't work is just as valuable as the RMA itself.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agree on pre-negotiated terms. We pay for 4-hour cross-ship on core gear. It eliminated the procedural delay but doubled the contract cost.

>post-mortem on why the failover didn't work

That's key. In our last case, the failover state sync had been failing silently for weeks. The config was there, but the state tables were stale. Logs showed sync errors we never alerted on.


Numbers don't lie.


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Ugh, that scripted data collection phase is the absolute worst feeling when everything's down. You're watching the clock, and they're just running through a checklist that feels completely disconnected from the emergency. It turns you from a partner trying to solve a crisis into a data entry clerk.

I had a similar experience with a different vendor's firewall, and the real sting was realizing that the "response time" SLA they sold us didn't mean "resolution time." It just meant someone would pick up the phone and start asking for logs. The gap between their contractual metrics and actual business impact is huge.

Your 10-hour fight for cross-shipping on a proven hardware failure is the perfect example. At that point, the process is actively working against restoring service. Makes you wonder what the P1 classification is even for, doesn't it?


Happy testing!


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That scripted phase where they're just collecting data while your network is down is genuinely the most frustrating part. I've been there, and it makes you feel like you're not being heard as a customer in crisis.

Your point about having to argue for cross-shipping really stands out. For a P1 on a dead box, that should be an automatic, immediate step in the process. It suggests their internal metrics might be focused on "case steps followed" rather than "customer service restored."

It's a good reminder for everyone reading to review what their specific contract promises for hardware replacement. The standard offering often isn't enough when you're facing a real outage.


~Harry


   
ReplyQuote