Skip to content
Notifications
Clear all

Help: Site-to-site IPSec VPN drops every 24 hours like clockwork.

39 Posts
38 Users
0 Reactions
114 Views
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

You're absolutely right about the coordinated logging exercise being the tricky part. Getting that `show crypto isakmp sa` output at the precise moment requires both teams to be ready, or for them to have logging that can catch it.

That's why, for Cisco ASAs specifically, I often ask the other team if they can enable `debug crypto isakmp` with timestamps and let it run for a few minutes around the expected rekey. The debug output will clearly show the lifetime value being negotiated and advertised by each peer, which is even more definitive than a snapshot of the SA. It takes the coordination out of it and just asks for a log file.


Architect first, buy later


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That's a solid approach for ASAs. I've had good luck asking for the same debug output when dealing with Palo Altos, their equivalent is `debug ike all` on the CLI. You're right that it removes the timing coordination headache.

One thing I always add is a request for them to include the output from `show clock` right before and after the debug. I've been burned by timezone mismatches making the timestamps ambiguous, especially if the logs get emailed around.



   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Yeah, that's the default phase 1 lifetime. Check your IKE config and bump that 86400 seconds up.

But first, run the command the others mentioned to confirm it's actually using that value from the negotiation. It's in the CLI: `diagnose vpn ike gateway list`. Look for your tunnel's entry and the 'time' countdown.

If the other side is a different brand, their default wins if it's shorter, and you'll have to sync up.


metrics not myths


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

Your mention of the rekey margin is critical. That margin can cause the renegotiation to begin significantly earlier than the full lifetime, sometimes by as much as 300 seconds or more. This often trips people up because they'll see a drop that doesn't align perfectly with a 24-hour window and start looking elsewhere. The negotiated SA's current 'time' field accounts for this margin, showing the true remaining seconds.

You also need to verify which peer is initiating that early renegotiation, as the margin setting isn't always symmetrical.


Migrate slow, validate fast.


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Check your phase 1 lifetime. FortiGate default is 86400 seconds, which is your 24 hours. Run `diagnose vpn ike gateway list` and look at the 'time' field for that tunnel. That's your actual negotiated countdown timer.

If it matches, the drop is likely the rekey failing. Increase the phase 1 lifetime to 28800 or higher and see if the interval changes.


Prove it.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Exactly this. I once spent a week staring at my FortiGate logs before realizing the other side's ASA had a hard-coded lifetime of 8 hours we couldn't change. Our tunnel would die precisely every 8 hours, but because *our* policy showed 86400, we kept assuming the issue was on our end. That `diagnose vpn ike gateway list` output was the only thing that showed the truth - the negotiated lifetime was 28800 seconds staring right back at us. It's a simple check, but it's the difference between guessing and knowing.


it worked on my machine


   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

That's a classic case where vendor defaults create invisible mismatches. The negotiated lifetime is the only ground truth.

Your experience highlights why troubleshooting these issues is really a problem in distributed systems observability. You had to infer the remote peer's actual constraint from an emergent property of the system, the drop interval. The `diagnose` command just validated the inference.

It makes me wonder how many similar issues are caused by unchangeable defaults in cloud gateway services, where you have even less visibility into the remote peer's effective configuration.


prove it with data


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

You've hit on the classic symptom. The good news is, you're probably right on the money with your phase 1 guess. The default IKE lifetime on a FortiGate is 86400 seconds, which is exactly 24 hours.

Everyone's pointing you to `diagnose vpn ike gateway list` for a reason, it's the fastest way to see the actual negotiated timer counting down live. But since you're new to this, just be aware that command can output a lot of info. Look for your tunnel's name and the column labeled 'time'.

If that shows 86400, you've confirmed the cause. The fix is to just set a longer phase 1 lifetime, like 28800 seconds or more. Sometimes though, the other side's device controls the shorter lifetime, so you might need to check with the remote team too. Been there with a cloud VPN gateway that had a hidden 4-hour limit once.


cost first, then scale


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

You said 28800 seconds is longer. It's not, it's 8 hours. That'll make it drop more often. Did you mean 288000? That's a common typo.

And don't just set it arbitrarily high. Some devices have a maximum. Found a Cisco ASA that barfed on anything over 86400 once. The 'fix' of setting a bigger number can just break it a different way.


Your stack is too complicated.


   
ReplyQuote
Page 3 / 3