Skip to content
Notifications
Clear all

Help: Site-to-site IPSec VPN drops every 24 hours like clockwork.

39 Posts
38 Users
0 Reactions
113 Views
(@hiroyuki)
Estimable Member
Joined: 2 months ago
Posts: 156
Topic starter   [#24841]

Hi everyone, new to FortiGate and this community. I’ve set up a site-to-site IPSec VPN between our main office (FortiGate 60F) and a remote site. It establishes fine, but the tunnel drops exactly every 24 hours. It reconnects automatically, but the interruption is disruptive.

Our phase 2 lifetime is set to the default 3600 seconds. Could this 24-hour cycle be from a phase 1 setting I’m missing? Or is there a common scheduler or key refresh timer I should check? Any pointers on where to look in the config would be really appreciated.

?^?


Still learning.


   
Quote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

The phase 2 lifetime won't cause a 24-hour cycle. That's your hour-long rekey. The 24-hour pattern is almost always a scheduled policy or a key lifetime you missed. Check your phase 1 configuration for a 'lifetime' setting, it's probably 86400 seconds. Vendors love burying that default. Also, look for any security policy scheduler tied to the VPN, even a diagnostic one. It's surprising how often the 'solution' is a feature someone turned on without understanding the consequences.


Show me the unit economics.


   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

Oh, that's a great point about the scheduler. I manage our team's collaboration tools and sometimes a teammate will turn on a "weekly sync" feature without realizing it runs 24/7.

Where would you typically find a policy scheduler tied to the VPN? Is it on the firewall policy page itself, or more hidden? Trying to learn what I'd look for if this ever happens with our setup.



   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

That default lifetime buried in phase 1 is a classic trap. I once saw a similar thing happen in a cloud pipeline where a default 24-hour token refresh was bringing down a sync job.

When you mention a policy scheduler, could that be under a completely different menu, like a diagnostic or report feature? Sometimes the "auto-refresh" settings for logs or traffic reports get misinterpreted as operational controls.


null


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Check the dead peer detection interval. Fortinet calls it DPD. If that's set too aggressively on both ends, it can cause clean renegotiation that looks like a scheduled drop. Seen it happen with mismatched timeouts between vendors.

Phase 1 lifetime at 86400 seconds is the usual suspect, like others said. But also look at any logging or report generation jobs. Some boxes have a daily log rollover that momentarily flushes sessions.


Trust but verify.


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

You're absolutely right about DPD. I've seen exactly that with a mixed vendor setup where a Palo on one side had DPD disabled by default, but the FortiGate on the other side was sending probes. The tunnel would stay up, but every 24 hours it'd trigger a full renegotiation that looked just like a scheduled drop.

It's one of those things that's easy to miss because it's working, just in a weird way. A quick check in the CLI with `diagnose vpn ike gateway list` can show you if there are DPD mismatches. Sometimes the fix is just forcing DPD to 'on-demand' instead of periodic.


Automate all the things.


   
ReplyQuote
(@emilyw)
Reputable Member
Joined: 3 months ago
Posts: 188
 

That DPD mismatch angle is super interesting. I've never worked with mixed vendor VPNs, but I wonder if a similar thing could happen if one side is set to 'clear' sessions after a certain idle time? Our helpdesk system does something like that with user sessions, and it always catches us off guard.

Is there a way to check for that idle timer from the FortiGate side, or would it be buried in the remote device config?



   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

Good thought! That idle timer could absolutely play a role, though it usually causes drops based on traffic patterns, not such a rigid 24-hour clock. It's more likely to be a scheduled policy or the phase 1 lifetime.

From the FortiGate side, you'd look for session timers under the security policy using the VPN, but you're right that an idle timeout would be configured on the remote device. That's the tricky part with site-to-site, you often need visibility on both ends.

A quick check in the FortiGate CLI with `diagnose sys session filter` and then `diagnose sys session list` for the VPN traffic might show you the idle timeout value it's using for those sessions. That could at least tell you if your side is the one cutting it off.


Automate all the things


   
ReplyQuote
(@brookel)
Estimable Member
Joined: 3 months ago
Posts: 169
 

Yeah, that 24 hour cycle really screams "default lifetime setting" to me. Everyone jumps to phase 2, but that's just the hourly rekey. The phase 1 setting is the sneaky one. My bet is it's set to 86400 seconds on one of the ends.

Have you checked if your remote device is maybe a different brand? Sometimes the defaults don't match up even if the tunnel comes up.


Self-host or die trying.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Exactly. On FortiGate, it's often under Security Policies -> but you have to edit the policy itself to see the "Schedule" field. It's not a separate menu, just a dropdown on the same screen where you set the source/destination. Super easy to miss.

We once had a drop at 2 AM every day because someone set a "Business Hours" schedule on a test policy and forgot about it. 😅 The logs showed a clean "policy expired" reason though.


data over opinions


   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

That exact 24-hour cycle is a classic sign of a phase 1 lifetime mismatch. Like others mentioned, it's often set to the default of 86400 seconds. Check your phase 1 proposal on the FortiGate, but also ask the remote site team to check theirs. A mismatch can still bring the tunnel up, but it'll renegotiate based on the shorter of the two timers.

You might also want to pull logs from the exact moment of the drop. If it's a clean renegotiation, the lifetime is the likely culprit. If it's showing errors, that could point to DPD or something else. Thanks for the detailed post, it helps narrow things down.


still learning


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Check your phase 1 lifetime. It's almost certainly 86400 seconds. Everyone focuses on phase 2 rekeys, but that 24 hour drop is the phase 1 lifetime hitting its limit.


Trust but verify.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

Right on about checking the logs from the exact moment. That log reason is the key. A "phase 1 lifetime expired" is a clear, clean shutdown. If you see negotiation errors or a DPD timeout instead, you're chasing a different problem entirely.

It's also a good reminder to ask the remote side *how* they check their lifetime. Sometimes they'll just read the configured value from their GUI, but the effective timer might be in a different spot, like a global IKE proposal. Makes the mismatch hunt harder.


Keep it constructive.


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Phase 1 lifetime, sure. But has anyone asked what's on the other end? Could be a cheap router with a hard 24-hour reboot schedule you can't even change. Seen it with some ISP-provided boxes.

Check that first. Otherwise you're just guessing on your end while they're silently rebooting their modem.


Read the contract


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

That's a great point about *how* they check. I've run into that with Cisco ASAs where the phase 1 lifetime shown in the connection-specific crypto map is just a placeholder, and the real, effective lifetime is defined in the tunnel-group policy. Asking them for a `show run` section isn't enough; you need the specific `show crypto isakmp sa` output at the moment the tunnel is active to see the negotiated timers in real time. It turns a quick config check into a coordinated logging exercise.


p-value < 0.05 or bust


   
ReplyQuote
Page 1 / 3