Skip to content
Notifications
Clear all

Troubleshooting: Site-to-site VPN drops every 24 hours like clockwork.

2 Posts
2 Users
0 Reactions
36 Views
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
Topic starter   [#3524]

We are in the process of consolidating several regional data centers into a centralized hub, utilizing Check Point Quantum Security Gateways (R81.10) as the VPN endpoints. The topology is a standard site-to-site mesh, using IKEv1 with Main Mode and SHA-256/AES-256 for both Phase 1 and Phase 2. The tunnels establish successfully and pass traffic without issue.

However, we are encountering a persistent and precisely timed operational fault: **every tunnel drops exactly every 24 hours**, with renegotiation succeeding immediately afterward. This results in a brief but consistent packet loss window. The logs indicate a clean, non-failure-based rekeying event, but the drop is observable at the network level.

Our current working theory involves a mismatch in lifetime timers, but our analysis of the configuration has not yet revealed the source. I have performed a granular review of the following domains:

* **IKE and IPsec Proposals:** Lifetime values are explicitly set to 86400 seconds (24 hours) on both ends for both Phase 1 and Phase 2.
* **Check Point Time Synchronization:** All gateways are synchronized to the same NTP source, with less than 50ms skew.
* **Dead Peer Detection (DPD):** Enabled with default settings (`send every 10 seconds, trigger 5 failures`). Logs do not show DPD-triggered resets preceding the drop.
* **NAT-T:** Enabled uniformly. The drops occur regardless of whether traffic is traversing a NAT device.

The diagnostic output from `vpn debug` on one of the hubs shows the following pattern preceding the drop:

```
ike[2421]: IKE_SA rekeying: IKE_SA (SPI: 0x8a3b...c7) is about to rekey.
ike[2421]: IKE_SA rekeying: New IKE_SA (SPI: 0x1f9d...e2) successfully created.
ipsec[2423]: CHILD_SA rekeying: CHILD_SA (SPI: 0xcb12...) is about to rekey.
ipsec[2423]: CHILD_SA rekeying: New CHILD_SA (SPI: 0x854f...) successfully created.
```

Despite the logs indicating a clean rekey, `vpn tu` shows the tunnel interface flapping, and our monitoring records the associated latency spike and packet loss.

My specific questions for the community are:

1. Has anyone observed similar *exactly* 24-hour drops in a Quantum VPN environment where lifetimes are explicitly matched? Could there be a hidden, non-configurable internal timer or a default from a pre-R80.x template manifesting?
2. Are there known interactions between the IPsec lifetimes and other Check Point daemons (like `cphad` for cluster members) that might force a re-establishment instead of a seamless rekey?
3. What would be the most definitive method to validate the *actual* negotiated lifetimes on the active SAs, beyond what is stated in the SmartConsole? Is there a `vpntool` command that provides this in a parseable format?

I am prepared to share anonymized `vpn debug` outputs and relevant `cpconfig` sections if it aids in a comparative analysis. The precision of the interval suggests a timer, but the source remains elusive.



   
Quote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Great detail in your analysis so far. The 24-hour pattern is the biggest clue, and you're right to zero in on those lifetime timers. Even with them explicitly set to 86400 seconds, there's a nuance I've seen bite people: the timer often starts at session establishment, not at midnight. So if your tunnels all came up around the same time during initial config, they could be expiring in sync.

You mentioned checking Dead Peer Detection - that's a good next step. An aggressive DPD timer could theoretically cause a blip, but the clockwork timing still points to rekey. Have you checked if there's a system-level "session timeout" or "key lifetime" policy somewhere that might be overriding the tunnel-specific proposals? Sometimes there's a global setting hiding in a different menu.


Stay curious, stay skeptical.


   
ReplyQuote