We've been migrating from a mesh of site-to-site IPSec tunnels to using BGP over IPSec for dynamic routing on our FortiGate 600E clusters (running 7.2.5). The basic functionality is there—prefixes are exchanged, and traffic can flow. The problem is failover and convergence time is abysmal, bordering on unusable for any service with real-time requirements. We're seeing 45 to 90 seconds of blackhole during a simple path flap, which defeats the entire purpose of a dynamic routing setup.
Our core setup is textbook:
* iBGP between the FortiGates at each datacenter (route reflector clients).
* eBGP multihop over the IPSec tunnels to the remote site FortiGates (ASN 65001).
* BGP timers are tuned down (`set keepalive-timer 3`, `set holdtime-timer 9`).
* IPSec phase1 is using AES256-SHA256, DH group 14, keylife 28800 seconds.
* IPSec phase2 is using AES256-SHA256, replay detection enabled, perfect forward secrecy disabled.
The tunnels themselves stay up. The issue seems to be the detection of the underlying IPSec tunnel state and the subsequent BGP reaction. I've confirmed the next-hop for the BGP routes is the tunnel interface IP. When I hard-shut the primary ISP link on one side, the IPSec tunnel doesn't immediately transition to a down state. BGP holds onto the routes until the dead timer finally expires, which is where we're losing minutes.
Here are the specific diagnostics and what I've tried without success:
```
# Relevant config snippets
config router bgp
set as 65000
set router-id 10.255.255.1
set keepalive-timer 3
set holdtime-timer 9
config neighbor
edit "10.10.10.2" # Tunnel interface IP of remote peer
set soft-reconfiguration enable
set remote-as 65001
set ebgp-multihop 255
set update-source "ipsec_tunnel_interface"
set route-map-out "prefer-link-a"
next
end
end
# For the tunnel interface itself
config system interface
edit "ipsec_tunnel_interface"
set remote-ip 10.10.10.2 255.255.255.255
set interface "wan1"
next
end
```
Attempted fixes that made no measurable difference:
1. Setting `set auto-negotiate enable` and `set add-route disable` on the phase1 config.
2. Experimenting with Dead Peer Detection intervals (set to 2 seconds on both ends).
3. Trying `set network-import-check disable` in the BGP config.
4. Creating a static route for the remote tunnel endpoint with a distance of 1, hoping the route withdrawal would be faster.
The suspicion is that the FortiGate's kernel is not tightly coupling the BGP session state to the IPSec SA state. In a pure network device, if a physical link drops, the BGP session instantly resets. Over an IPSec tunnel, the "link" is logical, and the failure detection seems lazy. Has anyone actually gotten sub-10-second convergence with BGP over IPSec on this platform, or is this a fundamental architectural limitation? I'm looking for concrete debug commands or a known working configuration template, not a suggestion to open a TAC case.
latency is a liar
You're on the right track suspecting the tunnel state detection. The tuned BGP timers are irrelevant if the underlying IPSec interface takes tens of seconds to transition to 'down'. FortiOS often treats the VPN interface as administratively up until the dead peer detection (DPD) timeout expires.
Check your `set dpd-retryinterval` under the phase1 config. The default is often 60 seconds. You need to aggressively lower this, along with `set dpd-retrycount`, to force a quicker interface state change. Be aware this increases control plane traffic.
Also, verify you're not using *automatic* metrics on your tunnel interfaces. A hardcoded, consistent metric ensures the route table withdrawal is immediate upon interface loss.
Data is the only truth.