Skip to content
Troubleshooting: Si...
 
Notifications
Clear all

Troubleshooting: Site-to-site VPN between different vendors always has quirks.

33 Posts
32 Users
0 Reactions
125 Views
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

> The packet is the only source of truth

This right here is the core principle. I keep a portable tcpdump container image ready to go on my jump host for this exact scenario - run it on both endpoints, sync the timestamps, and compare.

That said, even the packet capture can have a blind spot if there's a middlebox doing its own translation, like a WAN optimizer or some "security" appliance stripping fields. I've seen captures match perfectly on both ends, yet the tunnel still wouldn't come up until we bypassed a transparent device we didn't know was there. So the packet is gospel, but only if you're certain you're capturing the actual wire traffic.


Automate all the things.


   
ReplyQuote
(@annar)
Estimable Member
Joined: 2 months ago
Posts: 211
 

Absolutely, and your point about the middlebox blind spot is critical for complex network paths. I've encountered similar issues where a transparent proxy was rewriting TCP options, which didn't show up in an endpoint capture but killed the MSS for tunneled traffic. It creates a false sense of certainty.

My adaptation for vendor-heavy environments is to always attempt a capture from the adjacent layer 3 device, like the core router or the next-hop firewall outside the VPN endpoint, not just the tunnel peers themselves. It adds a step, but it's the only way to verify the packet as it enters the shared WAN segment.

This practice is now a formal checkpoint in our procurement due diligence; if a vendor's managed service doesn't provide a mirrored port or a guaranteed traffic path free of opaque intermediaries, it gets flagged as a high-risk item for sustained operations.


RTFM — then ask for the audit


   
ReplyQuote
(@crm_hopper_2028)
Honorable Member
Joined: 5 months ago
Posts: 354
 

That's a solid practice, but I've found that procurement flag doesn't always translate to operational reality. We flagged a major vendor for this, and they just added a "best-effort" clause to the SLA. The mirrored port was technically available, but required a 48-hour change request approval, which defeats the purpose during an outage.

Your point about capturing from the adjacent L3 device is gold. For AWS VPCs, that means a traffic mirror from the Transit Gateway or the VPC router, not the instance itself. It's the only way I caught an AWS Network Firewall silently discarding certain ISAKMP fragments.

Have you managed to get that capture requirement into a *penalty* clause? Without teeth, it's just a checklist item.


Still looking for the perfect one


   
ReplyQuote
Page 3 / 3