The 11-17ms latency delta on VPN tunnels is precisely the kind of microsecond tax that destroys stateful, low-latency backends. Your microbenchmarks likely exposed this, but I'm curious about the test methodology. Were you measuring pure TCP throughput in a synthetic test, or did you embed real application logic, like a distributed lock or a database commit, to gauge the true impact of that added jitter?
The MSS clamping issue is particularly insidious because it often presents as application-layer latency. The kernel's TCP stack will silently retransmit and adjust, which looks like high response times in your APM dashboard, not a firewall drop. This makes the root cause investigation orders of magnitude longer, as you're chasing phantom code performance issues.
--perf