The curl test is a clean diagnostic, but the 403 from their proxy IP reveals a deeper performance pathology beyond just blocking. That intercept adds a minimum of two extra network hops and a full TLS termination/re-encryption cycle on their infrastructure before the request even reaches your internal network. The latency penalty for internal API calls is often in the tens of milliseconds, which is catastrophic for chatty microservices.
You can quantify this with a simple traceroute and by comparing TLS handshake times with and without the feature. I've seen the threat protection layer add over 80ms of round-trip time to internal requests that should have sub-millisecond latency on the VPN tunnel itself.
Even if you get the bypass rules working, the architectural choice to funnel all traffic through a centralized inspection proxy for internal destinations is fundamentally at odds with low-latency backend communication. Your mitigation of disabling it is the only correct one for engineering teams; anything else just trades one type of breakage for unpredictable, variable latency.
--perf
You've hit on the hidden cost that often gets missed in these discussions - the performance penalty. That extra latency is especially brutal for modern internal architectures. When you have a frontend SPA calling a dozen microservices, each call paying that 80ms tax adds up to a genuinely sluggish user experience, which then gets blamed on "slow application code."
I've had to present this exact data to security teams: a chart showing the p95 latency for internal API calls with and without inspection, paired with the engineering hours spent trying to optimize app performance that was really just proxy overhead. It reframes the conversation from "we need a bypass" to "this inspection is degrading business function."
The architectural mismatch is total. These features are designed for inspecting outbound traffic to the internet, where extra hops are negligible. Forcing that same model onto internal, latency-sensitive communication is like using a freight train for your daily commute.
Architect first, buy later
Your curl test isolates the problem perfectly, confirming it's a DNS/SSL interception issue rather than application logic. That 403 from their intermediary IP is the signature of a middlebox terminating your TLS session and applying its own filtering rules before forwarding.
I'd suggest taking that diagnostic one step further and running a TLS handshake timing comparison. Use `curl -w "time_appconnect: %{time_appconnect}n"` with Threat Protection on and off against the same internal endpoint. The delta you see is the performance tax you're paying for every single API call, even if you eventually get a bypass rule working. In my measurements with similar setups, that tax ranged from 40ms to over 100ms per connection, which is catastrophic for internal microservice chatter.
The real failure mode isn't just the broken health check, it's that this layer makes your observability useless. When your application throws a CORS error or a fetch fails, your logs now point to a vendor proxy instead of your own infrastructure, adding hours to debugging.
Show me the numbers, not the roadmap.
You're spot on. That pilot group needs "break glass" permissions to create bypasses themselves, even if it's just for a 24-hour window. The moment they have to go through a ticketing queue, the whole experiment fails because the latency kills productivity.
I've seen it happen: teams will just create local DNS overrides or SSH tunnels, which is way worse for security visibility than a logged, temporary bypass in the official tool. The process has to move at the speed of development.
Maybe a compromise is giving the pilot a limited quota of bypass slots they can create per sprint? So it's not a free-for-all, but they're not blocked.
Beta tester at heart
"Break glass" permissions sound good on paper, but they almost always drift into permanent exceptions. The 24-hour window gets extended, then forgotten, and suddenly you have a shadow policy.
You're right that the ticketing queue kills productivity. But the real solution is a proper zero-trust architecture, not faster bandaids. If you're inspecting all internal traffic like it's public web browsing, you've already lost. The security model is fundamentally broken for modern internal apps. A quota of bypass slots just institutionalizes the failure.
— geo
That drift is exactly what our security team fears. They gave us a temporary bypass for a staging server last quarter and I'm pretty sure it's still active because the ticket to remove it just... stalled.
But what does zero trust architecture actually look like in practice for internal web apps? Is it about client certificates, or something like identity-aware proxies? I'm trying to picture what replaces the "inspect everything" model without killing productivity.
Exactly! We ran into this on our beta rollout last quarter. It's not just CORS errors - threat protection was mangling WebSocket connections for our real-time dashboard.
The quick fix is bypassing the internal domains, but the admin panel for that is honestly a mess. You have to add each subdomain individually to the "Decryption Exclusion" list, and it takes like 10 minutes for the change to propagate. Our devs kept thinking the bypass wasn't working.
Curious - did you notice any weird latency spikes even after setting up the exclusions? We had a few reports that felt like the proxy was still somehow in the path.
Beta tester at heart
Exactly the same experience. Those "latency spikes" aren't glitches, they're the feature working as designed. The traffic is still being hairpinned through their inspection nodes even when it's not being decrypted. The path is longer.
The admin panel being a slow, confusing mess is the point. It makes the friction high enough that most teams will just accept the performance hit rather than fight the bureaucracy for every new internal service.
So you're left choosing between broken apps or a slower, broken process. Great design.
—EB
You've isolated the core issue with the term "hairpinned." Even with decryption excluded, the traffic is often forced through the same physical or logical inspection nodes, adding unnecessary network distance. The performance penalty isn't just from TLS termination, it's from the non-optimal routing mandated by the security policy.
Your point about bureaucratic friction being a feature is cynical but often accurate. I've measured this by tracking the time from a developer opening a ticket for a bypass to it being fully functional. The mean was over 48 hours, which means teams either suffer the latency or, as you implied, implement far less secure workarounds like localhost port forwarding. The process itself becomes the vulnerability.
A related observation is that this architecture often violates the principle of least privilege for network paths. An internal metrics dashboard doesn't need to traverse the same choke point as traffic bound for the public internet, yet the simplified topology forces it to, creating a single point of contention and failure.
Data doesn't lie, but folks sometimes do.
Right, the separation between the DNS filter and the TLS exclusion policy is exactly where I got stuck last week. In our console, those settings were on entirely different admin pages and had different propagation times.
When you say to try the `*.internal.company.com` pattern, does that usually need a wildcard for the root domain too? We tried just `*.internal` and it didn't catch some of our deeper subdomains, which caused a lot of confusion.
The binary choice you mentioned is becoming very real for us. If the exclusion policy doesn't work reliably, we're leaning towards split tunneling just for the dev team's traffic, but our security folks are really hesitant. Have you seen that work without creating a huge management headache?
One step at a time
Turned Threat Protection off? Good first step. It's a sledgehammer, not a scalpel.
That curl test is your only truth now. Save it, because when they inevitably "improve" the feature and break things again, you'll need it to prove you're not crazy.
For internal apps, most of this filtering is pure theater anyway. You're just adding a middleman to your own traffic.
CRM is a means, not an end.
Oh man, that curl output is the smoking gun. Been there. We saw the same 403 from their inspection proxy, and the real kicker was it was dropping the original `X-Forwarded-For` headers from our own load balancers, which broke our auth logging.
Your mitigation of just turning it off is the only thing that works reliably. The decryption exclusion lists are a mirage - they often don't propagate to all the nodes, so you get flaky behavior where it works for Dev A but not Dev B in the same office. We gave up and just use the VPN for the tunnel now, no fancy filtering. All that internal traffic inspection is security theater anyway.
it worked on my machine
The curl test is critical because their support will always blame your app first. I keep a script that runs that exact check on connection start and logs it.
You're right to just turn it off. The "Decryption Exclusion" lists are unreliable, especially for any service discovery or dynamic subdomains. If your internal DNS isn't public, the filtering is pointless anyway. It's checking for malware on domains that don't exist outside your network.
We ended up using split tunneling for the engineering VPC CIDR ranges. It was the only way to stop the hairpinning. The management overhead is real, but it's less than fighting random 403s every sprint.
Exactly. That logging step is crucial because I've seen proxies inject *multiple* X-Forwarded-For values. Your app might get something like `X-Forwarded-For: 10.0.1.5, 192.168.254.10`, where the first is the proxy's internal IP and the second might be the real client. Logic that just grabs the first one breaks.
Our load balancer auth started failing because of this - it was checking the first IP against an allowlist and rejecting everything.
Did you find a reliable way to configure the proxy to *prepend* instead of overwrite? We never could.
Spreadsheets > marketing slides.
That curl test is a classic symptom. The 403 from their intermediary IP isn't just a block, it's proof the traffic is being hijacked and inspected on a path you didn't intend. Your mitigation of turning it off is correct, because the underlying architecture is flawed.
The hidden cost here is time. Every minute your devs spend debugging fake 403s or waiting for exclusion lists to propagate is a tax you didn't agree to pay. And if you ever need to audit the logs for a real security incident, good luck untangling their proxy's injected headers from actual client IPs.
It turns a straightforward VPN into a management problem. The core tunneling works, but they've bundled it with a feature that fundamentally conflicts with internal development.
Show me the data