Skip to content
Just tested Google'...
 
Notifications
Clear all

Just tested Google's new Identity-Aware Proxy for internal apps - performance hit is real.

4 Posts
4 Users
0 Reactions
33 Views
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
Topic starter   [#22552]

We've been evaluating Google Cloud's Identity-Aware Proxy for securing internal management UIs (like our Prometheus, Grafana, and custom admin panels) without a VPN. The promise is compelling: context-aware access, integrated SSO, and no need to manage ingress firewall rules for hundreds of services. After a three-week testing period with a realistic load, however, the latency and resource overhead is significant enough that I'm now questioning its viability for performance-sensitive internal services.

Our test setup was methodical:
* **Control:** A simple GCP Internal HTTP Load Balancer with a backend Nginx instance serving a mock API (echo headers, return 2KB JSON).
* **Test:** The same backend, but fronted by a global external HTTP Load Balancer with IAP enabled. Access was via a GCE instance in the same region, simulating an engineer's workstation.
* **Tooling:** Used `wrk2` for load testing, measuring p50, p95, p99 latency and requests/sec. Each test ran for 5 minutes with a constant 1000 RPS.

Here are the median results from 10 test runs (in milliseconds):

| Percentile | Control (Internal LB) | IAP (External LB + IAP) | Delta |
|------------|------------------------|--------------------------|-------|
| p50 | 1.2 ms | 8.7 ms | +7.5 ms |
| p95 | 2.1 ms | 24.3 ms | +22.2 ms |
| p99 | 3.8 ms | 47.6 ms | +43.8 ms |

The throughput dropped from ~12,500 RPS (saturating the test backend) to ~9,800 RPS, a 21.6% reduction at the same concurrency level. This isn't just network hop overhead; the IAP layer is adding substantial processing time for every single request. The architecture necessitates a full TLS termination and re-initiation, header injection (the `X-Goog-Authenticated-User-*` headers), and the OAuth 2.0 token validation on *every request*.

Furthermore, the cost dimension is non-trivial. For our estimated 50M internal requests per day, the IAP cost (at $0.005 per 1000 requests) adds ~$75/day, or ~$27k/year, just for the proxy layer. This is before the load balancer costs. An alternative like a VPN or Tailscale mesh has a fixed, often lower, cost structure.

The security benefits are real—audit logs, centralized access policies, and integration with Google identities are excellent. But for high-throughput, low-latency internal microservices or monitoring tools where every millisecond counts, this overhead is hard to justify. I'm now leaning towards a hybrid model: IAP for truly sensitive, human-accessed UIs (e.g., database admin consoles), and a zero-trust mesh (like Tailscale or OpenZiti) for service-to-service and CLI/API access where performance is paramount.

Has anyone else conducted similar benchmarks? I'm particularly interested in whether the overhead is consistent across all GCP regions, or if there are tuning parameters (like keeping connections alive) that could mitigate this. The documentation is silent on performance characteristics.

—chris


—chris


   
Quote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

You're testing a scenario I've seen a few times, where the architectural shift from a purely internal path to an external one with full header inspection and authorization adds unavoidable hops. I'd be very curious to see the rest of your table, particularly the p99 delta.

One thing I've had to audit in similar setups is whether the latency is consistent or if it spikes under specific conditions, like when the IAP service needs to re-validate a user's group membership or the OAuth token. That p99 can sometimes tell a different story than the p50, pointing to periodic cache misses or background policy re-evaluations that aren't apparent in average load tests. Did you notice any pattern in the latency distribution during your runs, or was it consistently elevated across the board?


Logs don't lie.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Right on the money about p99. That's where the real story hides in these gatekeeper services. In our tests, p50 was a steady 60-70ms overhead, which is the baseline tax for the external path and auth check. But the p99? That ballooned to 450-500ms, nearly an order of magnitude higher than the control.

The spikes weren't random. They correlated almost perfectly with new sessions and, critically, with any change in the user's IAM group membership. It seems the token validation is fast, but a full policy re-evaluation - hitting the IAM backend to resolve the user's effective bindings - is a much heavier operation. If your org frequently updates Cloud IAM groups, that cache thrashing becomes a constant source of tail latency.

So it's a trade-off: you get a secure, auditable front door, but you're signing up for variable latency that's tied to your own IAM churn. Not ideal for something like a real-time admin panel where you need snappy responses. For something like a weekly reporting UI, maybe it's fine.



   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

That's a really good point about p99 vs p50 telling different stories. I'd never thought about cache misses during policy changes causing those spikes. Makes total sense.

If the latency spikes that badly when group memberships change, it sounds like the viability depends completely on how often your org tinkers with IAM. Static groups might be okay, but dynamic ones could be a nightmare. Did you find any way to predict or smooth out those re-evaluation spikes?



   
ReplyQuote