Skip to content
Notifications
Clear all

Has anyone quantified the increased latency from all the runtime security checks?

19 Posts
17 Users
0 Reactions
81 Views
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
Topic starter   [#22128]

Hey everyone, I've been looking into Zero Trust/ZTNA solutions like Absolute Secure Access for my company's cloud migration.

Everyone talks about the security benefits, which make sense. But I'm trying to understand the performance cost for our developers and remote users. Has anyone actually measured or benchmarked the added latency from all the continuous posture checks and traffic inspection?

I'm especially curious about the impact on real-time apps or pulling large artifacts from S3 buckets. Is it just a few extra milliseconds, or something more noticeable? Any real-world numbers or experiences would be super helpful!



   
Quote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

Great question, and it's smart to think about the user experience alongside security. The few extra milliseconds per check is usually accurate for a simple posture validation.

However, I've seen noticeable latency emerge from a combination of factors, not just the check itself. For instance, if the policy engine is in a different region than your S3 bucket, that geo-latency can compound. Real-time apps can feel the impact more, especially if the solution is doing deep packet inspection on every single packet rather than just at session start.

The vendor's architecture makes a huge difference here. Some are much more optimized for developer workflows than others.


Keep it constructive.


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

Yeah, that's a really practical concern. User1079 is spot on about the compounding factors.

I'd add that the "continuous" part is key for measurement. A single check at login is negligible. But if you're pulling dozens of S3 artifacts in a CI/CD pipeline, and each one triggers a new posture check or the connection is being re-evaluated every 30 seconds, those milliseconds start to add up in a way users will notice. You're looking at the aggregate over a work session, not just a single transaction.

Some solutions have a "developer mode" or can whitelist trusted internal workflows to relax the frequency of checks for known-safe patterns. Might be worth asking vendors about that kind of granularity.


~Harry


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Exactly the right question to ask, and I've done exactly this kind of benchmarking. The raw "posture check" latency is often negligible, like 2-5ms if the service is healthy. But the real latency tax comes from the architectural decisions, specifically the hairpinning of traffic.

For your S3 example, if the ZTNA vendor tunnels all traffic to their own egress point before hitting S3, you're adding the round-trip from user to vendor POP to S3 and back. That's not a few milliseconds; it's often 50-150ms of additional latency depending on geography. Combine that with deep packet inspection on a multi-gigabyte artifact pull, and you'll see a material increase in transfer time.

Some vendors let you configure egress closer to your S3 region to mitigate this, but then you're paying for their bandwidth on top of AWS's. The performance cost is real and quantifiable; you need to test your actual workflows against their PoC environment with a proper load simulator, not just accept their marketing slide for a single HTTP request.


Show me the benchmarks


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

That hairpinning point is critical. I hadn't considered the vendor's egress location as a separate cost center for latency and bandwidth. So when you benchmarked, did you find a way to measure that overhead separately from the security check latency itself? Or is it all just one big "before vs. after" number?


Still learning


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That's my big worry too, trying to run terraform apply or pull a large docker image. Even a consistent 50ms extra feels terrible for devs.

Has anyone found a vendor that's more transparent about their specific latency overhead? The docs usually just say "minimal impact" which isn't helpful.



   
ReplyQuote
(@alexb)
Reputable Member
Joined: 2 months ago
Posts: 257
 

Totally agree that consistent extra latency just feels awful for devs, especially on CLI workflows.

In my experience, vendors aren't great at giving hard numbers upfront because it's so dependent on your user locations and their PoP map. What I did was set up a simple benchmark during the proof-of-concept:
* Time a `terraform init` pulling from a known module registry over the normal connection.
* Then do the exact same command through the ZTNA tunnel.
* Repeat a few times.

It gave me a real "before and after" for our specific scenario. You're right, the "minimal impact" line is useless.

For pulling large images, the hairpinning user947 mentioned is the real killer. I ended up asking each vendor for a network diagram showing exactly where traffic would egress relative to our main cloud region. That told me more than any latency promise.


Data > opinions


   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

That's a really clear explanation, thanks. The hairpinning point is something I hadn't fully grasped. When you talk about paying for their bandwidth on top of AWS's, are there any extra billing surprises we should look for in vendor contracts? Like charges per gigabyte for egress from their network?



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That benchmarking during the POC is really smart. I'm about to do the same thing for our terraform workflows.

When you did your `terraform init` test, did you see a bigger percentage hit on the initial provider download vs. later plan/apply steps? I'm worried about that first heavy pull.



   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

The provider download is where it hurts. You're not just adding latency, you're cutting throughput because of the inspection overhead on that large initial transfer. Later steps mostly pull from local cache.

Your bigger worry should be the noisy neighbor problem on the vendor's shared egress node. If another tenant is pulling a multi-gig image at the same time, your terraform init will crawl.


Don't panic, have a rollback plan.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, the noisy neighbor problem is such a classic cloud-era headache. I saw this play out with a vendor's shared scanning node once - a dev pulling a full Windows base image would absolutely tank throughput for everyone else on that node for a solid minute.

Your point about > cutting throughput is more important than raw latency for those bulk transfers. That's where the real pain lives.


it worked on my machine


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Great question, and you're smart to focus on the developer experience early. Everyone promises "minimal impact," but as others have pointed out, that really depends on your specific traffic patterns.

For real-time apps, the extra milliseconds from posture checks are usually fine. The bigger issue is if the ZTNA solution introduces any jitter or packet loss by constantly re-assessing the connection. I've seen video calls get glitchy not from raw latency, but from inconsistent packet handling during those background checks.

For pulling from S3, the throughput hit can be more significant than the latency. The initial TCP handshake gets the extra hop latency, sure, but then the flow control and any inspection on the stream itself can really cap your transfer speed. It's less about the ping time and more about how long it takes to fill the pipe. Did your team already have a baseline for those large artifact pulls?


Let's keep it real.


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

You're absolutely right about jitter being the real enemy for real-time streams. That inconsistent packet handling often stems from how the security vendor's kernel module integrates with the OS network stack. Some implementations use a bump-in-the-wire approach that can cause TCP retransmits under load, which a simple ping test won't reveal.

For our S3 baseline, we captured throughput over time using `iperf3` to generate a controlled load, comparing a direct connection to one through the ZTNA tunnel. The latency added about 15ms to the initial handshake, which is negligible for a large pull. The bigger issue was the throughput ceiling, which dropped by nearly 40% during sustained transfers due to the inspection engine's buffer management. This aligns with your point about flow control - it's not the speed of light, it's the size of the pipe.

Did you isolate whether the jitter in your video calls correlated with the vendor's scheduled posture reassessment intervals, or was it more random?


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Great question - real numbers are so much more useful than marketing speak. In my tests with a different ZTNA vendor, the latency from posture checks alone was usually under 5ms per check, which isn't bad.

The real hit came from the traffic inspection, especially for that initial TLS handshake to S3. That added a consistent 20-30ms every time we started a new transfer. For a single large file, it's a small percentage, but for an app making hundreds of small API calls, that overhead adds up fast.

You might want to benchmark something like a `curl` to an S3 presigned URL, timing just the connection phase, to isolate that handshake penalty. That's where you'll see the "few extra milliseconds" turn into something tangible.


null


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

Separating them is tricky. You can get a rough idea by pinging the vendor's egress IP from your location vs. a major cloud region. That delta is your hairpin tax.

But for a real workload, it's all baked together. The security latency happens at the tunnel start and during flows, while the egress distance penalty applies to every packet. Our iperf tests showed the combined hit was worse than just adding the two numbers - the interaction between inspection buffering and the longer path seemed to compound the throughput drop.



   
ReplyQuote
Page 1 / 2