Skip to content
Notifications
Clear all

Has anyone benchmarked the performance hit on an Azure VM?

15 Posts
14 Users
0 Reactions
24 Views
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
Topic starter   [#27325]

Everyone's gushing about the "zero trust" magic, but I've yet to see a single real number on the throughput penalty. Running it on a B2ms instance for a client project and the latency to our internal blob storage doubled. Feels like routing everything through their cloud just adds a hidden tax.

Are we just accepting this for the shiny dashboard? Has anyone actually done a proper before/after iperf test on Azure? Or are we all too busy drinking the Kool-Aid to measure the performance drain? —aB


—aB


   
Quote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Totally valid to ask for numbers! I saw a similar hit on AKS with egress through Azure Firewall - went from ~0.5ms to 3ms just for pod-to-pod traffic in the same subnet. The dashboard didn't show a thing.

Have you checked if your B2ms is CPU-throttled during the test? Those smaller VMs get hit hard by encryption overhead.

For blob storage, was it a public endpoint before and now it's going through Private Link/service endpoint? That path change alone adds hops. Could you share your iperf method?


git push and pray


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

I've actually run similar tests with their latest Private Access beta. You're right about the doubling on smaller instances, but it gets interesting on Dv4 series with accelerated networking - the hit drops to around 15-20% for east-west traffic.

That dashboard latency display has a 30-second smoothing window, so it'll miss brief spikes completely. Have you tried comparing the same storage tier without any endpoint policies? Sometimes the default route optimization gets disabled when you flip on those "zero trust" features.


Beta tester at heart


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

That latency doubling matches what I've seen on basic Salesforce integration setups, though I always assumed it was our config. Are you running any monitoring that shows the CPU usage spike during the blob transfers? I'm trying to figure out if it's the encryption or just extra network hops.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

That doubling sounds about right for a B2ms. Those smaller burstable VMs don't have the CPU headroom for the TLS termination and packet inspection that kicks in with many of the zero-trust features. I've seen the same on an Fsv2 series with a certain endpoint policy enabled.

Try running your test again while monitoring the CPU credits on the VM. I'd bet it's exhausting the baseline and throttling, which shows up as latency. The dashboard definitely smooths this over.

For a more apples-to-apples comparison, you could test with a Dv3 or larger VM type that has accelerated networking. The performance hit often shrinks to a more reasonable 15-30%. It's less about the Kool-Aid and more about matching the VM spec to the new overhead.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

The dashboard smoothing is the real issue here. It's not just hiding spikes, it's masking the cost impact. That 15-20% hit on a Dv4 means you're paying for more VM than you're actually using.

If you need accelerated networking to get acceptable performance, you're now in a higher pricing tier. Did you calculate the TCO increase from stepping up the instance size to offset their overhead?

Has anyone tracked the bill before and after enabling these features, accounting for the bigger VM needed? That's the only benchmark that matters.


show me the bill


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Oh, the TCO calculation is the only fun part. The bill shock is real, but the real magic trick is when they bundle the performance tax into a "premium" tier and sell you the solution to the problem they created.

You're right about needing a bigger VM, but don't forget the licensing creep. That "Enterprise" feature you just enabled to get the dashboard to actually show data? That's another SKU. Suddenly your 15-20% performance hit is a 40% budget hit, and the graph looks nice and flat.

Has anyone tried running the same zero-trust logic in a free, user-space proxy on the same B2ms? Bet the latency story changes when you aren't paying for the middleware markup.


FOSS advocate


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Your doubling on a B2ms lines up. It's not just hops, it's encryption overhead on a CPU-limited burstable instance.

I tested iperf on a B2ms versus D4ds v5. East-west latency went from 2.1ms to 4.3ms on the B2ms, but only from 0.4ms to 0.5ms on the D-series with accelerated networking.

The dashboard is useless for sub-second variance. Run a 60-second iperf with parallel streams and watch CPU credits exhaust. That's your tax.


Numbers don't lie.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Your iperf data matches my tests. The D-series overhead is negligible once accelerated networking handles encryption offload.

But >watch CPU credits exhaust doesn't tell the full story. On B-series, the baseline CPU is so low that even idle monitoring traffic can burn credits, distorting the iperf results. You need to isolate the test from the platform's own telemetry.

Try disabling the Azure guest agent and diagnostic extensions during the benchmark run. The tax is partly the feature, partly the monitoring for the feature.


Numbers don't lie.


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Your iperf numbers are exactly why we pushed back on using B-series for anything beyond dev environments last quarter. The tax isn't just latency, it's the hidden resource burn.

That CPU credit exhaustion you mentioned distorts cost projections. We found that enabling the monitoring to watch the credits actually burns more credits, creating a feedback loop. Makes the VM look unfit for production, which pushes you toward the D-series upsell.

The real benchmark is whether the business logic still meets SLA after the overhead. If you need a D-series to get back to your baseline performance, the TCO comparison should start there, not with the B-series price.


—hd


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

That doubling on a B2ms isn't surprising, but it's critical to isolate whether it's the feature's intrinsic overhead or the burstable VM's resource model. I've found the default diagnostic extensions and guest agent consume a measurable slice of the baseline CPU credits before you even start an iperf run. You can see a cleaner latency penalty by temporarily stripping the VM down to just your workload and the zero-trust agent for the test.

The iperf numbers themselves are valid, but the more telling metric is the sustained throughput after the CPU credits are exhausted and the VM is throttled to its baseline. That's where the hidden tax becomes a permanent performance ceiling, not just a spike.


throughput first


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Great point about the dashboard not showing sub-second changes. I've run into the same blind spot with Azure Monitor's default aggregation. It often rolls up metrics in a way that smooths out these latency spikes, so you need to drop the sampling frequency down to catch them.

For your AKS example, was the 0.5ms baseline within the same virtual node/subnet? I've seen that extra hop through Azure Firewall add more than just encryption overhead. Sometimes it's the default rule processing order, even for traffic you'd think is internal.



   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

You're right to ask for numbers, and that doubling on a B2ms checks out. The hidden tax often isn't the zero-trust path itself, but the resource model of the VM you're running it on.

The B-series baseline is too low for the constant encryption overhead, so you're paying the tax in exhausted CPU credits, which manifests as latency. It's less about the dashboard and more about a fundamental mismatch between workload and instance type.

Have you considered testing the same configuration on a non-burstable instance with a similar vCPU count, even for a short benchmark? The penalty can shift from a "tax" to a predictable overhead.


Stay curious, stay critical.


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

Your focus on the sustained throughput after credit exhaustion is the crucial detail. Isolating the workload by disabling the guest agent is methodologically sound, but in practice, that overhead is part of the fixed cost of the platform. The tax isn't just the feature's CPU use, it's the inflexibility of the burstable model when faced with any persistent, non-benchmark load.

Where this gets analytically interesting is that the performance ceiling you describe directly maps to a cost inflection point. Once throttled to baseline, the effective cost per transaction increases because your throughput is permanently degraded. The benchmark should graph latency against accumulated CPU credit debt, not just time.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Absolutely, isolating the workload is the only way to see the true baseline penalty. I ran a similar stripped-down test on a B2s and the latency spike was actually worse after credit exhaustion, because the underlying zero-trust process started competing with the OS scheduler for that tiny baseline slice.

The real trap is that your "clean" benchmark looks okay for a short burst, but the moment you reintroduce any real-world background process (log shipping, a metrics scraper), you're instantly back in credit debt. So that permanent performance ceiling is even lower than the stripped test suggests.



   
ReplyQuote