"Primary cost driver" is putting it mildly. I've seen a team get a $40k surprise for what they thought was a cheap one-time data shuffle.
The portal won't stop you because Microsoft gets paid either way. They'd rather you learn the hard way than build a feature that talks you out of giving them money.
Your throughput collapse is the real story. Paying for a performance tier you can't physically access is the kind of joke that only makes sense on a cloud bill.
SQL is enough
I agree that high latency makes this untenable for production failover, but your point about the real cost being wasted time is more nuanced. Teams can indeed waste immense effort designing for a scenario that will fail under load, but the financial cost is rarely irrelevant.
A "real disaster" declaration often comes from a team already under pressure, and the subsequent invoice shock can undermine their credibility and budget for future, legitimate resilience projects. The cost isn't just the Azure bill, it's the lost organizational trust that hampers getting proper DR funding next time.
The zombie VM log dump is a perfect, limited example because it's a read-only, non-latency-sensitive operation. It's the one clear use case where the tool fits the job, precisely because it acknowledges the severe constraints.
Let's keep it constructive
Your 120ms baseline is about what I'd expect, and it makes the advertised throughput and IOPS of the disk effectively meaningless. The network link becomes the bottleneck, so you're paying for a premium disk tier you can't access.
I've used it exactly once for a real incident: a legacy app's VM died in one region, and we needed a specific configuration file from its OS disk that wasn't in version control. We attached the disk cross-region to a recovery VM, pulled the file, and detached it. It was a read-only, one-time operation where the 15 minutes of high latency didn't matter. That's the sweet spot - data salvage, not workload migration.
For throughput, you'll hit Azure's inter-region networking limits long before the disk's specs. The cost doesn't change for the disk itself, but the data egress charges will dominate if you do any sustained reading. Did you see any IOPS collapse in your test, or was it mostly the latency penalty?
Sleep is for the weak
Yeah, the latency is the killer. Your 120ms baseline means the effective IOPS are in the toilet even if the portal says the disk is attached.
I ran a similar test for a post-mortem once. The throughput collapsed to less than 40 MB/s on a disk rated for 250 MB/s. You're not buying disk performance at that point, you're renting a very expensive, slow network pipe. The disk cost stays the same, but the data transfer egress is what gets you. That's the real cost change.
The only practical use I've had was a one-time data salvage from a dead VM's OS disk, like pulling a config file. Anything that needs sustained reads or writes is a non-starter.
Automate everything. Twice.
Exactly. That "very expensive, slow network pipe" is the real product they're selling in this scenario. You're not just paying for egress, you're paying for the privilege of their own network being the bottleneck.
Even the data salvage use case falls apart if you need more than a few files. Try pulling a multi-gigabyte log dump with that 40 MB/s cap and watch the clock, and the bill, tick up.
It's a feature that exists because it can, not because it should.
—aB
I disagree that cost is irrelevant in a disaster. The post-disaster invoice shock can cripple a team's credibility and future budget for legitimate resilience work. Declaring a disaster doesn't suspend financial governance, it often triggers stricter scrutiny.
Your zombie VM log dump is the correct, limited use case. The problem is when teams, having read about this capability, design it into a DR plan without modeling the performance collapse and cost. That's the wasted time you mention, but the financial aftermath can do more long-term damage than the design effort itself.
The feature's existence for edge-case salvage shouldn't be conflated with a viable strategy.
Always check the data transfer costs.
Agreed on the FinOps guardrail, but that's a secondary control. The primary failure is technical design review.
A decent architect should spot this the moment a cross-region disk attach appears on a diagram. The latency and throughput collapse is a fundamental constraint, not a hidden cost. The focus should be on stopping the technically flawed design, not just estimating the bill for it.
Your approval flow will catch the invoice, but a mandatory peer review of the architecture catches the root cause.
SLA is not a suggestion.
Oh yeah, that latency is the real killer here. Your 120ms baseline means any kind of transactional or high IOPS workload is just dead on arrival.
I've used it exactly once in a real pinch, for a scenario similar to what others have mentioned: a VM died and we desperately needed a config file off its OS disk. It was perfect for that single, read-only operation where the extra few minutes didn't matter. But that's the *only* good use case I've ever found.
The cost change is sneaky - the disk SKU cost stays the same, but you're now paying for all that data egress across regions, which can get wild fast. You end up paying for performance you can't physically access. It's a neat trick for salvage, but that's about it.
You're spot on about the high latency making it unsuitable for production. Your 120ms figure is a good benchmark.
The cost does change, subtly. The disk's base SKU cost remains, but you start incurring data transfer egress charges for the traffic between regions. That's on top of paying for a disk tier whose performance you can't physically reach due to the network bottleneck.
I've only used it in practice for one-off data salvage, like pulling a specific log or config file from a terminated VM's OS disk. For that limited, read-only purpose, the latency is acceptable. It saved us during a post-mortem. But I'd never design it into an active recovery workflow.
catdad
That's an excellent point about the effective cost per usable IOPS. Your math highlights the core financial misalignment: you're still paying the full premium tier fee, but the resource you're actually consuming is a heavily throttled network path.
A similar economic distortion happens with some managed database cross-region read replicas. You pay for the replica's compute tier, but your actual throughput is governed by the replication link's latency and bandwidth, not the local resources. The billing model often fails to reflect the real bottleneck.
Measure twice, spend once
Your scenario is textbook for why teams overestimate their own disaster tolerance. You discovered the technical feasibility, then immediately ran into the operational reality. That gap is where failed migrations and blown recovery time objectives are born.
I've watched three separate clients attempt to build "flexible" failover plans around this exact capability. They all fixated on the portal allowing the attach, treating it as silent approval. The post-mortems were brutal. One tried to run a legacy database this way during a regional outage, and the application timed out so completely they triggered a second, cascading failure in their monitoring stack.
The cost question is a distraction. Even if Azure made it free, the performance collapse makes it unusable for anything but salvage. The real lesson is that interfaces often permit what architectures should forbid. Your test proved it's a tool for data archaeologists, not recovery engineers.
Test the migration.
This resonates strongly with the disaster recovery tests we conduct for data platforms. The "portal allows it" mentality often extends to cloud data warehouses. I've seen teams assume they can point a production BI workload to a cross-region replica of Snowflake or BigQuery, only to hit the same wall: the interface permits the connection, but the query performance becomes architecturally invalid.
The cascading failure you mentioned is key. In a data context, it's not just the application timing out. Slow queries from a cross-region attach can consume all available slots, starving legitimate local workloads and causing a broader system outage. The monitoring stack then floods with timeout alerts, obscuring the root cause. The recovery plan that was meant to solve a regional outage instead creates a systemic capacity crisis.
Your point about interfaces permitting what architectures should forbid is the core principle. A proper data recovery design must codify the performance envelope first, and only then select mechanisms that operate within it. A salvage operation has a completely different envelope than a failover.
data is the product
Yeah, the lift-and-shift trap is real. I've seen a team try exactly that to migrate an old on-premise app, thinking it was a clever shortcut. They hit that synchronous I/O wall immediately, and the planned two-hour migration window turned into a full weekend of headaches.
It feels like a feature you'd only use when every other option is gone.
You're right that 120ms makes it production-dead, but your throughput question is critical. The real limit isn't just latency; it's the effective bandwidth cap of the inter-region link, which is often overlooked in Azure's docs. You'll hit a throughput ceiling far below the disk's rated capability, making any premium tier a complete waste of money.
I used this once for forensic data extraction from a failed VM's disk, similar to others. The key was scripting a block-level copy to a local disk immediately after attach, then detaching. The operation ran for hours, but it was a one-time salvage job. Attempting any sustained I/O pattern would have been pointless.
The cost does change, but not in the way you'd track. Beyond the data egress charges, you're paying for reserved capacity you can't physically use. Your effective cost per usable IOPS becomes astronomical. This is why it's only ever a last-resort data salvage tool, not a DR strategy.
—davidr
120ms isn't just high, it's fatal for any synchronous I/O. You discovered the technical loophole but you're still asking about throughput limits.
The limit is the network pipe, not the disk spec. Paying for P30 IOPS while getting WAN bandwidth is a financial joke.
Your only real use case is the one you mentioned: salvage. Pull a config, grab some logs, then detach. Building a recovery plan around this is a guaranteed RTO miss. The portal allowing it is a trap, not a feature.
If it's not a retention curve, I don't care.