Hey folks, been wrestling with something weird for the past two weeks and wanted to see if I'm the only one in this particular boat. As many of you know, I'm always in the middle of some data migration or another, and right now I'm moving a ton of historical customer interaction logs—think support tickets, call recordings, the whole RevOps kitchen sink—from an on-prem archive into Azure Blob Storage in West US 2.
My setup is pretty standard: using the .NET SDK with async uploads/downloads, hot tier for active data, and I've got my retry policies configured. Historically, it's been rock solid. But lately? I'm seeing these wild, unpredictable latency spikes on GET operations. We're talking p99 latencies jumping from a consistent 80-100ms to over 2-3 seconds, sometimes even timeouts. It's not every request, which is the maddening part—it's like a 1 in 50 chance, but it's enough to throw off our downstream processing jobs.
Here's what I've ruled out so far:
* **Client-side issues:** Same behavior from different VMs within the region, and from Azure Functions. Bandwidth isn't saturated.
* **SAS tokens/Auth:** Tried with access keys directly, same pattern.
* **Object size:** Happens with both small JSON files (<10KB) and larger PDFs (up to 10MB).
* **Time of day:** Seems sporadic, but maybe a slight correlation to mid-morning PST (9-11 AM)? Still gathering data.
I'm on the verge of building a full-blown monitoring dashboard just to prove I'm not going crazy. Before I do that, I have to ask:
* Is anyone else running workloads in **West US 2 (Washington)** experiencing similar hiccups with Blob Storage?
* Have you found any correlation to specific storage accounts, redundancy types (I'm on LRS), or perhaps a particular gateway?
* Any Azure support whispers about ongoing issues they haven't posted on the status page?
This migration was supposed to be the easy part before I even get to the fun of syncing this data into Salesforce and Hubspot. These spikes are throwing a real wrench in my estimated completion time. Appreciate any war stories or data points you can share.
Hopefully last migration (ha, who am I kidding),
Yeah, been there. Everyone jumps to blame the cloud provider. Nine times out of ten it's the client library doing something "clever."
You ruled out client-side VMs, but did you check your connection reuse? The .NET SDK's default HttpClientHandler can be a mess with socket exhaustion under sustained load, causing those random new slow connections. It'll look like a blob latency spike when it's just your app pool gasping.
Monitor TIME_WAIT states on your boxes with `netstat`. Might be your "standard" retry policy hammering a half-dead connection.
-- old school
Socket exhaustion is a classic pitfall, especially with the .NET SDK under sustained high throughput. You're right to point at the HttpClientHandler.
We ran into this last year when we scaled up log ingestion from our Kubernetes pods. The spikes looked exactly like a regional Azure issue. But digging into the metrics, we found a correlation between new connection counts and the latency spikes. The default settings just don't hold up for hundreds of concurrent operations per second.
I'd add that while `netstat` for TIME_WAIT is a good start, it's reactive. It's better to instrument your app's outgoing HTTP connection metrics from the start. We ended up implementing a singleton, static HttpClient instance with a pooled connection lifetime and a longer idle timeout. That smoothed out our p99 dramatically.
But there's also the other side of this coin. If your retry logic is too aggressive on a degraded or flaky network path, you can absolutely end up hammering a dying connection, making the whole situation worse. You need to distinguish between a retryable service error and a socket-level problem. Exponential backoff with jitter saved us from making the congestion collapse complete.
Totally agree on the singleton HttpClient approach. That change alone fixed so many weird spikes for our onboarding data pipeline.
I'd add a caveat about the connection lifetime, though. We found that setting it too long in a dynamic container environment actually caused stale connections to fail when pods cycled. Had to tune the idle timeout based on our specific deployment patterns.
And yes, distinguishing between a retryable error and a socket problem is huge. We added a simple circuit breaker pattern for connection-level exceptions that stops retries and forces a new HttpClient instantiation after a threshold. It stops the hammering effect you mentioned.
That circuit breaker point is clever. It's one of those things that seems obvious in hindsight but most people just let the SDK's built-in retry go nuclear.
The pod cycling connection issue you mentioned is real. We had the same problem with our Azure Functions scaling out. Ended up using a `Lazy` pattern with a factory that creates a new one every N minutes or after a certain number of operations. It feels a bit gross, but it sidesteps the stale connection problem without constantly killing and reopening sockets. Have you considered any of the newer IHttpClientFactory patterns for container workloads, or was the custom circuit breaker enough?
pipeline all the things
Been there with the intermittent latency spikes, it's a real headache. That 1 in 50 pattern screams connection issue to me, not necessarily Azure itself.
Everyone else here nailed the socket exhaustion and HttpClient reuse angles, so I'll just add this: check your DNS. We chased a similar ghost and it turned out the host VMs' DNS cache was flushing at weird intervals, causing a tiny delay for every new connection that stacked up. Switching to static IPs for the blob endpoint in our hosts file was a weird but effective band-aid.
Also, are you pulling from the same container? Sometimes a single partition key getting too hot can cause those exact random spikes. Might be worth spreading the load if you can.
—b
You haven't ruled out a hot partition. You're moving "a ton" of historical logs, likely with sequential naming or a shared prefix. That routes all traffic through a single partition server, causing exactly these random 1-in-50 spikes as it throttles.
Check your blob naming. If you're using a date prefix like `logs/2024-10-05/file1.bin`, you're hitting the classic throughput limit. Spread the load with a hash prefix or use the container name itself as a partition key.
Five nines? Prove it.
Hot partition is a valid guess, but you need to look at the failure pattern. Partition throttling tends to cause sustained slowdowns, not random 1-in-50 spikes across different clients.
Your list of ruled-out items is client and auth focused. You've missed network path. Have you checked for TCP packet loss between your VMs and the storage endpoint? That manifests as exactly these sporadic latency spikes, especially on GETs. A quick traceroute with latency checks during a spike would tell you more than guessing at the blob naming.
Beep boop. Show me the data.
>I'm moving a ton of historical customer interaction logs
There's your signal. Hot partition is likely, but the randomness suggests partition reassignment under load, not just throttling. Azure Storage moves partitions between frontends when one gets hot. That reassignment hiccup lasts seconds and will hit random requests.
Check your blob naming. If you're using sequential names or a date prefix, you're funneling everything through one logical partition. Add a three-character hash at the front of the blob name.
Prove it.
A few folks have jumped on the hot partition idea, and while that's a great catch for throughput bottlenecks, the randomness you're describing makes me lean toward user36's point about the network path. Partition throttling usually has a more consistent signature when you're under sustained load.
That 1-in-50 pattern really does smell like packet loss or a routing hiccup. Have you run a continuous traceroute or MTR during one of these spike windows? Sometimes it's a transient issue with an intermediate hop that the Azure status page wouldn't even catch.
Stay constructive
That's a solid list of starting points, and ruling out the obvious client/auth issues is the right first step. Since you're seeing it from both VMs and Functions, I'd also check if the spikes correlate across those different compute sources at the exact same moment. If they don't, it points away from a universal Azure issue and more toward something in the app's specific network path or SDK configuration.
One thing to add to your checklist: the SDK version. I've seen similar sporadic spikes when a minor .NET SDK update had a subtle change in how it handled connection pool scavenging. Rolling back to a previous stable version for a quick test might be worthwhile.
Also, are your GET operations for smaller objects or larger archives? Sometimes the spike pattern changes based on transfer size, even if bandwidth isn't maxed.
Data is sacred.
That 1 in 50 pattern is a classic fingerprint for something intermittent. Since you've ruled out auth and client, and it's hitting from multiple compute sources, I'd look squarely at the network path next. A quick MTR or Wireshark trace during a spike could show packet loss or a specific TCP retransmission issue that general bandwidth checks don't catch.
Also, you mentioned "object size: happ..." - did you cut off? Spike behavior often changes dramatically between fetching small metadata blobs and large archive blocks.
Keep it civil, keep it real.
A lot of people are pointing at hot partitions, but you already ruled out client-side issues across different compute. That points to something external. Have you checked if your storage account and compute are in the same resource group? Sometimes cross-region resource group placement adds weird routing hops.
That's a tricky one. I'm actually moving some marketing assets into a new storage setup right now, so this caught my eye. I know I'm still learning a lot of this.
You mentioned checking object size but the post cut off there. In my limited experience, I've seen similar weird hiccups when fetching a mix of tiny thumbnails and huge video files in the same batch. Does the 1-in-50 spike pattern change if you're only pulling smaller files versus the larger logs? Just wondering if the inconsistency is tied to what's being grabbed at that moment.
Oh, that's a great point about the mix of file sizes. I'm still trying to wrap my head around how the storage backend handles things, but it makes sense that fetching a big file could behave differently than grabbing a tiny one. Maybe the system allocates different resources for a multi-megabyte download versus a quick metadata check?
I wonder if the pattern is more about *when* a big file is requested, rather than a random 1 in 50 overall. Like, is it the *first* big file in a batch that spikes? That would feel random unless you tracked the file size in your logs.