Correlating with transaction volume is smart, but that metric can be just as opaque as the Server/E2E split. A drop in volume during a spike could be throttling, or it could just be your client backing off. You're still inferring cause.
The real issue is that Storage Analytics logs have a hefty ingestion delay themselves, sometimes 5-10 minutes. By the time you see the pattern, the stamp might have already rotated the problematic node. It's a post-mortem tool, not a diagnostic one.
I've had more luck enabling the newer **Resource logs (diagnostic settings to Log Analytics)** and charting `TransactionResponseType` against latency. Seeing a cluster of `ServerBusyError` or `ClientTimeoutError` types, even at low count, often points you to the real culprit faster than minute-aggregated transaction volume.
That's a really practical point about the Resource logs. We switched to them for similar reasons, but I'd add that the key is setting up the diagnostic setting to stream directly to a Log Analytics workspace, not just the storage account itself. The delay shrinks to under a minute in my experience.
You're spot on about `TransactionResponseType`. Even a handful of those errors is a signal. We found the same pattern with `ClientOtherError` sometimes, which turned out to be the service side closing connections early under load, not our client.
automate everything