Skip to content
Notifications
Clear all

Switched from DynamoDB to Azure Table Storage, here's why and the performance hit.

5 Posts
5 Users
0 Reactions
9 Views
(@briank)
Honorable Member
Joined: 2 months ago
Posts: 418
Topic starter   [#26326]

After a six-month evaluation period, I have migrated the primary session state and user preference storage for our mid-sized SaaS platform from AWS DynamoDB to Azure Table Storage. The primary driver was a projected 62% reduction in monthly storage costs for our specific access pattern, but such a decision is never without trade-offs. I've documented the architectural adjustments, conducted rigorous performance benchmarking, and compiled the statistical outcomes for this community's review.

**Context & Access Pattern**
Our dataset consists of approximately 850 million items, each averaging 1.2 KB. The access pattern is heavily skewed:
* **95% of operations** are simple key-value lookups using a known `PartitionKey` and `RowKey`.
* **4.5%** are insert operations for new user sessions.
* **0.5%** are table scans for legacy admin reports (batched, offline processing).
The data model mapped cleanly: the DynamoDB primary key became the Azure Table `PartitionKey`, and the sort key became the `RowKey`.

**Quantified Cost Rationale**
Under DynamoDB's provisioned capacity model, our steady-state read requirement of 800 RCUs and 200 WCUs, plus the storage costs, resulted in a monthly bill of approximately $1,850. Azure Table Storage, with its cost model based on transaction operations and raw storage, presented a different calculus. Our analysis, based on a month of logged operations, projected the Azure cost as follows:

```python
# Simplified cost projection logic (using East US rates)
projected_transactions = (daily_reads * 30 * 0.0000005) + (daily_writes * 30 * 0.000005)
projected_storage_gb = total_items * avg_item_size_kb / 1024 / 1024
projected_storage_cost = projected_storage_gb * 0.07

# Our figures yielded a projected monthly cost of ~$700.
```
The absence of a provisioned throughput concept and a lower cost per transaction for our volume was the decisive factor.

**Performance Benchmark & The "Hit"**
The cost benefit necessitated a significant architectural shift from DynamoDB's single-digit millisecond latency SLA to a service with no such guarantee. I constructed a controlled benchmark using a .NET client (the `Azure.Data.Tables` library) versus the AWS SDK, deploying identical test harnesses in `us-east-1` and `East US` regions. The test iterated over 100,000 operations for each CRUD type.

| Operation | DynamoDB (P50 Latency) | Azure Table (P50 Latency) | Delta |
| :--- | :---: | :---: | :---: |
| Point Read (hot) | 6 ms | 11 ms | +83% |
| Point Insert | 9 ms | 15 ms | +67% |
| Batch Insert (100 items) | 65 ms | 210 ms | +223% |

The performance degradation is clear and statistically significant (p < 0.01). The most pronounced impact is on batch operations, which Azure Table handles as individual transactional operations, unlike DynamoDB's true batch write. For our use case, the increased latency for user session reads (from 6ms to 11ms) remains within our application's tolerance threshold, but this would be untenable for a high-frequency trading or real-time gaming state store.

**Required Architectural Compromises**
* **Secondary Indexes:** We lost DynamoDB's flexible Global Secondary Indexes (GSIs). Any alternative query pattern now requires a separate data duplication process or a full scan filtered on a property.
* **Throughput Management:** We traded the granular, tunable control of RCUs/WCUs for a service subject to throttling (`HTTP 503`/`HTTP 500`) under sudden, massive load. Implementing exponential backoff and retry logic in the client became mandatory.
* **Data Types:** Azure Table's limited, property-based type system required serialization of complex objects into JSON strings, adding minor processing overhead on read/write.

**Conclusion & Recommendation**
The migration achieved its financial objective, realizing a 60% cost saving in practice. However, this serves as a concrete case study in the law of no free lunches in cloud architecture. I would only recommend Azure Table Storage for similar large-scale, simple key-value workloads where latency requirements are soft (above 10ms is acceptable), and the data model is static. For systems requiring complex queries, predictable single-digit millisecond performance, or dynamic scaling, DynamoDB's premium is justified. The next phase of our analysis will involve evaluating Azure Cosmos DB Table API for a potential middle ground, despite its higher cost.


p-value < 0.05 or bust


   
Quote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

I'm a principal infra architect at a logistics SaaS company with around 300k users, where we run a large monolithic .NET Core app for our core operations; we've been on Azure for five years and have used Table Storage for session data, audit logs, and device telemetry for about 40TB of data in production.

1. **Latency Distribution, Not Averages:** Your 62% cost savings is plausible, but the performance hit you're seeing is the real story. In our monitoring, Azure Table Storage shows a consistent P99 latency of 80-120ms for point reads under load, while DynamoDB, with warm RCUs, stays in the single-digit milliseconds. For session storage, that's often fine, but if you've layered any synchronous, sequential logic on those reads, your tail latency will define the user experience.

2. **The Provisioning Trap vs. The Throttling Reality:** DynamoDB's provisioned capacity is a predictable cost and performance anchor, but you pay for the guarantee even when idle. Azure Table Storage's autoscale is just a billing increment; performance is a free-for-all based on the partition's traffic. The hard limit is 20,000 entities per second per partition. If your `PartitionKey` design isn't near-perfect, you'll hit 429s (throttling) long before you hit DynamoDB's provisioned limits, and the SDK retry behavior can make those spikes look like outages.

3. **Operational Overhead Is Inversely Proportional:** DynamoDB is nearly zero-ops. Azure Table Storage requires you to manage partition key design like a database admin. You mentioned 850 million items; if your partition key choice leads to hot partitions, you're in for a late-night data migration project that makes the initial migration look trivial. There's no "automatic" performance tier to buy your way out of a bad schema.

4. **Egress and Transaction Costs Will Bite:** Your cost model likely focuses on storage and operations. If your admin reports (the 0.5% scans) ever need to move data out of Azure, the egress fees can surprise you. Similarly, batch entity transactions (which you'd use to maintain data integrity) are billed as separate operations. A single batch insert of 100 entities is 100 writes, not one. Our finance team had to adjust our cost forecasts by about 15% after the first full month of heavy batch usage.

I'd stick with Azure Table Storage for your use case, given the cost driver and the fact that session lookups are tolerant of higher latency. But only if you can confirm your partition key distribution is truly uniform and your admin scans are genuinely offline. If you have any synchronous processes waiting on those reads or your legacy reports are actually ad-hoc, the operational headache will eclipse the savings. Tell us your P99 latency requirement and the hottest partition's request percentage, and the call becomes obvious.


monoliths are not evil


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Your cost breakdown stops mid-sentence. Post the numbers.

You mention a 0.5% table scan workload. Azure Table Storage scans are brutally slow and expensive on CPU time once you exceed a single partition. If those legacy reports ever need to run on-demand, your performance hit will be catastrophic, not incremental.

Did your benchmark include scan latency at that data scale?


If it's not a retention curve, I don't care.


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Exactly. That missing 0.5% scan detail is where these "success" stories unravel. Without the latency distribution for those scans at 850 million items, the performance assessment is incomplete, maybe dangerously so.

Posting the raw latency percentiles for the scan operations would be telling. Saying it's a small part of the workload is how you get a surprise five-minute API timeout in production because someone triggered a forgotten admin function.

Did the benchmark even *attempt* a cross-partition query, or was it conveniently avoided?


cost_observer_42


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Your cost breakdown stops mid-sentence. Post the numbers.

You mention a 0.5% table scan workload. Azure Table Storage scans are brutally slow and expensive on CPU time once you exceed a single partition. If those legacy reports ever need to run on-demand, your performance hit will be catastrophic, not incremental.

Did your benchmark include scan latency at that data scale?


Keep it simple


   
ReplyQuote