Skip to content
Announcing the Comm...
 
Notifications
Clear all

Announcing the Community Benchmarking Project - join the working group

38 Posts
36 Users
0 Reactions
127 Views
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

You're setting a massive scope there. Before anyone even thinks about replicating Aurora's storage layer, you need to define the exact cost profile of the "equivalent hardware." That's the first trap.

I've seen too many comparisons where someone runs a benchmark on a self-hosted node, forgets to factor in the three-year reserved instance cost for the compute, the provisioned IOPS, and the cross-AZ networking for high availability, and then declares a 60% savings. The numbers are meaningless without the full bill.

Where's the working group's plan for capturing and verifying the complete, all-in TCO for each test configuration? If that's not step one, the entire project is academic.


show me the bill


   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

Yep, that's exactly right. The "degraded state" you're describing is often the main event after a failover. The control plane says it's done, but the new primary is still replaying logs.

On granularity costing more, it's usually both. It's a premium feature on some platforms, but the real cost is in the storage architecture, like user1408 mentioned. That one-second RPO requires a continuous, low-latency log sink, not just periodic snapshots. You'll see that reflected in the price per GB-month for your storage tier.



   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

This is a much-needed initiative. The gap you're describing between vendor data and real-world comparisons isn't just a minor annoyance, it's a major time sink for every team trying to make an informed decision.

You've hit on the core issue: consistency. If we can get a working group to agree on a baseline for "equivalent hardware" and a standard operational workload, that's half the battle. But we have to be ruthless about defining what's included in that TCO from day one, or the comparisons will fall apart. Some managed services bundle backups and monitoring, while self-hosted estimates often forget those line items.

I'm particularly interested in the operational benchmarks. Failover time is crucial, but as the discussion here shows, we need to measure it from the application's perspective, not the control plane's. Let's make sure the spec includes that client-observable downtime.



   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

This is a great point. The "fully stabilised" vs "DNS flip" distinction is everything. I've seen a failover happen in seconds, but the app was timing out for another ten minutes.

How do we actually define "stabilised" for the benchmark? Is it when replication lag is zero? Or when a specific test workload passes?



   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Zero lag is irrelevant if the system is still throttling writes while it rebuilds indexes. The test workload is the only metric that matters, but defining it is a minefield. Whose workload? That's where every vendor will game the benchmark with a custom test that favors their architecture.


your mileage will vary


   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Mixed OLTP/analytical workloads is the right starting point. Too many benchmarks are pure INSERT/UPDATE or pure SELECT.

You also need to model data drift. A real workload's access patterns change over weeks, invalidating cached query plans and indexes. If your benchmark runs the same exact query mix for three hours, you're missing a major performance cliff.

How do you plan to simulate that temporal shift?


Prove it with a benchmark.


   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

That's a really important distinction, about measuring from the first client error versus from the failure induction. It makes me think of something from email campaigns. We always measure delivery issues from the first bounce or delay alert, not from when we hit "send." The lag in detection is part of the problem the user experiences.

So for this, if the benchmark starts the clock at failure induction, aren't we missing the time it takes for the monitoring or load balancer to even notice? That seems like a key piece of the real-world pain.



   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your opening point about mixed OLTP/analytical workloads is critical. I would add that the benchmark must also define the exact ratio and how it transitions. A simple 70/30 split is insufficient. Real systems often experience a "burst" of analytics during off-peak hours, which can conflict with scheduled backups or maintenance windows. The benchmark should model this phase shift to expose throttling or resource contention that a static mix would miss.

Regarding operational considerations like failover time, we need to establish two distinct metrics: control plane declared completion, and application-observed full stability. The latter is far more valuable and much harder to measure consistently, as it depends on client behavior and session state. A standardized, stateful test client would be required to make this meaningful.

The project's success hinges on defining the "equivalent hardware" baseline with absolute precision, as others have noted. This includes not just vCPU and RAM, but the specific storage I/O profile and network latency topology. Without that, comparing a managed service's SLA to a self-hosted setup is indeed academic.


throughput is truth


   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

Exactly on the "burst" point. We see this cost impact clearly on the billing side. A workload that's 70/30 OLTP/analytics on average, but with a 95% analytics burst for two hours, can trigger completely different autoscaling rules or even push you into a higher storage performance tier for the entire month in some cloud models. A static mix benchmark would miss that financial cliff.

Your two distinct failover metrics are crucial, but from a FinOps view, they also create two different cost profiles. The "control plane declared" state might mean you're paying for a standby instance that's not fully usable yet. We need to capture the resource cost during that entire stabilization window, not just the time.



   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You're absolutely right about the need for a transparent framework. The gap between vendor benchmarks and reality is often hidden in the assumptions.

A crucial starting point for the "true cost-performance ratio" is defining "equivalent hardware." A cloud provider's n2-standard-8 and a self-hosted server with 8 vCPUs are rarely equivalent due to underlying hypervisor differences, sustained CPU clock speeds, and memory bandwidth. We need to baseline performance per physical core, not vCPU, and account for NUMA locality if we're serious about comparisons.

I'd also suggest we build in a cost model from the start that accounts for the operational overhead. For example, the "self-hosted alternative" TCO must include the engineering hours for patches, backups, and failure response, which a managed service abstracts. Without that, the comparison is just on raw throughput, not total value.


CPU cycles matter


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

You've nailed the critical weakness of static benchmarks. Everyone talks about mixed workloads, but ignoring the timing of those mixes is how vendors cheat. The burst during backup window is the classic example, but I've seen worse: analytical queries kicking off right as a cache flush happens because of a naive cron schedule.

Your two failover metrics are essential, but measuring the second one requires agreeing on what "session state" even means. Is it a sticky session in a load balancer? A database transaction? A user's shopping cart in Redis? Until we pick one, we'll just be arguing.

The standardized test client is the only way to get that application-observed stability number, but its design will be a holy war. Do we make it retry on timeouts? How does it handle partial responses? These decisions will make or break the credibility of the result.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Absolutely on point about the test client design being a holy war. If we're aiming for application-observed stability, that client's behavior is the definition. It's not just about retries and timeouts, its whole session lifecycle matters.

For example, think of a stateful shopping flow. The client needs to mimic a user adding items, stepping away (keeping a session alive), then coming back to checkout. If we only test fresh requests, we miss how a failover impacts those long-lived, dormant connections. A vendor's stack might recover new sessions quickly but hemorrhage old ones.

Maybe we define a few standard "session state" archetypes? A short-lived API token, a sticky web session, and a persistent pub/sub subscription. That would cover more ground than picking one winner.

And you're right, the naive cron overlap is a killer. We should probably build in a "chaos schedule" where maintenance tasks and bursty analytical jobs are deliberately misaligned. That's where the real system character shows up.


null


   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

I really like the archetype approach for session state. It makes the benchmark framework more modular, and we could even add more archetypes later without rewriting everything.

One thing I'd add: the test client itself shouldn't be monolithic. It should be a collection of small, composable modules - one for each archetype. That way, someone could configure a workload to simulate, say, 30% sticky web sessions, 10% persistent pub/sub, and the rest short-lived tokens. It'd give us more realistic and flexible mixes.

The chaos schedule idea is gold. We could implement it as a simple probability model that occasionally triggers, say, a cache flush during peak analytics. It would expose those hidden bottlenecks.

You're spot on about long-lived connections. I've seen systems pass all the uptime checks but still lose session data on failover because they weren't testing the dormant states.


Clean code, happy life


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your opening premise is precisely why I find most published benchmarks so frustrating to use in practice. The gap between vendor marketing and real world operational reality is more than just methodological. It's often an intentional choice to obscure the complex trade offs a practitioner must make.

Focusing on operational considerations like failover and point in time recovery is a critical, often overlooked dimension. Many benchmarks treat the database as a black box under perfect conditions. But the true cost of a managed service often reveals itself during these operational moments. For instance, a vendor might tout a five minute Recovery Time Objective, but the fine print reveals that's only for a single node failure, and the application layer reconnection logic isn't part of that metric. A standardized benchmark that measures from the first client error to full application stability, as others have mentioned, would expose that gap immediately.

This working group must also grapple with the reproducibility challenge for these operational tests. Simulating a network partition or a storage failure in a consistent way across different cloud providers and on-premises setups will be a significant technical hurdle. Without that, we'll just end up with another set of incomparable numbers.


Let's keep it constructive


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Exactly. That fine print is where the real costs hide. A five minute RTO is meaningless if it takes fifteen minutes for your application pool to drain and reconnect gracefully. You end up paying for a standby instance you can't even use during the transition.

Your point about reproducibility is the real monster. Good luck simulating a true storage failure or network partition in a consistent way across AWS, Azure, GCP, and on-prem. The vendors' own failure injection tools are black boxes with unknown fidelity. Without a way to standardize the *failure condition itself*, we're just benchmarking their simulation, not the system.



   
ReplyQuote
Page 2 / 3