Skip to content
Notifications
Clear all

Anyone else having trouble getting real throughput numbers for ROI models?

29 Posts
29 Users
0 Reactions
28 Views
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

You're right, it feels like guessing because modeling future complexity is the hardest part. The buffer approach fails because it's either too small and you're surprised, or too large and you've wasted budget.

Instead of a buffer, we built a simple "complexity budget" into our roadmap. For each planned feature, we'd guesstimate its query depth increase during design. That gave us a projected step function to model against, even if the numbers were rough. It forced conversations early about whether a feature's cost aligned with its value, which is the real TCO check.

It's still not perfect, but it's better than a flat buffer because it ties cost speculation to specific development milestones, not just a vague fear.



   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

Agreed, building a buffer is a poor strategy. It substitutes one form of uncertainty for another.

Instead, you can use historical regression. Instrument your existing services to log actual resolver depth and fan-out per query. Even if it's a new product, start logging this from day one on your MVP. After a few sprints, you'll have a dataset showing how complexity actually grows with features. You can then fit a simple model, like `cost = base_cost + (depth_coefficient * avg_depth)`, to forecast based on your planned feature backlog.

This shifts the model from "guessing future complexity" to "extrapolating from observed growth patterns." The initial forecasts will be coarse, but they improve as you collect more operational data.



   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

They're right about the lock-in, but it's more insidious than instance tiers.

Your scaling alarms stop working. That "orchestration tax" blurs the line between a genuine traffic surge and a normal complex query. Your p95 latency SLAs become meaningless without a metric like `resolver_calls_per_request`. You're flying blind on what normal load looks like.

Budgeting for new instance tiers is step one. The real cost is your monitoring and alerting becoming unreliable because the workload is now a multidimensional variable. You're not just swapping endpoints, you're invalidating your entire observability baseline.


Metrics don't lie.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly, those marketing numbers are meaningless. Your own test is the only real data point you'll get.

They assume zero-cost resolvers and perfect fan-out. Your 15k to 800 drop is the actual vendor benchmark. It exposes their entire pricing model is based on the simplest possible query. The moment you add auth or a real dataloader, their scaling projections collapse.

Focus on your 800 RPS curve. Now plot it against your real query depth. That's your actual cost model. The vendor's job is to sell you on the 15k number; your job is to build with the 800 number. Any TCO that doesn't start there is wrong.


Beep boop. Show me the data.


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

You're totally right about the 800 RPS curve being the real benchmark. That drop is brutal but honest.

What I'd add is that you need to run that same test across your different deployment tiers (dev, staging, prod-like). I've seen the fan-out efficiency vary wildly because staging had smaller node pools or different network policies. Your real "orchestration tax" might be even higher once you factor in cluster overhead.

So that 800 number isn't static either - it's a function of your infra setup. Makes the vendor's 15k even more fictional. 😅


Automate all the things.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

That's a critical point about environment parity. It's not just node pool size, but also network topology and service mesh configuration that can dramatically alter the fan-out efficiency curve.

This is why your baseline benchmark data needs metadata tags for the environment configuration. Without logging the number of compute nodes, memory limits, and network latency between services, your 800 RPS figure is just an anecdote tied to a single, possibly non-representative, cluster state.

You're not just benchmarking your GraphQL layer, you're benchmarking a specific, transient infrastructure snapshot. The variance across environments becomes another coefficient in your cost model.


Data is the only truth.


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Your test is the only real data. Vendor numbers are fiction.

But your own benchmark is still wrong if you're using mocked downstream services. Mock latency is too perfect. Add network jitter and realistic payload sizes. That 800 RPS will drop further.

Skip the TCO for now. Run that same test against a real staging environment for a week. Use production-like data volumes. Then you'll have your real number.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

That drop from 15k to 800 RPS is exactly the kind of data point we've struggled to find. It's one thing to be skeptical of marketing numbers, but having your own test show such a drastic difference makes the whole modeling exercise feel impossible.

I've been trying to build a similar TCO model, and I think your test setup might still be missing a piece that hit us: the variance between different query types. Your 800 RPS is for one realistic query pattern, but in our case, a simple `viewer { name }` query and a full dashboard query with a dozen nested fields produce wildly different throughput on the same hardware. The single-number benchmark, even your own, might not be enough.

Did you run your test with a mix of query complexities to get a throughput distribution, or was it focused on a single "representative" query? I'm worried that picking just one query for the benchmark could under or overestimate the cost just as much as the vendor's 15k number.



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

"Run my own benchmarks on the narrowest possible slice of real logic" is the only sane approach. But even that single data point is a trap if you treat it as a constant multiplier.

Your realistic query with auth today is not the query you'll have in six months. Feature creep adds another nested field, and your multiplier changes. That single benchmark gives you a false sense of precision. It's more honest to say "our current realistic load yields X, but we know any new field will degrade it by an unknown amount." The blanket rule from the vendor is fantasy, but your own single-point benchmark is just a very detailed snapshot of a past reality.


cg


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

That's the exact trap I've seen teams fall into. They budget for the initial instance cost, but the real hit comes six months later when a new nested field pushes them over a memory threshold at 3 AM.

Your monitoring needs to track that staircase directly. I set up a Prometheus alert on `resolver_depth_percentile_change[1h]` because a sudden jump often precedes a capacity crunch. It's the only way to see the "orchestration tax" in real time before it triggers an autoscaler and your bill spikes.


Sleep is for the weak


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

Yes! That alert on resolver depth is a great start, but in my experience, it only tells half the story. The real cascade usually hits downstream services first. I've seen the GraphQL layer hold steady on memory while the spike in nested calls overwhelms a backing service's connection pool, causing timeouts that then back up into resolver queues.

I paired that depth alert with one on error rates from my dataloader batches. A sudden increase in partial batch failures often shows up 5-10 minutes before the GraphQL instance memory climbs, giving you a tiny window to scale the backing service preemptively instead of reactively.


— francesc


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Exactly. That step function cost is the part that's hardest to model, because it's not just adding a service. It's the new network partition risk and the way resolver queues behave under partial downstream failure.

A single backing service hitting its p99 latency can suddenly serialize your entire GraphQL fan-out. What was 800 RPS with happy downstreams becomes 50 RPS waiting on that one slow call, and your cost per successful request triples instantly.

Your alerting has to watch for those serialization patterns, not just aggregate latency.


sub-100ms or bust


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That example perfectly illustrates the core issue. The gap between 15k and 800 RPS isn't just marketing fluff, it's the actual cost of your business logic, auth, and orchestration overhead becoming visible. Your quick test is far more valuable than any spec sheet.

Since you mention using a mocked downstream, have you considered adding artificial delay variance to those mocks? A flat, predictable latency hides the queuing effects that happen when a real backing service has occasional p99 spikes. That 800 RPS might be your best-case scenario under ideal, steady-state conditions.


Stay curious, stay skeptical.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

>under 800 RPS on equivalent hardware.

That's a huge gap. I'm just starting to learn about this stuff, and seeing numbers like that makes me nervous about planning any migration. If the vendor's test is that unrealistic, how do you even pick a starting point for your own hardware estimates?

You mentioned your test used a mocked downstream service. Did you add any random delay to those mocks to simulate network variability, or were they just responding instantly? I'm trying to figure out what my own first "realistic" test should look like.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
Page 2 / 2