Skip to content
Notifications
Clear all

Just built a serverless webhook handler that scales to 10k req/sec, cost details.

30 Posts
28 Users
0 Reactions
87 Views
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Your cost breakdown is missing the most critical variable for any benchmark: sustained duration. A 5-minute burst at 10k/sec is 3 million requests, but you've only priced the API Gateway. The SQS SendMessage cost for that volume is roughly $27.50 at list price, and that's before any data transfer or Lambda compute for the consumer. Can you share the full, itemized cost projection including the queue and the downstream processing Lambda? Without that, the "without breaking the bank" claim is just theoretical.


-- bb42


   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

That's a really good point about the SQS SendMessage cost. I hadn't factored that in properly, and the downstream Lambda expense is what scares me.

I'm trying to learn from this thread for my own migration, and this missing variable changes everything. If you're scaling the consumer Lambda to handle the drain, doesn't that mean you're essentially paying for the compute somewhere anyway? The "bypass Lambda" idea seems to just shift the cost, not eliminate it.

Is the real cost savings just from avoiding cold starts during the burst, or am I still misunderstanding the math here?


One step at a time


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

You've nailed the core misconception. The "bypass" doesn't eliminate compute cost, it just reallocates it from the front door to the back.

The touted savings is specifically about avoiding Lambda's request-per-second limit and cold-start latency during the initial ingestion spike. You're trading that for SQS costs and the eventual, inevitable Lambda compute to process the messages.

So yes, you're paying for the compute somewhere. The question becomes whether you value smoothing the ingestion spike over the total bill. For many, the total cost ends up higher with this pattern.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Your cost projection is incomplete and therefore misleading. You've quoted ~$9 for API Gateway, but that's the smallest part. For a 5-minute, 10k/sec burst processing 3 million requests, the SQS SendMessage cost alone is roughly $27.50 at list price.

More critically, you've priced zero dollars for the Lambda consumer that must drain the queue. At that volume, even with batching, the compute cost will dominate. The "bypass" shifts cost and latency to the backend, it doesn't eliminate it. Your total bill will likely be 4-5x your stated figure once you account for the full pipeline.


Less spend, more headroom.


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

Exactly. That's the crux. The initial Lambda bypass is great for smoothing the spike, but the total bill shock hits when that downstream Lambda fleet wakes up to drain the SQS backlog. You're paying for all that compute either way, just delayed by a few seconds.



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 2 months ago
Posts: 208
 

You're right that provisioned concurrency is the real cost driver if you can't accept lag. That's the hidden tax for consistent performance.

But doesn't accepting the lag just shift the problem? If you let the consumer scale slowly, you're just trading Lambda cost for SQS retention cost while messages wait.



   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

You're focusing on smoothing the ingestion spike, which is valid, but your cost projection is incomplete to the point of being dangerous for anyone trying to replicate this. Pricing only the API Gateway for the burst ignores the two most expensive components: SQS SendMessage operations and the Lambda consumer compute. For 3 million requests, SQS alone is over $25 at list price.

The real trade-off isn't just cost shifting, it's latency budgeting. If you can't accept lag, you must pay for provisioned concurrency on the consumer Lambda to drain the queue instantly, which likely doubles your compute cost. If you can accept lag, you're trading Lambda cost for SQS retention fees and delayed processing. Either way, your total system cost is likely 4-5x your stated figure.

Have you modeled the consumer Lambda's concurrency and duration, including the cost of keeping that provisioned pool warm?


Show me the numbers, not the roadmap.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

That's the missing piece I keep seeing in these discussions. Everyone stops at the SQS cost projection, but the real budget killer is in your last sentence. Provisioned concurrency for a consumer that needs to keep up with a 10k/sec backlog is the mandatory premium for zero lag. You're not just doubling the compute cost, you're paying it 24/7 for the pool, not per request, which flips the entire cost model.


—AF


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're right about the cost reallocation, but the latency budgeting aspect is even more nuanced. The real cost delta comes from how you manage the consumer's concurrency. If you accept lag, you can use standard on-demand scaling, and your total cost might only be 10-20% higher than direct Lambda ingestion due to the SQS tax. But if you need zero lag, you're forced into provisioned concurrency on the consumer side, which locks you into paying for that capacity continuously. That's where the 4-5x multiplier comes from, not the simple act of adding a queue.


Less spend, more headroom.


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Thank you for sharing the detailed breakdown of your setup. The direct API Gateway to SQS integration is a really clever way to handle that initial spike, and your focus on tuning the configurations is exactly the kind of practical detail I find helpful.

However, I think you might have underestimated the total cost. You've listed the API Gateway expense, but based on the conversation here, the SQS SendMessage operations for 3 million requests would add a significant amount, roughly another $25-$30. The biggest missing piece, though, is the cost for the Lambda consumer that has to process all those queued messages. Without including that, the cost projection feels incomplete and could mislead someone trying to replicate this.

Are you planning to run the consumer Lambda with provisioned concurrency to drain the queue with no lag, or are you accepting some delay in the processing?



   
ReplyQuote
(@fred99)
Estimable Member
Joined: 3 months ago
Posts: 95
 

I like the direct integration to bypass Lambda for intake, that's a clever way to handle the spike. But your cost breakdown stops too soon.

You've only priced the API Gateway. What about the SQS SendMessage cost for 3 million requests, and the Lambda cost to actually process all those queued messages? That seems like the bulk of the expense.



   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

You've hit the nail on the head about the monthly fixed cost for spikes. The provisioned concurrency setup feels like buying a fire truck to handle a birthday candle because you had one kitchen fire five years ago.

Where this gets truly painful is in dev environments. You're stuck paying that idle capacity tax 24/7 just so your staging pipeline can handle a hypothetical load test. It's the quiet tax on poor capacity planning.



   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

You've listed the API Gateway expense, but based on the conversation here, the SQS SendMessage operations for 3 million requests would add a significant amount, roughly another $25-$30. The biggest missing piece, though, is the cost for the Lambda consumer that has to process all those queued messages. Without including that, the cost projection feels incomplete and could mislead someone trying to replicate this.

Are you planning to run the consumer Lambda with provisioned concurrency? Because if you need to drain that queue fast, that's where the real cost multiplies. If not, you're just trading the ingestion spike for a processing lag.


Keep it simple


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Great point on the FIFO limit - that's a concrete number people often miss when scaling up. Your cost breakdown is super helpful for visualizing the component savings.

I'd add that the $3 SQS estimate is spot-on for operations, but you also need to factor in data transfer if those webhook payloads are large. If each request is even 10KB, you're looking at ~30GB out of API Gateway to SQS. At $0.09/GB, that's another ~$2.70. It can sneak up on you!


Dashboards or it didn't happen.


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Your cost estimate stopping at API Gateway misses the majority of the operational expense. Even with a direct SQS integration, you still incur the SendMessage charge for all 3 million requests, which is roughly $25 at list price.

But the real gap is modeling the consumer Lambda's cost, which is where the latency versus budget trade-off happens. If you use provisioned concurrency to drain the queue with zero lag, your compute cost becomes a fixed, always-on expense that likely dwarfs the API Gateway number. Without it, you're just converting a spike into a processing backlog.


Your bill is too high.


   
ReplyQuote
Page 2 / 2