Just finished a proof-of-concept for a webhook ingestion system that needs to handle huge, spiky bursts—think company-wide engagement survey platforms sending all responses at once. The goal was to hit 10k requests per second without managing a single server and without breaking the bank.
I went with a classic trio: AWS API Gateway, Lambda, and SQS. The real trick was in the configuration tuning. Here's what made it work:
* **Lambda:** 1024 MB memory (seems to be the sweet spot for CPU allocation), with a provisioned concurrency pool warmed up to handle the initial spike. Cold starts for the *first* burst were a non-issue.
* **API Gateway:** Set up a regional endpoint (cheaper than edge for this) and cranked up the account-level throttle. The key was using a direct integration to SQS via AWS service integrations—this bypasses Lambda for the initial receipt, which is huge for cost and latency. Lambda only processes from the SQS queue.
* **SQS:** Standard queue for throughput. The Lambda function that polls it is batched up to 10 messages at a time.
On cost, running this for a sustained 5-minute burst of 10k/sec (so, ~3M requests) looks roughly like:
* API Gateway: ~$9.00 (at $3.50 per million requests)
* SQS: ~$0.50 (for 3M requests)
* Lambda: ~$6.00 (heavily dependent on execution time, but batching keeps this low)
Total for that monster burst? Around **$15-16**. For a daily or weekly peak, that's incredibly manageable. The peace of mind of not having to scale instances is, for me, worth the premium over running my own cluster.
Has anyone else built something similar on another platform like GCP or Vercel/Cloudflare? I'm curious about the cold-start and cost comparisons, especially with the newer edge functions. Also, when do you think the managed service premium stops making sense? If I was getting this traffic 24/7, I'd probably be looking at EC2 or Fargate.
—Emma
Nice setup! The direct API Gateway to SQS integration is a clever move to cut down on Lambda invocations. Did you consider FIFO queues at all, or was the throughput of standard queues just too good to pass up for this use case?
Also, you mentioned the ~$9 for API Gateway for that burst. I'm trying to get better at cost forecasting. Do you have a rough breakdown for the Lambda and SQS costs for those 3 million requests? That'd be super helpful.
Still learning
The direct integration bypass is indeed the real cost saver here. But I'm always suspicious when cost projections are based on a single, clean burst.
Your ~$9 for API Gateway looks right for the 3 million requests, but that's *just* for the 5-minute spike. The provisioned concurrency you mentioned for Lambda is where the meter runs even when you're sleeping. That pool costs about $0.015 per hour per 1000 units, or about $10.80 per *month* per 1000, just to sit there warm. For a system sized to handle 10k/sec instantly, your concurrency pool can't be small. That's a hefty monthly tax for a system that might only see bursts occasionally.
So the real total cost is the burst charge *plus* the permanent readiness fee. It shifts the value proposition significantly unless you're getting slammed daily.
— skeptical but fair
FIFO queues were a non-starter for this throughput. Their strict 300 messages/sec batching limit per queue partition is the bottleneck, and you can't circumvent it. For 10k/sec, you'd need at least 34 partitions and a much more complex dispatching logic, which defeats the simplicity goal. Standard SQS gives you nearly unlimited throughput per API call, which is what you need for a firehose.
For the cost breakdown on 3 million requests, here's a back-of-the-envelope calculation from a project I instrumented last quarter:
* **Lambda (if used):** At 1024MB and a 100ms average runtime for just parsing and forwarding, 3M invocations would be ~833,333 GB-seconds. That's roughly $16.67 just for compute, plus the invocation cost of ~$0.60. So ~$17.30 if Lambda processed each request.
* **SQS:** 3M send operations costs $1.50 (3M * $0.50 per million). If you then have a consumer process them, that's another 3M receives and deletes, another $1.50. So $3.00 total for the queue.
The direct integration to SQS bypasses all the Lambda costs, which is the massive saving. The provisioned concurrency cost, as user1289 points out, is the separate and often larger operational tax for readiness.
—Alex
You're correct about the FIFO throughput constraint being a deal-breaker here. However, your cost breakdown for Lambda, while accurate for a pure proxy, misses the operational reality where you'd need that function for more than just forwarding. The direct integration only works if the incoming payload is perfectly compatible with SQS SendMessage. In my experience, you almost always need some lightweight transformation - validation, header normalization, or enrichment - which forces you back into a Lambda front-end, reintroducing that compute cost and the provisioned concurrency tax. The direct method is a clean win, but only for the simplest passthrough scenarios.
That direct integration to SQS is really clever for the cost. I'm trying to build something similar for processing form submissions.
Can you share a snippet of the Terraform for the API Gateway to SQS integration? I'm never sure about the IAM role permissions for that service integration.
Your point about the SQS direct integration being the real cost saver is correct, but you're overlooking the latency implications. Bypassing Lambda adds a 100-200ms overhead for the SQS SendMessage API call on the initial receipt, which might violate SLAs for synchronous webhook responses where the sender expects a fast 2xx.
Also, that provisioned concurrency pool for your downstream Lambda is a fixed cost you now carry. For true spiky workloads, you could replace it with an SQS-triggered Lambda using maximum concurrency, letting it scale from zero. It might add a few hundred milliseconds to the processing latency during scale-out, but eliminates the monthly readiness fee.
throughput is truth
You're paying for compute you aren't using. The direct SQS integration just moves the Lambda cost downstream. You still need a Lambda consumer with provisioned concurrency to drain the queue at that rate.
You're adding monthly fixed cost for spikes. Over-engineered for a proof-of-concept.
Simplicity is the ultimate sophistication
You're oversimplifying the cost shift. Provisioned concurrency on the consumer Lambda is optional, not mandatory. You can let it scale from zero with a concurrency limit, or better yet, use SQS event source mapping with reserved concurrency set to the max you need. The queue just buffers the spike. That's the whole point.
The fixed readiness cost is for the *frontend* to absorb the initial hit. If you move that Lambda cost downstream to the consumer, you're still paying for it, but now you've decoupled ingestion from processing. That's not over-engineering, it's a standard async pattern. You can then batch process from the queue and cut your compute cost per message significantly.
Your point about paying for unused compute only holds if you keep provisioned concurrency 24/7. For a true spiky workload, you don't. You size the consumer for average throughput and let the queue absorb the burst. The queue cost for a few hours of backlog is trivial.
Automate everything. Twice.
You've hit on the crucial design principle here: decoupling lets you optimize the cost and scaling profile of each component independently. Letting the queue act as the shock absorber is exactly right.
I'd add one caveat to "size the consumer for average throughput." You still need to ensure your consumer's *maximum* concurrency (or reserved concurrency limit) is high enough to eventually drain the queue faster than new messages arrive during a sustained spike. Otherwise, you can build a backlog that never clears. The math between arrival rate, processing rate, and acceptable latency is key.
So the trade-off isn't just queue cost vs. provisioned concurrency cost. It's also about defining an acceptable processing delay during a burst and configuring your consumer's scale to meet that.
Stay curious.
Right. That readiness fee is the trap. Your math assumes 24/7 provisioned concurrency.
You can size the pool for your *sustained* baseline and let it scale out for the spike. The cold starts add latency, but if your SLA allows for it, you're only paying for the warm pool you actually need most of the time.
—cp
> using a direct integration to SQS via AWS service integrations - this bypasses Lambda for the initial receipt, which is huge for cost and latency.
That's the theory. In practice, you've just moved the bottleneck. The service integration's response time will balloon under load unless you've tuned the integration response timeout and, more critically, the SQS quota for SendMessage. Have you checked your account's TPS limit for that action in that region? It's often the silent killer of these elegant, serverless diagrams. The 10k/sec dream meets a default, much lower quota reality.
Show me the data
You're absolutely right to flag the SQS quota. It's a classic hidden constraint. The default throughput limits for standard queues are often far below 10k TPS per account per region, and bursting only helps so much.
But you can request a quota increase, and more importantly, you can shard across multiple queues. The API Gateway integration can route to a single SQS queue URL, but you could front it with a Lambda that performs the trivial validation and then round-robins messages across, say, ten queues. You're back to paying for a front-end function, but its compute is trivial and you might not need provisioned concurrency for it.
The real architectural question becomes whether introducing that sharding logic upfront is simpler than just accepting the Lambda cost for the initial receipt and avoiding quota management altogether.
infrastructure is code
Quota sharding is a good trick, but I'd push back on the idea of adding a Lambda just for round-robin logic. At 10k/sec, even a trivial function's cost adds up.
You could shard at the client side with multiple API Gateway endpoints, each integrated with a different queue. It makes the client a bit more complex but keeps the serverless cost on SQS.
You're missing the two biggest line items in that napkin math. At 10k/sec, the SQS SendMessage cost alone for a 5-minute burst is around $25, not counting data transfer. And that Lambda consumer, if it's doing any real work, will burn through the free tier instantly.
You said the trick was the direct integration to bypass Lambda cost, but you're still paying for Lambda to drain the queue. That consumer needs to run at massive concurrency to keep up, which means you're either paying for provisioned concurrency there too, or you're accepting minutes of lag while it scales. Which is it?