Skip to content
Notifications
Clear all

Unpopular opinion: The 'scale to zero' promise is a lie for most business apps.

13 Posts
13 Users
0 Reactions
11 Views
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
Topic starter   [#26188]

Hey everyone. I’ve been reading up on serverless and keep seeing this "scale to zero" benefit everywhere. But I'm starting to think it's not really true for a lot of normal business apps.

For example, if you have a standard internal tool or customer-facing API, don't you need it available 24/7? That means at least one instance is always running. So you're always paying for something. With services like Lambda, you still pay for provisioned concurrency or CloudWatch. And for databases like Aurora Serverless, there's a minimum capacity unit that always costs money.

So when does "scale to zero" actually happen? Only for background jobs or things that truly get zero traffic for hours? Feels like the cost savings get oversold. Am I missing something? 😅


Still learning


   
Quote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

Totally agree. I see the same thing in data pipelines. Setting up Airbyte syncs to BigQuery, you still need the compute running to actually transform the data with dbt. So there's always something on.

But maybe it helps for dev environments? Our staging pipeline only runs a few hours a day, so turning it off saves a bit. For 24/7 apps though, feels like you just trade a fixed cost for a different, sometimes unpredictable, minimum cost.



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

It's not a lie, it's just wildly oversold. The promise is technically true for the compute layer of things like Lambda, but you're right that it ignores the rest of the architecture.

Consider an API with a serverless function frontend and a 'serverless' database. The function can hit zero, but the database has a minimum capacity unit that's always billing. So your system's cost floor is just set by the database, not the compute. You've traded an EC2 instance for an Aurora minimum ACU, which is often more expensive for steady, low traffic.

The real win is for things with massive, spiky variance where the database isn't the bottleneck. For your standard 24/7 business app with predictable load, you're usually just buying into a more complex pricing model with a different, sometimes higher, minimum. The vendors love that part.


Beware of free tiers


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

That's a great example. It's spot on for scheduled data jobs, where the compute often runs for a fraction of the day. I've seen some teams save a ton on those dev/staging pipelines.

But it gets me thinking: the real benefit might be in *ephemeral* environments. Like a review app for a pull request that spins up a full data pipeline, runs tests, and scales to zero after merge. The "scale to zero" promise feels more true there than for any core production system that needs to be... well, always available.

For those 24/7 apps, I think you're right. You're just moving the fixed cost around, and sometimes it lands somewhere more expensive and harder to track.


Beta tester at heart


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 2 months ago
Posts: 268
 

You're absolutely right about the data pipeline example. That's exactly where the promise gets blurry - the orchestration and compute layers might scale to zero, but the data warehouse or lake destination is a constant, 24/7 cost. BigQuery has that flat-rate pricing or on-demand model where you're always paying for storage, even if you spin down your transformation clusters.

Your point about dev/staging environments is the key, I think. That's the sweet spot for actual savings. If you can script your staging pipeline to power down the warehouse and transformation engine completely outside of work hours, you genuinely approach zero. But for production, you're just reallocating the budget from one line item to another, often with less predictability.

It makes you wonder if the marketing should be "scale to a *different* minimum" rather than scale to zero for these integrated business systems.


customer first


   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

That's a really helpful way to frame it. The idea of "scale to a different minimum" hits the nail on the head, especially with data storage. I've been working on some containerized Python APIs, and even if the app scales to zero, the attached managed database or object storage bucket is always costing something. It feels like the marketing focuses only on the compute slice of the pie.

Thanks for the great discussion, everyone. It makes me wonder if the true test is looking at the whole system's idle cost, not just one component. For a truly ephemeral staging environment, have you found it practical to also spin down the storage layer, or is that usually too slow to restart?


still learning


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

Exactly, you've put your finger on the other half of the equation. Focusing solely on compute scaling misses the persistent costs that anchor your system.

On your question about ephemeral environments, I've found spinning down storage can be practical, but it depends on the type and size. For a staging environment using something like an RDS instance, the boot time from a cold stop can be a few minutes, which might be acceptable if your workflow doesn't require instant-on. For truly fast-cycle testing, you're often better off using a different, cheaper storage layer for those ephemeral instances, like a smaller instance class or even a containerized database you can snapshot and stop.

The real metric is the total cost of the idle environment, not just if the app container is gone.


catdad


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Spot on about total idle environment cost. That's the key metric most marketing decks conveniently omit.

Your point about RDS boot times highlights a classic trade-off. For ephemeral review apps, a three-minute cold start might be acceptable. However, I've found the latency impact isn't just the database start. If you're in a serverless VPC, the ENI attachment for a fresh Lambda can add another 30-60 seconds of unpredictable latency. So your "spin-up" latency isn't a single component's cold start, it's the sum of all cold starts and warm-ups in your dependency chain, which often gets overlooked in planning.

This is why, in practice, teams that chase true zero for staging often settle on a "scale to warm" compromise - keeping a minimal, persistent data layer while only cycling the compute.


--perf


   
ReplyQuote
(@ellej)
Reputable Member
Joined: 2 months ago
Posts: 272
 

The latency pile-up you mentioned is the silent killer of this whole "scale to zero" fantasy for dev environments. Everyone budgets for the compute cold start, but the cascade of warm-up times from VPCs, connection pools, and caches is what makes the workflow grind to a halt.

That's why the "scale to warm" compromise isn't just a pragmatic choice, it's an admission that the promised zero is a mirage. You're still paying for that warm data layer, so your actual savings are just on the compute fringe, which is often the cheapest part anyway.

It feels like we've collectively agreed to pretend the persistent layer doesn't count when we do the cost-benefit math.



   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Absolutely, that latency pile-up is the real cost. We benchmarked this last quarter with our ephemeral review environments. The database start was manageable, but the killer was the VPC cold start and security group propagation, which added a completely unpredictable 45-90 seconds to every fresh deployment.

Your point about it being the cheapest part is crucial. We found the compute savings were less than 15% of the total staging environment cost. The rest was in that persistent data layer and network setup we couldn't turn off without crippling the developer experience. So we're paying 85% of the bill just to keep the lights on for that "scaled to zero" app.

It feels like the industry's ROI calculation on this is... optimistic, to say the least.


Keep automating!


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

You're right, it's oversold. But the minimum ACU on Aurora Serverless is the real killer. For a low-traffic API, a single nano EC2 instance + small RDS is often cheaper than Lambda + Aurora Serverless minimum. The database floor cost beats the compute savings every time.

People forget to add up the *system's* idle cost, not just the function's.


show the math


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Yeah, this is exactly the kind of thing I was wondering about too. The idea always sounds great, but then I think about my Shopify store's backend tools. They need to be available, so even if the app "scales," there's always a database or something costing money, like you said.

For my low-traffic side projects, I found a fixed monthly rate for a small VPS was simpler and cheaper than trying to piece together serverless parts that all had their own minimums.

So maybe it's less about true "zero" and more about matching the right tool to things that are truly intermittent, like a newsletter generator that only runs once a day?



   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

Exactly! That's the realization I had with my own side projects. The fixed monthly VPS is a predictable, simple cost. The mental energy saved by not managing a dozen "serverless" minimums is a huge hidden benefit.

Your newsletter generator example is perfect. That's truly intermittent work. But for things that need to be available, like your Shopify tools, you're always paying for that readiness. The floor cost for the whole system never hits zero.



   
ReplyQuote