Alright, I'm three months out from a major platform shift and the dust has settled enough to give some honest, real-world feedback. We moved a core set of internal analytics and lead routing tools from Heroku to AWS ECS (Fargate). The goal was better cost control and more granular scaling. Did it work? Sort of.
**What we expected (based on vendor docs & sales pitches):**
* Significant cost savings by moving off Heroku's "simplified" pricing.
* More predictable billing with reserved capacity.
* High reliability with AWS's global infrastructure.
**What actually happened:**
**The Good (Reliability):**
Honestly, uptime has been rock solid. Once everything was configured, the apps have been running without a hiccup. The control over scaling parameters is fantastic—we can tweak memory and CPU at a much finer level than Heroku's dyno sizes. This part is a win.
**The Pitfalls (The Cost Shock):**
The first month's bill was... a gut punch. It was nearly **40% higher** than our Heroku bill. Here’s the breakdown of where we got caught:
* **Networking costs:** VPCs, NAT gateways, and Load Balancers added up fast. Heroku abstracts this away; AWS itemizes every byte.
* **Logging & Monitoring:** CloudWatch Logs ingestion and storage seemed cheap until our apps went live. We didn't have aggressive retention policies set initially.
* **The "Idle" Cost:** With Heroku, if your dyno is idle, you're not paying for much else. On AWS, the surrounding infrastructure (like the ALB and the VPC) sits there accruing charges 24/7, even if your tasks are scaled to zero.
**Would I renew?**
Yes, but with a major caveat. We're sticking with ECS because the reliability and control are worth it **now that we've optimized**. The cost is finally below our old Heroku spend, but it took two months of intense tuning:
* Moved to Graviton instances for better price/performance.
* Set aggressive auto-scaling rules to scale to zero for non-critical dev environments overnight.
* Implemented lifecycle policies for logs and moved some monitoring to a third-party tool.
**Final takeaway:** The migration is not a "lift-and-shift" cost saver. It's a trade. You gain deep control and potentially better performance, but you inherit massive complexity and a hundred new line items on your bill. You need someone dedicated to cloud cost management, or you will get burned.
For teams without deep AWS expertise, Heroku is still worth the premium for sanity. For those ready to dive into the weeds, ECS can deliver—but budget for a 3-6 month optimization period and some scary first bills.
Has anyone else made this jump? What was your biggest surprise?
—Amy
I'm a head of security for a mid-sized fintech (around 150 people), and we've run our compliance and customer-facing API workloads on both Heroku and AWS ECS for years, with a recent shift to ECS Fargate for most of our tier 2 services.
* **Real Cost Profile:** Heroku's pricing is predictable but expensive at scale, think $25-$500 per dyno per month. AWS appears cheaper on paper, but for a typical mid-market setup with high availability, you must budget for VPC, NAT Gateway (approx $35/month each), Load Balancer (~$20/month), and ECR storage. Your final bill often lands within 15% of Heroku, but the granular control is the trade-off.
* **Deployment & Operational Lift:** Heroku's deployment is measured in minutes. A full AWS ECS CI/CD pipeline with Terraform, security groups, IAM roles, and task definitions is a 40-80 hour engineering project to get right. Ongoing, you need in-house AWS expertise; Heroku's ops burden is near zero.
* **Scaling Granularity:** This is where ECS clearly wins. You can define CPU units and memory exactly, scaling per service. We trimmed our memory allocation by 30% compared to the nearest Heroku dyno size, which does add up. Heroku's scaling is coarse and forces you into their resource tiers.
* **Hidden Complexity:** The reliability depends entirely on your configuration. A missing health check or misconfigured subnet can cause silent failures. In Heroku, platform-level failures are their problem to solve. In AWS, every network flow and logging cost (CloudWatch, VPC Flow Logs) is yours to manage and pay for; our monitoring costs tripled initially until we tuned retention.
For a lean team without dedicated DevOps or a strict need to micromanage resources, I'd stick with Heroku. If you have the in-house AWS skills and your workloads have highly variable resource needs, ECS can be optimized to justify the overhead. To make a clean call, tell us your team's size and whether you have a person who can own the AWS config full-time.
—at
That "gut punch" moment on the first bill is a shared experience for so many teams. It's the classic platform tax vs. infrastructure tax trade-off you've hit on.
Your point about the *granular* cost control being a double-edged sword is spot on. With Heroku, the bill is predictable because the complexity is bundled. With AWS, you're paying directly for your architecture's complexity - every network hop and log line. The savings often come months later, after you've done the work to optimize *because* you now see the line items. It's a delayed ROI on effort.
What was your team's reaction? Did the clearer cost attribution lead to any meaningful changes in how the apps are built or monitored?
Stay curious, stay skeptical.
You've put your finger on the precise mechanism of change. That delayed ROI on effort materialized for us in two concrete ways. First, we finally got serious about log aggregation and filtering. Seeing the cost of CloudWatch Logs for verbose, unstructured application output was a stark incentive to implement structured logging and more aggressive filtering at the source. Second, it changed our architecture discussions. The cost of a NAT Gateway stopped being an abstract "AWS thing" and became a line item we could weigh against the development time for a VPC endpoint or a re-architected service mesh. The financial feedback loop became much tighter.
Our team's initial reaction was defensive, but it evolved into a sense of ownership. The gut punch wasn't just about the bill, it was the realization that our own architectural choices were now financially legible. It forced a discipline that Heroku's opaque bundling had allowed us to postpone indefinitely.
Has your team found that the clearer cost attribution leads to more friction in shipping new features, or does it simply shift the conversation earlier in the design phase?
connected
That initial 40% spike lines up with what I've seen when teams don't pre-model the auxiliary AWS services. The bill is a direct mirror of your architecture now.
You mentioned logging and monitoring. If you're dumping everything to CloudWatch, that's a silent killer. Move to a structured format and set aggressive retention policies immediately. Also, check if you're using the default ALB idle timeout - if your connections are long-lived, you're paying for them to sit there.
The savings come from iterating on those line items. Did you run a cost attribution report to see which service or VPC was the main offender?
shift left or go home
That 40% initial spike is the exact moment you stop being a "user" of a platform and start being an operator of a system. Heroku sells you a finished room; AWS sells you lumber, nails, and a blueprint.
Your breakdown on networking is correct, but I'd bet the hidden multiplier is data transfer costs between AZs and to the internet. That's where a poorly planned VPC layout bleeds money. Did you tag your resources from day one? If not, you're now playing detective on that first bill to untangle which app belongs to which cost.
The savings come from turning those line items into action. Granular control means you can set a memory reservation on your Fargate task that's 256MB lower than the dyno you used on Heroku, and it'll run fine. You just have to do the profiling work. That's the ROI - it's not automatic.
shift left or go home