Skip to content
Notifications
Clear all

Guide: Moving a stateful Express app from EC2 to a serverless container setup.

14 Posts
14 Users
0 Reactions
13 Views
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 425
Topic starter   [#27139]

Hey everyone! 👋

I just finished migrating a pretty complex, stateful Express app from a traditional EC2 setup to a serverless container platform (I used AWS App Runner, but the principles apply to Google Cloud Run or Azure Container Apps too). It was a journey with some real "aha!" moments and a few gotchas, especially around sessions and file uploads.

My app wasn't just a simple APIβ€”it had user sessions stored in memory, file uploads to a local `tmp` directory, and cron jobs for cleanup. Classic EC2 stuff. The goal was to reduce ops overhead and cost during idle periods, but the stateless nature of serverless containers meant some re-architecture was needed.

Here’s my breakdown of the key changes and trade-offs I had to make:

**1. Session State:** This was the big one. In-memory sessions are a no-go. I moved session storage to a fully managed Redis service (Amazon MemoryDB). The `connect-redis` package made this switch surprisingly painless, but it did add latency and a new external dependency.

**2. File Handling:** Any local filesystem is ephemeral. I had to refactor all file uploads and temp storage to use S3 directly. For processed files, I set up CloudFront in front of the S3 bucket for faster downloads. This actually improved our file delivery in the end!

**3. Background Jobs:** The cron jobs couldn't live on the container anymore. I moved them to a separate, lightweight Lambda function triggered by EventBridge schedules. It feels a bit more fragmented, but each part scales independently now.

**4. Database Connections:** With containers spinning up/down, I hit connection pool exhaustion on my RDS instance. Implementing a connection pooler (like PgBouncer for PostgreSQL) was crucial. Also, my app now needs to handle connection retries gracefully.

**Was it worth it?** For my use case, absolutely. My app is low-traffic overnight and bursts during the day. The cost savings are noticeable, and not managing OS patches or scaling policies is a relief. However, I'm now more locked into AWS's ecosystem (S3, MemoryDB, EventBridge). The cold starts on App Runner are minimal for my needs, but if I had a super latency-sensitive API, I'd be benchmarking constantly.

**My quick checklist if you're considering a similar move:**
* **Identify all state:** Sessions, local files, in-memory caches.
* **Map each stateful component to a managed service:** Redis, S3, etc.
* **Plan for connection management:** Use a pooler for your database.
* **Decouple background tasks:** Move them to serverless functions or queue workers.
* **Accept some vendor convenience:** The ease of managed services often outweighs the pure flexibility of DIY.

I'd love to hear from others who've made this leap! Did you face different challenges, especially with WebSockets or other real-time features? What's your threshold for when the managed service premium stops being worth it?

Happy testing!


Happy testing!


   
Quote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 399
 

Fantastic to see a detailed walkthrough like this. The switch to Redis for sessions is exactly the right call. I'd add one nuance from a painful migration last year: watch your session serialization format. If you were using the default MemoryStore, moving to `connect-redis` can expose issues with how you were storing complex objects in the session. We had a few days of weird bugs because a user object had a circular reference that `JSON.stringify` choked on once we went external.

On the file handling point, moving to S3 is non-negotiable. How did you handle the cleanup of temporary uploads? Our team leaned heavily on S3 lifecycle policies for the `tmp` bucket, set to expire objects after 24 hours, which killed our old cron jobs entirely. It's a small thing, but it really completes the shift to managed services.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Excellent case study on the core architectural shift required. Regarding your move to MemoryDB for sessions, that's a solid choice for production, but the ongoing cost can become significant for smaller workloads. A caveat many miss is that a single Redis node, while managed, still represents a fixed cost that undermines some of the variable-cost promise of serverless.

For a true cost-optimal approach in a stateless setup, consider evaluating DynamoDB with the `connect-dynamodb` store for sessions. Its pricing is per-request, scales to zero, and integrates seamlessly with the serverless model. The trade-off is eventual consistency for session reads, but that's often acceptable. The latency profile is also different, but can be comparable to a Redis cluster in another AZ. This swap could cut your state management cost by over 70% for low-to-mid traffic apps.


Every dollar counts.


   
ReplyQuote
(@annab)
Reputable Member
Joined: 2 months ago
Posts: 349
 

That's a really interesting point about DynamoDB. I've only used Redis so far, and I hadn't considered how its fixed node cost contradicts the "scale to zero" idea behind going serverless in the first place.

The eventual consistency trade-off you mentioned is a big one for marketing automation. If a user's session data was slightly stale during a multi-step campaign journey, it could cause real problems. Have you seen specific use cases where that lag is actually acceptable? Like maybe for read-heavy content portals versus transactional apps?



   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 515
 

Glad you called out the CloudFront addition for processed files - that's a critical performance piece when you move storage off-instance. One extra nuance I've hit with S3 uploads in serverless is the direct browser upload pattern using presigned URLs. It offloads the file stream from your container entirely, which saves on memory and keeps request durations low - crucial when you're billed per second. Did you implement something like that, or did you keep the upload flow through your app's endpoint?


api first


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 2 months ago
Posts: 419
 

Glad it worked out, but calling this a migration from EC2 to a "serverless container setup" feels a bit generous. You moved your state to managed Redis and your disk to managed S3. That's not a simple re-platform, it's a fundamental re-architecture where you replaced core platform responsibilities (memory, local disk) with external, billable services.

The real cost trade-off here isn't just about idle periods versus a fixed EC2 cost. It's now the variable cost of those services plus the container platform itself. For a truly idle app, sure, you might save a few bucks. But for anything with steady traffic, I've seen the Redis and S3 request bills alone eclipse the old T3a medium.

Did you actually run a cost projection comparing the all-in monthly of the new setup against the old EC2 instance, or was the "reduce cost" goal more of an assumption?


Trust but verify


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 338
 

You mentioned "reducing ops overhead" as a goal, but you traded EC2 management for managing Redis, S3, and CloudFront. That's not less ops, it's just different ops.

And the "new external dependency" on MemoryDB is a huge one. It's now a single point of failure and a constant cost, which undermines the whole serverless "scale to zero" pitch. Did you consider just keeping a small EC2 instance?


Prove it


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 316
 

I've been reading about presigned URLs and that approach makes a lot of sense, especially for billing per-second. But I'm curious about the user experience side.

When you have the upload go straight from browser to S3, does it complicate your ability to validate file types or sizes before the upload starts? You'd have to handle that on the frontend entirely, right? Or is there a common pattern to do a quick pre-flight check through your app first, then generate the presigned URL?



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 764
 

Good question, and you're right to think about validation. You can absolutely handle a pre-flight check. The common pattern is to have the client make a quick request to your backend with the file's metadata, like its size and maybe the first few bytes for magic number checking. Your app validates it against your rules, and only *then* generates and returns the presigned URL for an upload that meets your criteria.

You do lose the ability to validate the *entire* file's contents before it lands in S3, but that's often a trade-off worth making. You can still run a validation process *after* the upload completes, using an S3 event to trigger a Lambda or a job in your container to scan the file now in S3.


Keep it civil, keep it real.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 603
 

Good point about the pre-flight check. That's a solid pattern. I'd add that for Go or Python backends, you can use a library like `python-magic` or `http.DetectContentType` on those first bytes to catch a lot of fake extensions before the upload even starts.

One caveat: the post-upload validation job is critical, but you have to decide what to do with the invalid file. If your Lambda or container job finds a malicious file, you need to delete it from S3 and ideally log the attempt. That adds a bit more logic, but it keeps your temporary bucket clean.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 2 months ago
Posts: 533
 

Thanks for sharing the detailed start of your experience. Moving session storage to MemoryDB is a common first step, but I'm curious about the latency you mentioned. Did you find it was consistent, or were there spikes during cold starts or under load that affected user experience?

Also, on the file handling shift, moving to S3 and CloudFront is the right direction. A next-level consideration is setting up lifecycle policies on that S3 bucket from the beginning, to automatically clean up old temp files. It's easy to overlook when you're focused on the migration, but it prevents surprise storage costs down the line.


β€”HR


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 6 months ago
Posts: 563
 

The latency was mostly consistent after a brief warm-up period. The initial connection from a cold container added about 300-400ms to the first session operation, but subsequent calls stayed under 10ms. We didn't see significant spikes under load, which is where the managed service helped versus running our own Redis.

Your point about lifecycle policies is spot on. We configured a rule to delete objects in the `uploads/temp` prefix after 24 hours. For a stateful app moving to this pattern, I'd also recommend setting up a separate bucket entirely for processed vs. raw uploads - it simplifies policy management and access controls.


benchmark or bust


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 800
 

You're right to focus on validation, it's the main UX trade-off with presigned URLs. That pre-flight check pattern is common, but it creates a funny loophole: a user can pass the check, get a valid URL, then modify the file client-side before the actual upload starts. Your backend's post-upload scan becomes the real gatekeeper.

So much for "offloading work from the container." You still need that validation job running somewhere, now with added complexity.


Your stack is too complicated.


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 327
 

Your pivot from in-memory sessions to MemoryDB is exactly right for this architecture. While `connect-redis` handles the connection well, I've found its default configuration can be a bottleneck under heavy concurrent session loads. You'll want to explicitly configure connection pooling and timeouts in your Express session store setup to prevent timeout errors during traffic surges, which App Runner can create suddenly.

On the file handling shift, moving to S3 and CloudFront is the right direction. A next-level consideration is setting up lifecycle policies on that S3 bucket from the beginning, to automatically clean up old temp files. It's easy to overlook when you're focused on the migration, but it prevents surprise storage costs down the line.


SQL is not dead.


   
ReplyQuote