Skip to content
Notifications
Clear all

What's the best way to handle long-running jobs in a serverless world?

6 Posts
6 Users
0 Reactions
17 Views
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
Topic starter   [#19900]

The best way? Don't use serverless for long-running jobs. You're trying to hammer a nail with a feather. The whole point of these platforms is short, stateless bursts. They'll time you out, bleed your wallet dry, and hide the real infra cost.

If you're forced into this madness, here's the least-bad approach:
* Split the job into smaller steps and chain them with a proper queue (SQS, not some managed wonder-tool).
* Use a state machine (AWS Step Functions, etc.) to orchestrate. At least the logic is visible.
* Keep the actual compute in containers on Fargate or a real VM. Serverless should just trigger and monitor.

Example of a sane trigger from a Lambda (the only part that should be serverless):
```bash
#!/bin/bash
# Kicks off real job, doesn't try to be one.
JOB_ID=$(aws batch submit-job --job-name "real-work" --job-queue my-queue ...)
echo "Track ${JOB_ID}, not my problem."
```

Anything else is overcomplicating a solved problem. We had cron on servers for decades. It worked.

-- old school


-- old school


   
Quote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

I'm a senior platform engineer at a fintech company handling around 10k transactions a minute. Our data pipelines are all long-running batch jobs, and we've run them on Lambda, Fargate, Batch, and good old Kubernetes. Here's the real breakdown.

* **Real cost per job minute:** Lambda billed to the 100ms at ~$0.20 per 1M GB-seconds sounds great until your job runs 45 minutes. A single 2GB Lambda running for an hour is about $0.03. A 2vCPU/4GB Fargate spot container for an hour is ~$0.024, and you don't have to architect around a 15-minute timeout.
* **State and observability:** Serverless functions are black boxes. You get logs dumped to CloudWatch after the fact. With a container (Fargate/EKS/GCE), you can exec in, inspect the filesystem mid-job, and get real-time stdout. This matters when a 6-hour job fails at hour 5.
* **Cold start vs. constant readiness:** A Lambda with large dependencies (like a Python data library) can take 45-60 seconds just to start. For a 5-minute job, that's a 20% tax every cold invocation. A warmed container on Fargate or a VM starts processing in under 2 seconds, every time.
* **Infrastructure coupling:** Lambda forces you into the vendor's workflow, runtime, and permission model. A container image runs the same way on ECS, your local dev box, or a bare-metal cluster. We moved a job from Lambda to our on-prem K8s during an AWS outage with a config change.

I'd recommend AWS Batch with Fargate spot compute environments for any job over 5 minutes. It handles the queue, respects dependencies, and uses cheap, ephemeral containers. If you absolutely must stay in a "serverless" billing model, tell me your maximum acceptable job duration and your average job memory footprint. I'll tell you which headache to choose.



   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Spot on with the cost and timeout breakdown. The vendor lock point is critical.

Lambda forces you into their event shape and runtime API. We tried to move a batch job off AWS and the Lambda-specific logging and init code became a week of refactoring. With a container image, you just change the orchestrator.

The 45-60 second cold start for data libs is real. Our workaround was pathetic: a ping service to keep functions warm, which defeats the whole "serverless" promise and adds more cost. Containers on Fargate were cheaper and predictable.


—cp


   
ReplyQuote
(@integration_maven)
Reputable Member
Joined: 6 months ago
Posts: 261
 

While I agree with the core premise of using serverless as a trigger, I think dismissing all stateful orchestration is too reductive. There's a valid middle ground.

Your example uses AWS Batch, which is excellent for predictable, large jobs. But for data pipelines that are inherently stepwise - extract, transform, load, notify - a state machine *is* a proper queue, just with more explicit control flow. The cost for Step Functions or equivalent is trivial compared to the dev time saved debugging an SQS DLQ chain. The logic being "visible" is the main benefit; you can see exactly which step failed and retry it without replaying the whole sequence.

The real anti-pattern is trying to make Lambda do the heavy lifting itself. Using it as the glue, exactly as you suggest, is correct. I'd just argue the glue can be a bit smarter than a simple bash script without falling into vendor lock. You can define the state machine in CDK or Terraform and keep the business logic in your container.


IntegrationWizard


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

You're making an assumption that the state machine logic being "visible" is the same as it being *portable*. It's not. The vendor lock is in the definition language and the service's execution model itself. Terraform can define the resource, but good luck moving that Step Functions workflow to Azure or GCP without a total rewrite. An SQS queue is just a queue.

The cost being "trivial" is a classic trap. It's trivial for one pipeline. It's a massive, opaque line item when you have a hundred of them, and debugging that "explicit control flow" through a proprietary visual UI becomes its own special hell. At least with a dead-letter queue I can point a generic tool at it.


Question everything


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Spot on. Using serverless as the trigger and letting purpose-built services do the heavy lifting is the only pattern I've seen work at scale without crazy complexity.

Your cron on servers point is key. We replaced a dozen fragile Lambda-based data jobs with a single, lightweight scheduler container on Fargate. It just polls a database for work and kicks off Batch jobs or state machines. The cost dropped by about 40% because we weren't paying the "stateless tax" of Lambda memory timeouts and constant warm-up tricks. The cron container itself costs pennies and never times out.

It feels less glamorous than a "full serverless" pipeline, but it's boringly reliable. Sometimes the old ways work for a reason.


Keep automating!


   
ReplyQuote