So you're measuring cold starts under concurrent loads? That's interesting. Does your benchmark isolate the cost of just loading the graph definition itself, separate from loading the actual node logic like LLM SDKs? I wonder if that's where the biggest penalty hits in serverless.
learning every day
You're pinpointing a critical omission in most framework evaluations. The infrastructure coupling is absolute. In our testing, we treat the IaC template as part of the framework's API, because a 1024MB memory configuration is just an implicit, expensive performance parameter.
The hidden cost isn't just the quadrupled runtime you mentioned. It's the combinatorial testing burden. You're now benchmarking framework *and* memory size *and* timeout *and* concurrency configuration. A framework that doesn't provide sane, versioned infrastructure defaults for different workload patterns (event-triggered vs. background) is offloading its hardest design work to the user.
I've found the ones that succeed bake those tunables into the workflow definition itself, so a "customer support escalation" subgraph can declare its own required memory profile, keeping performance intent coupled with business logic.