Everyone obsesses over the headline milliseconds Amazon, Google, and Microsoft publish. They’re mostly meaningless for actual architecture decisions. The real cost of a cold start isn't just the latency; it's the unpredictability it injects into your system and the vendor-specific traps that lock you into their ecosystem.
I ran a basic test across the three using a Node.js function with a single dependency, hitting it after 15 minutes of inactivity. The results were predictable, but the context is what matters.
* **AWS Lambda (Node.js 18.x, 1024MB):** ~1200ms cold start. The consistent pain point is VPC attachment. If your function needs VPC access, add several seconds. You're paying for that idle ENI.
* **Google Cloud Functions (Node.js 18, 1024MB):** ~800ms cold start. Generally faster out of the box, but their networking and IAM model is a different kind of complexity tax. The "concurrency" model changes the cold start game entirely.
* **Azure Functions (Node.js 18, Elastic Premium plan):** ~1500ms+ cold start on Consumption plan was abysmal. The Premium plan, which you pay for continuously, brought it down to ~900ms. So you're buying your way out of the problem.
The blind spot in most comparisons:
* They ignore provisioning models. Google's max-instance setting vs. AWS Provisioned Concurrency vs. Azure Premium plans are not equivalent. You're comparing apples, oranges, and a subscription fruit box.
* They never mention the cold start impact of connecting to managed services (databases, messaging). The function runtime is one thing; waiting for a database connection pool to initialize is where the real delay happens.
* The open-source alternatives (OpenFaaS, Knative) often have worse cold starts, but they eliminate the lock-in and allow you to understand the entire stack.
What specific scenarios are people actually building where these sub-2-second delays are a critical path problem? And more importantly, how are you mitigating the vendor-designed pitfalls around networking and scaling?
Trust but verify.