Alright, let's cut through the marketing. We're on AWS (Lambda, API Gateway, Dynamo, the usual suspects). Our "stack" is a collection of disjointed services that somehow need end-to-end validation before we push to prod.
I've seen the usual suspects paraded around:
* **Cypress:** Beloved, but feels like bringing a sledgehammer to a serverless problem. The browser focus is overkill when 80% of our "E2E" is hitting HTTP endpoints and checking data states.
* **Playwright:** More flexible, can do API testing, but again, are we just using 20% of its capability and lugging around a browser engine for fun?
* **The classic Jest/Supertest combo:** It works, but calling it "E2E" feels like a stretch when it's often run in the same isolated environment as unit tests.
I'm leaning towards something that treats our cloud resources as a first-class citizen. Criteria:
* Can deploy/teardown test stacks (SAM or CDK).
* Can run tests against the *actual deployed* HTTP endpoints and directly against AWS services (e.g., scan a DynamoDB table after an API call).
* Doesn't require a PhD in custom scripting to manage test data and cleanup.
So, what's the real-world verdict? Is the answer:
1. Biting the bullet with **Playwright** for its mixed-mode testing?
2. A dedicated **AWS-integrated runner** (like something from the AWS DevOps suite I'm cynically ignoring)?
3. A **simple custom script** with the AWS SDK and a good HTTP client, because all these frameworks are over-engineered for our use case?
Show me your battle scars and configs. Bonus points for how you handle idempotent test data.
You're spot-on about the sledgehammer feeling. I've been down that road.
For your criteria, the answer is almost always a small custom script using the AWS SDK itself, orchestrated by a test runner you already know. Use Jest or Vitest as the runner, but make the actual test actions direct API calls (API Gateway) and service calls (DynamoDB scan). The key is your infrastructure-as-code: deploy a dedicated test stack with SAM/CDK before the suite runs.
It sounds obvious, but it's the cleanest fit. You avoid bringing a browser to a serverless fight and keep everything in one language/context. The messy part is test data seeding, but you can use lambda-backed custom resources for that.
Your instinct about first-class citizen treatment is exactly right. I've seen teams get tangled in the framework itself, losing sight of the goal, which is validating the live cloud system.
The real-world answer is often a pragmatic mix. You'll probably end up with a custom script core, like the next poster suggested, but you can lean on some tools to reduce the PhD-in-scripting feeling. AWS's own `aws-cdk-testing` constructs and the `@aws-lambda-powertools/testing` utilities can handle a lot of the heavy lifting for deployment, assertions, and cleanup. They let you write tests that look like your infrastructure code.
My caveat: the messy part isn't the test execution, it's orchestrating the whole pipeline. How do you guarantee the test stack is fully deployed before the HTTP tests run? That's where you might need a tiny bit of scripting glue in your CI/CD, maybe using the AWS CLI's wait commands.
Architect first, buy later
The orchestration point is critical. I've observed that teams often underestimate the eventual consistency of CloudFormation deployments, leading to flaky tests. While AWS CLI wait commands help, they don't handle resource-level readiness, like a Lambda function's initial invocation latency.
A pattern I've documented is using the `DescribeStacks` and `DescribeXxx` APIs in a retry loop with exponential backoff, paired with a canary probe, e.g., a simple `GetItem` to DynamoDB. The `@aws-lambda-powertools/testing` library's `integTests` node actually has utilities for this, but they're not well-advertised. You can see it in their GitHub examples under the `e2e` folder.
This shifts the complexity from CI/CD scripting into the test framework itself, which I find more maintainable.
Nullius in verba
Totally agree that a custom script with the SDK is the way to go. Your point about using Jest/Vitest just as the runner is key - it gives you the structure and assertions without forcing a weird paradigm.
My one caveat from doing this is that direct DynamoDB scans for assertions can get gnarly as your table design evolves. I started using them, but ended up building a tiny abstraction layer for test data queries. It felt like over-engineering, but saved so many headaches when GSIs got added.
The lambda-backed custom resource for seeding is a pro tip. Do you have a pattern for tearing down that test data? I've used DynamoDB TTLs as a "poor man's cleanup" to avoid orphaned data between test runs.
Happy testing!
You're asking the exact right question. That feeling of "we're using a browser tool to test a headless API system" is a massive red flag I've seen derail projects.
You nailed it with "first-class citizen." The real verdict I've lived through? It's the custom SDK script, but you don't have to build it from scratch. The Powertools library mentioned later is a huge step up from raw SDK calls. It gives you those test constructs that feel like your infrastructure, without the PhD. The trick is committing to the pattern fully, including the annoying orchestration layer, or the flakiness will kill morale.
I'd be curious, what's your CI/CD setup like? That often dictates how much of that "stack readiness" pain you have to solve yourself.
Agreed on the CI/CD setup being the deciding factor. We use GitHub Actions, and that orchestration layer is indeed where we felt the pain. We ended up writing a composite action that handles the stack deployment and your exact point about the canary probe, using the AWS CLI and a simple curl health check in a loop before the test suite even starts.
It feels like plumbing work, but it made the actual test scripts so much cleaner and reliable. Without it, the flakiness was a constant drain, exactly as you said.
Keep it real, keep it kind.
Hold on, you're leaning towards "something that treats our cloud resources as a first-class citizen." That's the trap. Frameworks that promise that are just pre-baked custom scripts with their own opinions and learning curve.
The real-world verdict you're asking for is a trick question. There isn't one. You're right to dismiss the browser-heavy tools, but you're hoping for a magic bullet that does the hard parts for you. The hard parts *are* the PhD scripting. The deployment orchestration, the data seeding, the cleanup, the handling of eventual consistency - that's the actual work of E2E testing on AWS. Any tool that abstracts it just moves the complexity to learning its specific DSL and quirks.
So your options are: embrace the custom script and own the complexity, or accept a ton of flakiness. The middle-ground "framework" usually gives you the worst of both.
FOSS advocate
You've nailed the exact friction point. You're hoping for a tool to solve the hard parts, but the hard parts *are* the PhD scripting. The messy orchestration, cleanup, and eventual consistency checks are the actual work.
I'll add this from the on-call trenches: the choice often boils down to where you want that complexity to live. A GitHub Action that manages the deployment and canary probes pushes it into CI. Using `@aws-lambda-powertools/testing` moves it into your test code. The former is a one-time, centralized pain; the latter makes your test logic more portable but every dev needs to understand the pattern. Which complexity is easier for your team to own?
Sleep is for the weak