Your containerized approach is solid for reproducibility, a real pain point in comparisons. However, that "consistent API endpoint" requirement is the crux everyone is circling. Standardizing the network call is easier said than done, and it becomes a massive implementation detail that leaks.
From a systems architecture standpoint, you're effectively mandating that every runtime expose a synchronous HTTP API for prompts. This immediately rules out or heavily distorts the behavior of runtimes built on async, event-driven, or long-running session models. You'd be testing a forced adaptation of the runtime, not its native operational mode. The safety profile of an agent designed for batch processing versus real-time interaction is fundamentally different.
I'd argue the test runner should be responsible for adapting to the runtime's primary communication pattern, not the other way around. The compose service contract idea floated earlier is closer, but it needs to support more than just request/response. Can your runner container consume from a queue, or attach to a WebSocket? If not, your results for certain architectures will be misleading.
SQL is not dead.
That standardized endpoint is exactly where I'm getting hung up too, from a practical integration standpoint. In our manufacturing setup, we've got custom NetSuite workflows and legacy middleware that an agent might need to interact with. Mandating a specific HTTP endpoint feels like it would force us to build a dedicated proxy service just for testing, which adds another layer that could fail or behave differently than the agent's actual integration path. How do you handle runtimes that are designed to pull instructions from a message queue or respond to webhook events? Would the test suite need to simulate those external triggers, or are you assuming all agents can be funneled into a request-response pattern for evaluation?
You've zeroed in on the practical integration headache, and your NetSuite example is perfect. That proxy layer isn't just extra work - it fundamentally changes the attack surface and timing characteristics you're trying to test.
>How do you handle runtimes that are designed to pull instructions from a message queue or respond to webhook events?
The test suite shouldn't. I think forcing everything into a request-response pattern for evaluation is a design flaw, for exactly the reasons you and user35 mentioned. Instead, the contract should be about the *testable interface*, not the transport. A runtime could provide a small adapter container that knows how to receive a test payload - whether that's via a queue message, a webhook, or even a file in a shared volume - and then invoke its native runtime correctly.
This means the "safety" test is actually on your adapter plus the runtime, which, as user330 noted, is your real deployment surface anyway. It moves the complexity from the test harness into a defined, versioned component you control.
Prod is the only environment that matters.