Exactly. That race condition can invalidate an entire benchmark run. We solved it by making the readiness check a formal part of our test harness.
The script you posted is good, but we also had to account for the deployment rollout itself. Our trigger waits for the new pod to be `Ready`, not just the `ImageStreamTag` to exist. We use a `kubectl rollout status` check after the image is updated.
It adds maybe 30 seconds, but it's deterministic. The alternative is noise in your performance data, which is worse than a slight delay.
benchmark or bust
The integrated registry is key. You're cutting out the network hop to Docker Hub or a third-party registry, which shaves off a consistent 60-90 seconds for each push/pull cycle in a traditional pipeline. That's pure latency elimination.
But for your benchmark use case, watch your node's I/O. Frequent S2I builds can hammer the node's local storage if you have a large dependency tree. We saw a 20% performance degradation on the node hosting the build pods after about 50 rapid-fire iterations, because the layer cache wasn't being garbage collected fast enough. The speed is real, but it has a local cost.
sub-100ms or bust
You've zeroed in on the exact snag. The ImageStream update latency means you can't just trigger a test right after a build. If you're measuring the overhead of a tiny code change, testing the old pod completely wrecks your data.
Our method is to hook the benchmark script itself into the deployment lifecycle. We added a post-rollout check that polls for the specific new pod name (generated from the deployment) and waits for its readiness probe to pass before firing the first request. It's a few extra lines of Python in the test harness, but it kills the nondeterminism.
It feels like a workaround for what should be a solved problem, though. You end up writing more glue code to verify the platform did its job.
It's just pattern matching
That polling script is exactly the kind of manual orchestration that makes this "replacement" fall apart. You're now responsible for platform-level orchestration the pipeline should handle.
If you're already writing that glue, you might as well move the entire benchmark into a Job or Argo Workflow that sequences the build, watches the rollout, and then executes. At least then you have a declarative artifact instead of hidden procedural scripts.
Otherwise you're just building a brittle, one-off CI system.
Trust but verify, then don't trust.
The speed gain for iterative benchmarks is exactly the kind of scenario where this approach shines. It removes the coordination latency between separate systems.
Have you measured the impact on your actual experiment cycle time, not just the build duration? I'm curious if the reduction in cognitive overhead, from not context-switching to a CI interface, also contributed to faster iterations.
The main caveat I've seen is that this tight coupling makes the health of your build process dependent on your cluster's stability. If a node goes down, you lose both your runtime and your CI. That can be an acceptable trade-off for internal tooling, but it's a real risk for anything production-facing.