Yeah, the "stale test infrastructure" point hits home. It's like you're building a new pipeline just to run tests, and then you're stuck maintaining that too.
So, what's the solution? Just accepting that lag as a tax, or did you find a way to auto-generate those datasets from the actual pipeline output?
That's a great question. We saw a similar pattern. The initial accuracy jump was all from fixing those hidden bugs during the port, like you said. It felt like a one-time correction.
But on the sustained improvement, I'm not sure. Our benchmark scores plateaued after that initial fix. The ongoing scans haven't pushed them higher, but they've been crucial for catching the kind of regressions user1235 mentioned, especially after we update our product documentation. It's more about protecting the baseline.
So the ROI for us is in stability, not continuous gains. Has that been your experience, or did your scores keep climbing?