> these tools can only navigate patterns they've seen enough of.
This is the core limitation for any statistical model. It's not just a problem for legacy migrations. I've seen it confidently suggest N+1 query patterns in a new service because the training corpus was full of eager loading anti-patterns from older tutorials. It optimizes for token likelihood, not for runtime performance or architectural integrity.
Your point about auditing holds even for greenfield work. If you can't treat its output as correct, you're just adding a review step for generated code, which often takes longer than writing the straightforward, performant version yourself. The latency hit comes from the constant context switch between reading its suggestion and tracing back to your actual data model.
--perf
That's such a great example with the N+1 queries. It highlights how the tool's "help" can actually cement bad patterns if you're not vigilant.
It makes me think the review step you mentioned is actually a new, hidden skill. You're not just checking for bugs, you're auditing for architectural drift and performance antipatterns that the model thinks are normal. That mental load is real.
Maybe the only safe use is for truly generic boilerplate you already know by heart, where the review is instant. Anything with actual logic means you're now wearing an extra hat.