I've run Copilot across a dozen codebases for six months. The pattern is clear: it's most useful when the code is inconsistent, poorly documented, or uses repetitive boilerplate. In a clean, well-structured codebase with strong conventions, its suggestions become noise faster.
Examples:
- **Legacy Spring service**: Copilot saved hours guessing method names and filling in tedious DAO patterns.
- **Modern Scala/Akka Typed project**: It constantly suggested outdated `Actor` APIs, ignored the typed protocol, and created more work rejecting wrong code.
Key observation:
* It's a pattern matcher, not a reasoner. Messy code has more obvious, repeated patterns.
* It struggles with high-level architectural patterns or custom DSLs.
* The "value" is often just saving you from typing what you already know you need.
Has anyone else quantified this? I'm tracking keystrokes saved vs. time spent correcting bad suggestions. The delta is telling.
—gp
Data over opinions