After hearing so much hype about AI-powered code review, I decided to put Continue's "Review Changes" feature to the test on a recent pull request. The TL;DR is that it genuinely saved me a significant amount of time, but it came with a crucial caveat.
I was reviewing a ~400 line PR that touched our API client and some related test automation. Instead of my usual line-by-line reading, I opened the PR in Continue and asked it for a general review. Within seconds, it generated a list of comments covering:
* Potential error handling gaps in a new method.
* Suggestions for consolidating two similar test helper functions.
* A spot where a null check might be redundant based on our internal library's guarantees.
This gave me a fantastic starting framework. I'd estimate it cut my initial review time by about 60%. I could jump straight into evaluating its suggestions rather than building the review from scratch.
However, the "double-check" part of my title is critical. I found I had to treat its output as a very smart, but junior, engineer's first pass. For example:
* It correctly flagged a "magic number," but its suggested constant name didn't align with our team's naming convention.
* It missed a subtle race condition possibility in an updated integration test because it wasn't aware of our broader system context.
* One suggestion about logging levels was technically sound, but went against our current incident postmortem playbook.
My workflow ended up being: generate the comments, then methodically go through each one to validate, contextualize, and often rephrase it before posting. This extra verification step is non-negotiable.
For teams considering this, my takeaways are:
* **Great for:** catching common pitfalls, suggesting standard improvements, and speeding up the initial pass. It's excellent for bulk.
* **Requires:** deep domain knowledge to vet its advice. You are still the responsible engineer.
* **Best for:** reviewers, not as a substitute for the PR author's own testing. It didn't replace our CI pipeline or my own test runs.
Has anyone else developed a specific workflow for using AI in reviews? I'm thinking of creating a small checklist for myself on what to always double-check.
gh2
ship early, test often