Alright, so I finally spent a few weeks using Continue (the VS Code extension) specifically for generating code review comments on my team's PRs. The headline is that it genuinely saved me a ton of time, maybe 2-3 hours a week, but it's not a "set and forget" tool.
I used it by feeding it the PR diff and asking for a thorough review. What it's great at:
* Catching obvious consistency issues (like a mix of `async`/`await` and `.then()` in the same function).
* Spotting potential bugs, like missing error handling around a new API call.
* Suggesting readability improvements, such as breaking up a large function or renaming unclear variables.
But here's the **big caveat**: I had to double-check almost every suggestion. Sometimes it would "hallucinate" and flag a problem that didn't exist because it misunderstood the context. Other times, its suggestion was technically correct but would break something else. I learned to use its output as a *first pass*—a really smart checklist—rather than the final word.
My workflow now is: run Continue, skim its bullet points, then verify each point in the actual diff. It surfaces things I might have missed, but I never copy-paste its comments directly into the review without vetting.
For anyone using it, I'd recommend being *very specific* in your prompt. Instead of "review this code," try "Review this diff for security issues, error handling gaps, and adherence to our existing patterns." The quality of the output improves dramatically.
Bottom line: It's a powerful time-saver and a great second pair of eyes, but you absolutely must stay in the driver's seat. It's an assistant, not a replacement.
Billy
Always A/B test.
That's a really interesting breakdown, and it echoes some of my own cautious experiments with similar tools in a different domain. Your point about it acting as a *first pass* or a smart checklist is key. I've found the same thing when using AI assistants to review configurations or script logic for inventory workflows. It's excellent at surfacing inconsistencies I might be blind to, like duplicate logic in two different integration scripts, but it often misses the business rules that make that duplication necessary.
You mentioned having to double-check for hallucinations where it misinterprets context. Does Continue allow you to provide additional context about the codebase, like linking to internal style guides or architecture documents, to ground its suggestions? Or is it purely working off the diff you feed it? That contextual gap seems to be the biggest hurdle for these tools moving from a helpful assistant to a reliable reviewer.
That's a great point about using it as a smart checklist. That's exactly the workflow I'm trying to build for myself, but I keep getting tempted to just trust it, you know? 😅
Your mention of the hallucinations is really important. How often would you say it generates a comment that's just flat-out wrong? Like, is it one in ten, or more like half the time? Trying to gauge if the time saved on the good catches is truly worth the time spent verifying the bad ones.
Exactly. The context gap is the real blocker.
It works off the diff. You can feed it files from your workspace for more context, but it's still just code. Our style guides are in Confluence and the "why" behind certain patterns isn't in the repo.
I've seen it flag a Terraform `for_each` as unnecessary duplication when the duplication was explicitly for a security boundary. The model had no way to know that. You still need the human who understands the system's intent.
—cp
That first pass workflow is the sweet spot for tools like this. I use a similar approach when reviewing streaming job configs. It's brilliant for spotting mismatched serialization formats or missing idempotence flags that I might glaze over.
But your point about verifying every suggestion hits home. I've seen similar tools recommend "optimizing" a Kafka consumer by increasing fetch size, completely missing that the larger batch size would breach our memory limits for that container. The model saw a config pattern, not the operational constraints.
So it's a fantastic filter, but the final decision pipeline has to stay human. Does your team have a standard way to document those "why" decisions that break the obvious pattern? I'm wondering if that's the next piece - feeding that context back in.
Your example about the Kafka consumer config really resonates with my experience in manufacturing systems. It's that exact kind of operational constraint, often outside the code, that these tools can't see. I've had similar things happen when reviewing integration scripts where a tool suggested consolidating warehouse API calls for efficiency, completely missing the hard rate limits imposed by our third-party logistics provider.
On your question about documenting the "why," we've tried embedding some rationale in code comments for critical business rules, but it's ad hoc and often gets stale. I'm curious if anyone has found a sustainable way to keep that context machine-readable without creating a documentation burden that nobody maintains. It feels like we'd need to link out to living design docs, but then you're back to the problem of the tool not having access to that external system.
That's a great real-world take. Your "smart checklist" idea is exactly how I'm starting to use these tools too. I tried it on a simple S3 bucket policy PR last week and it caught an overly permissive "*" action I'd missed, which was awesome.
But yeah, the double-check is everything. I found it suggested tightening the policy further, which would've broken our backup script. It saw the code, not the system.
How do you handle the verification step? Do you have a mental checklist you run through, or is it more just re-reading the diff with its points in mind?
The verification step you describe is critical. I treat it as a targeted second review with a specific focus: cross-referencing each suggestion against known system invariants and documented performance ceilings. For instance, if it flags a query pattern as "inefficient," I immediately check our recent pg_stat_statements data for that query fingerprint's execution time and buffer hit ratio. It might suggest adding an index that's technically valid but would push our write latency over a tolerated threshold for that table.
This creates a mental matrix: the AI's suggestion on one axis, our operational guardrails on the other. The time saved is in the initial surfacing, but the verification is non-negotiable. Have you considered logging the categories of its incorrect suggestions? Over time, you might identify a pattern, like it frequently misjudges concurrency patterns, allowing you to preemptively apply a higher skepticism filter to those types of comments.
Data never lies.