I've been systematically testing several AI coding assistants on a specific task: reviewing a common, real-world code snippet for security and performance issues. I ran the same prompt 100 times across a few different services (Claude, ChatGPT, Gemini, and one internal tool) and tallied the results. Out of 100 reviews, 22 were objectively wrong in their analysis, and 15 were what I'd classify as "dangerous" – suggestions that would introduce a vulnerability, a bug, or a significant performance regression if followed.
The prompt I used was a simple code review request for a Python function that processes user uploads. It contained a subtle path traversal risk and an inefficient nested loop. The goal was to see if the assistant would catch both.
What surprised me wasn't the error rate itself, but the nature of the failures. The "wrong" reviews often missed the issues entirely, giving the code a clean bill of health. The "dangerous" ones, however, were more concerning. For example, several assistants correctly identified the potential for a path traversal attack, but then suggested a "fix" using `os.path.normpath()` without sufficient validation, which is still vulnerable. Others suggested replacing the nested loop with a list comprehension that, due to a scope misunderstanding, would have broken the logic entirely.
I'm trying to understand the practical differences between these tools for real, high-stakes tasks like code review. Has anyone else done similar controlled comparisons, specifically between the more developer-focused assistants? I'm particularly interested in how tools like Cursor, Claude Code, and GitHub Copilot compare when the prompt requires spotting logical flaws rather than just generating new code. I'd like to know which one might be most reliable as a second pair of eyes, where a false positive or a bad suggestion could be costly.
My takeaway so far is that they can be useful for spotting *some* patterns, but you absolutely must have the expertise to validate their output. They can confidently suggest fixes that are subtly broken. I'm now wondering if any of them are statistically significantly better than others for this kind of analytical task.
Thanks!