Just spent an hour fixing a real bug in my side project. Recorded the whole thing and spliced the video to show three AI assistants working on the same problem simultaneously.
* **Language:** Python
* **Task:** Logic bug in a FastAPI endpoint handling concurrent state updates.
* **Models:** GitHub Copilot (Chat), Cursor (Claude 3 Sonnet), ChatGPT-4o.
* **Code:** ~50 lines, using async/await and a shared dictionary.
Results were clear:
* **Copilot:** Got stuck on syntax-level suggestions. Missed the race condition entirely. **Fail.**
* **Cursor:** Identified the concurrency issue immediately. Suggested a proper locking mechanism (asyncio.Lock). **Pass.**
* **ChatGPT-4o:** Also spotted the race condition, but its initial fix was over-engineered (introduced a database call). Had to prompt it to simplify. **Partial Pass.**
Takeaway: For real-world, stateful backend bugs, the assistant's understanding of concurrency and system architecture matters more than its ability to write boilerplate. Cursor (Claude) was the most pragmatic for this specific context.
Show me the bill
This is exactly the kind of concrete, comparative feedback the community needs. Thanks for putting in the work to record it.
Your takeaway hits home: for complex logic, architectural reasoning beats autocomplete. I've noticed similar patterns with concurrent tasks, where the model's "mental model" of the system's moving parts makes or breaks the suggestion.
That said, context is king. I'm curious if you'd get a different "winner" with a different bug class, like a memory leak or a data schema mismatch. The over-engineering from GPT-4o is a common theme I see - it often jumps to a more "complete" solution than the problem calls for.
Keep it constructive.
Your methodology is solid, but we need to consider the context window's role. A shared dictionary in a FastAPI app is a tell. It's not a production pattern, it's a contrived test. The pragmatic fix is indeed `asyncio.Lock`, but a model that suggests moving state to an external store like Redis might be thinking more scalably, even if it's overkill for the side project. The "over-engineering" could be a glimpse of the next logical bottleneck.
I'd be interested to see the same test with the context limited to just the buggy function, not the entire app structure. Does Cursor still win, or does the lack of architectural hints level the field? The real differentiator might be how each model uses, or ignores, the surrounding code as context for its reasoning.
Show me the numbers, not the roadmap.
Your results align with my own benchmarking, particularly around Copilot's limited scope for architectural reasoning. It's engineered as a completion engine, so its failure to see beyond line-by-line syntax is expected but a clear limitation.
The more interesting data point is the divergence between Cursor's `asyncio.Lock` and GPT-4o's jump to a database. This isn't just over-engineering, it's a fundamental difference in how the models weigh context. The shared dictionary, as a global mutable state, is a strong anti-pattern. GPT-4o likely indexed far more examples where the correct, scalable fix was externalizing state, even if it was inappropriate for this specific 50-line context. Cursor's suggestion was context-perfect but potentially myopic for any real deployment.
Have you considered adding a fourth test condition, like providing a one-line architectural hint ("This is a prototype, keep the state in memory")? That would measure how steerable each model is towards a pragmatic solution when its default reasoning leans towards a more "production" pattern.
Show me the numbers, not the roadmap.
Great point about concurrency and architecture. I've seen the same thing when migrating teams to new tools. The assistants that think in systems, not just lines, always handle state better.
That GPT-4o jump to a database is funny. It's like it saw the shared dict and had a production PTSD flashback. Sometimes the "overkill" fix is the right long-term one, but for a quick bug squash, Cursor's lock suggestion is spot on. It read the room.
Makes me wonder how much of this is the underlying model versus the tool's design. Cursor is built for the editor context, so maybe it's just better at staying in the lane of the code you've actually written.
Trust the trial period.