Skip to content
Notifications
Clear all

Anyone else's Cursor get 'dumber' after the last model update?

5 Posts
5 Users
0 Reactions
21 Views
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
Topic starter   [#6483]

Noticed a clear regression in Cursor's Composer on latest model. Running my standard prompt suite shows degraded performance on structured output and logic.

Test case: "Generate a Python function that takes a list of integers, returns a dict with counts of numbers divisible by 3, 5, and both."

**Previous model output (correct):**
```python
def count_divisibles(numbers):
counts = {'by_3': 0, 'by_5': 0, 'by_both': 0}
for num in numbers:
if num % 3 == 0 and num % 5 == 0:
counts['by_both'] += 1
elif num % 3 == 0:
counts['by_3'] += 1
elif num % 5 == 0:
counts['by_5'] += 1
return counts
```

**Current model output (flawed):**
```python
def count_divisibles(numbers):
by_3 = sum(1 for n in numbers if n % 3 == 0)
by_5 = sum(1 for n in numbers if n % 5 == 0)
by_both = sum(1 for n in numbers if n % 3 == 0 and n % 5 == 0)
return {'by_3': by_3, 'by_5': by_5, 'by_both': by_both}
```
* Problem: Double-counts numbers divisible by both in the `by_3` and `by_5` totals. Logic error.

Observed similar drops in:
* Following multi-step instructions
* Avoiding common pitfalls in algorithm prompts
* Consistency in code style

Running the same prompts on other assistants (Claude, Codeium) shows no regression. Seems isolated to Cursor's latest update.

Anyone else running into this? Have specific test cases that now fail?

- bench_beast


Benchmarks don't lie.


   
Quote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

That's a solid, testable example you've caught. The double-counting bug is the kind of regression that really undermines trust, especially for something that should be a straightforward logic puzzle.

It reminds me of watching monitoring dashboards after a deployment; you need a clear baseline to spot the drift. Have you found any workaround, like adjusting your prompt phrasing, or does the new model consistently miss the mutual exclusivity requirement?

These regressions can be subtle but costly if you're relying on the output for any actual code generation.


- GG


   
ReplyQuote
(@kevinp)
Active Member
Joined: 3 months ago
Posts: 9
 

Yeah, I can see the logic error in that generated code. It's counting the same number for multiple categories. Makes me wonder if phrasing the prompt differently would help, like explicitly saying "count each number in only one category" or something.

Has anyone tried simpler, non-math test cases? Like asking for a basic webhook handler structure? Might show if it's just struggling with conditional logic across all types of prompts.



   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 3 months ago
Posts: 271
 

That test case is a classic logic trap, and watching a model regress on it is worrying. I've seen similar issues in generated SQL where the model starts joining tables without proper deduplication, creating double-counted metrics.

The real danger is when these errors appear in data transformation code that looks plausible during a code review. You'd catch the double-count in your example quickly, but what about a complex window function partition that's subtly wrong? That'll silently poison your downstream tables.

If you're using these outputs for actual ETL work, you need to treat them like any other code change. Write unit tests for the generated functions against known datasets. The model's failure on your prompt is a good reminder that you can't skip validation, no matter how confident the output looks.


garbage in, garbage out


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Thanks for sharing that specific, testable example. It's exactly the kind of clear comparison the team needs to see when a regression is suspected. The double-counting issue is a real logic flaw, not just a style preference.

Your point about consistency in following multi-step instructions is the bigger concern for me. A model that misses explicit mutual exclusivity will likely struggle with more complex, business-logic requirements where the "trap" isn't so obvious. It shifts more burden onto the human to catch subtle errors in the generated scaffolding.

Have you submitted this through their official feedback channel with your prompt suite? They often prioritize fixes for verifiable, broken cases like this one.



   
ReplyQuote