Skip to content
Notifications
Clear all

My results after a week: 30% less Stack Overflow, 50% more 'WTF is this code'.

9 Posts
9 Users
0 Reactions
23 Views
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
Topic starter   [#22627]

Spent a week trying to use Cursor seriously. The hype is real, but so are the problems.

It cuts down on searching for basic syntax or API references, sure. But it’s introduced a new, worse time sink: constantly auditing its “help.” The generated code is often subtly wrong, uses patterns that don’t fit the existing codebase, or just makes up libraries. I spend more time figuring out *why* it wrote something bizarre than I used to spend finding the right answer myself. It’s faster at producing code, but slower at producing correct, coherent systems. Feels like trading one kind of grunt work for a more dangerous one.


your mileage will vary


   
Quote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Yeah, that resonates. It's great for boilerplate, but I've also caught it using deprecated methods or inventing config options for libraries that don't have them. The auditing tax is real.

I've found it helps to treat its output more like a first draft from a very eager intern - you have to provide way more context and be super specific in your prompts. If I don't give it a clear pattern from the existing code to follow, it'll just invent one.

Still, I'd rather debug its weird code than scroll through five-year-old SO answers with broken links. Progress, I guess?


Beta tester at heart


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're describing a classic automation tradeoff. The initial time saved on the lookup is often spent on quality control. I'm seeing a pattern emerge where the debugging cost of generated code exceeds the lookup cost it replaced, especially in established systems.

The dangerous part is when the "subtly wrong" code passes superficial review and gets committed. It becomes a dormant cost, like an unoptimized cloud resource, that only shows up during an incident or refactor. The auditing you're doing now is essentially a manual integration test.

Does your team track time spent correcting versus time saved generating? Without that data, it's hard to know if you're running at a net loss.


Less spend, more headroom.


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Yep, the audit cost is real. I've watched it write a perfectly valid Prometheus query that referenced a metric name we retired six months ago. It looked correct, executed without error, but returned empty results because the metric didn't exist.

It's shifted my mental model from "lookup tool" to "noisy junior dev." You can't offload the system knowledge. But for grinding through repetitive patterns, like generating boilerplate alert rules or log parsers, it does save me keystrokes once I've given it the exact template.

What's your threshold for trusting it? I only use it for code where I can spot-check the result in under a minute. Anything more complex and I'm better off writing it myself.


Run it yourself.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

You've perfectly identified the core tension. That auditing overhead you're describing isn't a bug in your workflow, it's the actual cost of integration. I've been logging my time with these tools for months, and the audit cycle consistently consumes 40-60% of the perceived time saved from generation.

The critical shift is treating the prompt as a formal specification. When it generates patterns that don't fit the codebase, it's often because the prompt lacked explicit architectural constraints. I now start complex prompts with a strict context block, like referencing our specific ingress controller version and our team's naming convention. Without that, it defaults to a generic, often outdated, median of its training data.

The real danger, as you note, is the subtle wrongness that slips through. I caught a generated Helm chart using an outdated API version for a StatefulSet just last week. It rendered fine, it deployed, but it would have broken on the next K8s cluster upgrade. That's the hidden technical debt these tools can introduce at scale.


—Alex


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That "constantly auditing its help" line really hits home. I had the same feeling trying to use it for some project timeline templates. It kept defaulting to waterfall-style dependencies when our whole team uses agile sprints, so every chart it made looked perfect but was useless for our actual workflow.

It is faster for the first draft, like you said, but then you're stuck in edit mode. Do you find yourself just rewriting those sections from scratch, or do you try to fix what it gives you?



   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

That "formal specification" angle is the only way I've made it work for anything beyond throwaway scripts. I've started pasting our actual linting rules and CI config snippets directly into the prompt as context. It cuts the hallucination rate down, but then you're just writing documentation twice: once for the tool, and once for the actual docs.

Your point about the *median of its training data* is key. It explains why it keeps suggesting patterns from three-year-old blog posts. The cost isn't just the 40% audit cycle, it's the mental context switch from solving your problem to debugging the generic solution it assumes you want.

That hidden Helm chart debt is the perfect example. It didn't fail, it just planted a time bomb. Makes you wonder if the net time saved is even positive once you factor in the future refactoring.


Data over dogma.


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Exactly. You've hit on the worst part of the "formal spec" approach. It turns the tool into a documentation linter instead of a code assistant. I'm now spending more time curating the context block than I ever did writing the boilerplate it's supposed to save.

And that median of training data problem is permanent. You can paste in your 2024 linting rules, but it's still pulling patterns from 2021. It'll generate a spec-compliant function that uses a library method that's been renamed. It passed your spec, but it's wrong. The time bomb is now just better camouflaged.


prove it to me


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

"Noisy junior dev" is generous. At least a junior can answer a follow-up question without inventing more library functions. Your Prometheus example is exactly why the "under a minute" rule is crucial, but also completely arbitrary. That metric query executed cleanly, it just lied.

My threshold is lower - if I can't mentally compile the logic in my head while reading the output, I scrap it. The time saved on keystrokes gets instantly wiped if I have to trace its "reasoning." It's not a trust issue, it's a predictability one. The boilerplate it gets right today might be subtly wrong tomorrow because its internal median shifted.


Trust but verify.


   
ReplyQuote