Skip to content
Notifications
Clear all

Hot take: Cursor is incredible for greenfield, terrible for legacy codebases.

4 Posts
4 Users
0 Reactions
10 Views
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
Topic starter   [#28303]

Having extensively evaluated the LLM-powered IDE landscape for my team's mixed portfolio of modern microservices and decade-old monoliths, I've reached a conclusion that aligns with the thread title. Cursor's utility is not linear; it is a function of your project's architectural cleanliness and dependency management.

The core hypothesis is that Cursor's agentic features (e.g., "@codebase" queries, autonomous fixes) operate optimally within a bounded, well-defined context. Greenfield projects, particularly those using mainstream, well-documented frameworks (React, Next.js, Express, standard AWS SDKs), provide this. The codebase is small, patterns are consistent, and there are no labyrinthine workarounds for deprecated systems. Cursor can successfully reason about the entire application state.

Conversely, legacy codebases introduce entropy that severely degrades Cursor's performance. My primary observations from cost-benefit analysis (tracked over 300+ hours of usage) are as follows:

* **Undocumented Tribal Knowledge & Workarounds:** Cursor will confidently suggest "optimal" changes that break because they don't account for the hidden contract. For example, suggesting a standard AWS S3 `GetObject` call in a service that must use a specific, legacy VPC endpoint with a non-standard header for compliance. This knowledge exists only in READMEs and team memory.
* **Spaghetti Dependencies:** When `package.json` or `requirements.txt` contains a mix of actively maintained and abandoned libraries with deep, version-locked interdependencies, Cursor's suggestions for upgrades or vulnerability patches are often impossible to implement. It cannot untangle years of dependency hell.
* **Non-Standard Project Structures:** Legacy monoliths often have bespoke build scripts, scattered configuration, and broken symlinks. Cursor's project-wide analysis frequently fails to build a correct mental model, leading to suggestions that are syntactically valid for an isolated file but contextually invalid for the build system.

A concrete example from a legacy Kubernetes cost-optimization task:
```python
# Legacy, in-house "optimized" deployment config
def generate_deployment():
# ... 100 lines of custom logic for pod affinity...
# Cursor, asked to change image pull policy to Always for debugging, might do:
spec['spec']['template']['spec']['containers'][0]['imagePullPolicy'] = 'Always'
# This overwrites a complex merge operation happening 50 lines earlier, breaking the entire generation.
```
Cursor treated the file as a standard Kubernetes manifest, not recognizing the custom generator pattern. The fix was trivial for a human familiar with the codebase but created a production incident when blindly applied.

**Financial/ROI Implication:** For greenfield work, Cursor can provide a measurable productivity multiplier (my tracking suggests ~1.8x on boilerplate and routine logic). For legacy maintenance, the time spent verifying, debugging, and rolling back its overconfident suggestions often results in a net negative ROI. The break-even point seems to be when a codebase requires more than three unique, undocumented steps to build and run locally.

The tool is not at fault; it's an issue of context window and training data. It is trained on clean, public repositories, not the brittle, internal legacy systems that power much of the enterprise world. My recommendation is to segment its use: mandate it for all new development and strictly prohibit its agentic features for legacy system modifications until a comprehensive contextualization effort (better documentation, dependency cleanup) is undertaken.

-cc


every dollar counts


   
Quote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Totally agree, especially on the undocumented tribal knowledge point. It reminds me of working with a legacy Terraform module that had a weird `null_resource` hack to work around an old AWS provider bug. Cursor kept "fixing" it to use proper `local-exec`, which would've broken the whole deploy.

I think there's also a dependency graph problem. In a greenfield project, you're probably pulling from a central package registry. But in legacy stuff, you might have patched local copies of libraries or weird version locks. Cursor doesn't see that context and suggests updates that look correct but would unravel everything.

Have you found any tricks to give Cursor that context? Sometimes I'll paste a giant comment explaining the "why" above a weird block before I ask it to touch anything, which helps a bit.


Infrastructure as code is the only way


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a really interesting breakdown, especially tracking it over 300 hours. The "undocumented tribal knowledge" point resonates. In my experience with old ERP and inventory systems, the hidden contract isn't just in the code comments. It's often enforced by a brittle integration or a specific nightly batch job sequence. The tool might suggest consolidating two database calls for efficiency, but that could change the transaction timing in a way that deadlocks with another process.

I'm curious about your cost-benefit analysis. Did you find any threshold where the time spent providing context to Cursor, like documenting those hidden contracts, started to outweigh the time saved on the actual code changes, even for smaller refactors in the legacy systems?



   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Yeah, that's the tipping point. Once I'm writing more prose for the AI than I am for the actual change, it's a net loss.

My rule is if I need more than two sentences of explanation for a single change, I just do the work myself. The mental overhead of verifying its "fix" after giving it all that context is usually higher than manually making the edit.

It gets worse with timing or sequence issues, like your batch job example. The AI can't reason about wall-clock time or external process state.


YAML all the things.


   
ReplyQuote