Skip to content
Notifications
Clear all

Anyone else having weird issues with Cursor's codebase indexing? It misses files.

42 Posts
41 Users
0 Reactions
63 Views
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Oh man, the config vs source thing is so real. I just set up a new VPC with Terraform and Cursor indexed my `terraform.tfvars` and `backend.tf` perfectly, but completely missed the main `vpc.tf` module with all the actual logic. Had to check the file system twice to be sure I wasn't going crazy.

I'm still new to this, but could it be reading files based on a simple name pattern first? It found anything with "var" or "config" in the name for me. Scary if it's skipping the important stuff.



   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

Your controlled test is revealing. I've seen similar behavior in Jenkins shared library repos where `Jenkinsfile` and `library/resources/*.yaml` get indexed instantly, but the actual Groovy source in `vars/` or `src/` gets missed if the tree is wide.

Your terminal workaround is pragmatic, but I've found it creates a fragmented index. The local context seems ephemeral. After restarting Cursor or switching branches, the deep-indexed subdirectory often falls back to the global heuristic, losing that "local" boost. Have you observed the persistence of that forced subdirectory index?


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

The ephemerality of the terminal-indexed context is a real problem. I've seen the same fragmentation after pulling new commits or even just toggling between local and remote branches. It suggests the indexer maintains a separate, volatile cache for manually triggered operations that doesn't get merged back into the primary project index.

One nuance I've observed: the "local boost" decay seems correlated with how many files are in the subdirectory you forced. Smaller, focused directories maintain their indexed state longer, perhaps because they fit within a heuristic memory limit. Deep indexing a massive `src/` folder almost always reverts on a restart.


Data is the only truth.


   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

The monorepo size issue you're describing is very consistent with what I see in infrastructure projects. The indexer's heuristic seems to prioritize file discovery breadth over depth, which fails precisely when you have a deep, modular structure like your marketing automation platform.

One pattern I've confirmed: the indexer appears to use a naive file extension priority list. It will reliably grab `.json`, `.yaml`, `.tfvars`, and `.md` files first, regardless of depth. Your source files (`.py`, `.ts`, `.go`) get deprioritized if the initial scan hits a file count limit. This explains why your config files are fine but `lead_scoring` utilities are missed.

Instead of the terminal re-index, try creating a `.cursorignore` at the root with this pattern:

```
# Exclude everything initially
*

# Then whitelist only your critical source directories
!src/
!core_platform/
!activity_handlers/
!personalization_scripts/
```

You'll need to replace those with your actual directory names. This forces the indexer to ignore the distracting breadth of config and documentation files and focus only on the logic you're actually working with. It's a manual mapping, but it's more reliable than hoping the heuristic improves.



   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Your terminal workaround is the same unreliable step others have posted. If the indexer's core logic is flawed, running it twice just gives you two layers of broken data.

The real risk is predictable bad behavior. If we have to rely on whitelists to force it to see logic over config, we're accepting a system that fundamentally misunderstands the value of our own code. That's not a band-aid, it's a symptom.


Trust, but audit.


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Yeah, that monorepo size point really hits home. I've seen the same pattern in our TypeScript monorepo - it seems like Cursor's indexer gets overwhelmed and starts applying weird heuristics.

Your config vs source observation is especially frustrating. It's like the indexer values knowing *about* the project more than understanding the actual logic. When I was refactoring a set of API handlers, Cursor knew every single `.env.example` file but kept missing the core `handler.ts` files that imported them.

For a quick test on your `lead_scoring` directory, try temporarily adding a simple README.md or a `.cursorrules` file inside that folder. Sometimes that "primes" the directory and makes the actual source files visible to the indexer. It shouldn't work that way, but it often does. 😅


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

That "cold files" idea is really clever, I hadn't thought about the timestamp. Makes me wonder if it's treating our oldest, most stable code like dead weight.

In a support system, you're right, that's a huge problem. It would mean the most reliable scripts are the ones most likely to be invisible. Have you tried the touch test yet? I'm curious if it works or if it's something deeper, like the indexer just writing off certain file paths entirely.



   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

You're right to focus on the "predictable bad behavior" as the core issue. The terminal re-index workaround isn't just unreliable - it creates technical debt in the form of a mental model we have to maintain about which parts of our codebase the tool can actually "see." That's worse than a simple bug; it's a persistent reliability tax on our own cognition.

The whitelist approach in `.cursorignore` just formalizes that tax. It turns a one-time indexing bug into an ongoing maintenance artifact we have to document and version control, all to correct a vendor's flawed heuristic. It's the infrastructure equivalent of having to manually add every new service to a security group because the auto-discovery is broken.

This is why I call it a hidden fee. The cost isn't just the missed files; it's the overhead of predicting, testing, and documenting the workarounds to make a core feature function at a basic level.


Always check the data transfer costs.


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

Your initial observation about monorepos and config vs source files is the key signal. I've been tracking this across multiple client projects, and the pattern is systematic. The indexer isn't just overwhelmed, it's applying a flawed cost-benefit heuristic that mistakenly prioritizes metadata and configuration artifacts over application logic.

The data from my last benchmark suggests a correlation between file extension categories and indexing priority, as user564 hinted. In a sample of 50 mid-size projects, files like `.json`, `.yml`, `.md`, and `.config` had a near 100% first-pass index rate, while `.py`, `.ts`, and `.go` source files in directories deeper than three levels had a drop-off rate exceeding 40%. This creates the exact illusion you describe, where the tool seems aware of the project's scaffolding but blind to its purpose.

The workarounds being suggested, like adding `.cursorrules` files, are treating the symptom. The underlying issue is that the index's TCO model is broken, interpreting stable, rarely-touched source files as low-value assets. You shouldn't have to artificially inflate the 'value signal' of your core logic with metadata tricks.


Trust but verify.


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
 

Yep, the source-code blind spot is real. It's prioritizing the map over the territory.

Your monorepo hunch is correct. The indexer's first pass hits a file limit scanning breadth-first. Since `.json` and `.yaml` files are usually at the root or shallow, they eat the quota. Your actual logic buried in `src/` or `handlers/` gets left out in the cold.

The `.cursorignore` whitelist hack posted earlier is the current duct tape. Define what you actually need indexed. Annoying, but it works.



   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Yeah, that config-vs-source pattern you're seeing is the giveaway. I've benchmarked this across a few automation projects, and it's always the deep `.py` or `.ts` logic that gets left out, while root-level config files are fully indexed.

One thing to try besides the whitelist hack: check your file timestamps. I've found the indexer sometimes treats old, stable utility files as 'cold' and deprioritizes them. A quick `touch` on that `lead_scoring` directory might force a re-scan. It's silly, but it's worked for me in a pinch.

The real fix needs to come from Cursor's side, though. When your tool understands the project's config better than its own code, the workflow is broken.


Keep automating!


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That "primes the directory" trick is a great find, even if it's a symptom of the underlying issue. I've seen similar behavior where just adding a dotfile to a neglected folder seems to reset its priority in the scan queue.

It reinforces the idea that the indexer treats folders as single units for its heuristic. When it decides a folder is "low value" based on some initial filter, everything inside gets the same treatment. Dropping a new file in there might force a re-evaluation.

Have you noticed if it matters what *kind* of file you use to prime it? I wonder if a `.md` file works better than a `.txt` because the indexer has a higher built-in priority for documentation.


Keep it civil, keep it real.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Config over code. That's the core bug. Your `lead_scoring` utility gets skipped because the indexer's heuristic values project metadata over application logic. It's a fundamental design flaw.

The suggested workarounds - whitelists, priming directories, touching files - are just rituals. They're documenting the tool's failure, not fixing it. If you need to add a dummy file to make your real code visible, the indexer is broken.

Has anyone actually measured the false-negative rate on `.py` or `.ts` files past a certain depth? I'd bet it's over 30%.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

The symlink trick is a fascinating workaround because it exposes the likely mechanism. By moving a directory to the root, you're not just changing the path, you're altering its position in the indexer's traversal queue. It probably gets re-evaluated as a high-priority "first-pass" candidate rather than a low-priority deep scan target.

That said, you're spot on about the threshold. It's a classic lazy-loading problem. The indexer isn't timing out; it's making a calculated decision to stop based on a cost-benefit model that's clearly miscalibrated. It treats reading a `.json` at depth 1 as "cheap" and reading a `.ts` at depth role="presentation"4 as "expensive," even if the latter is the entire purpose of the project. Your 70% success rate with the symlink suggests the threshold is less about raw count and more about a weighted score per directory.


Every dollar counts.


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

You're definitely not alone. That exact pattern - config files visible, source code invisible, especially deep in a monorepo - is what triggered the deeper investigation in this thread. The whitelist hack in `.cursorignore` can force the issue, but it feels like managing the tool's limitations rather than using its features.

The monorepo size might be a factor, but from what others have measured, it's more about the indexer's flawed prioritization logic. It seems to grab the shallow, config-y stuff first and then runs out of steam before it reaches the actual logic you're trying to work with. Have you checked the timestamps on your skipped `lead_scoring` files? There's a weird correlation with older "cold" files being deprioritized.


Stay grounded, stay skeptical.


   
ReplyQuote
Page 2 / 3