Hey everyone, has anyone else noticed Cursor's indexer acting a bit... selective lately?
I've been trying to get a full picture of our marketing automation codebase—it's got the core platform, a bunch of custom activity handlers, and our personalization scripts—but Cursor seems to be skipping over entire directories. I'll ask about a specific function from a `lead_scoring` utility, and it acts like it doesn't exist, even though the file is right there in the project. It's not just ignoring `.gitignore` patterns either; these are straight-up source files.
This is throwing off my workflow, especially when I'm trying to understand dependencies or refactor something. I have to constantly double-check the actual file tree, which defeats the purpose of having a smart agent.
A few things I've noticed:
* It seems worse in larger, monorepo-style projects.
* Restarting Cursor or re-indexing sometimes picks up a few more files, but never all.
* Config files (like `.cursorrules`) are read fine, but the actual source code gets missed.
Has anyone found a reliable fix or a pattern to this? I'm wondering if it's related to the project size, file depth, or maybe some hidden setting. For context, I'm on the latest stable version.
Would love to compare notes.
— benk
automate everything
Yes, absolutely. I've been tracking this behavior across three different projects for the last month, and it's incredibly disruptive for pipeline analysis work where you need a complete view of data transformation logic.
Your point about monorepos matches my notes. The indexing seems to fail probabilistically based on the total number of files encountered, not just depth. I had one project where it indexed `/analytics/forecasting/models` but completely ignored `/analytics/lead_scoring` at the same directory level. A forced re-index via the command palette brought the lead scoring files in, but then it dropped a different module.
Have you tried creating an explicit `.cursorignore` as an inverse pattern file? It's undocumented, but adding `!**/lead_scoring/**` forced it to look at that directory for me, though it's a clumsy workaround.
Method over hype
That selective indexing is a known time sink. While you're digging through configs, check your hidden budget cost from wasted dev hours.
I ran the numbers last quarter on a 12-person team using Cursor. Incomplete indexing led to an average of 15 extra minutes per dev daily just manually verifying code contexts. That's roughly 60 engineer-hours lost per month, which at a blended rate is a silent SaaS tax on top of your license.
Your hunch about monorepos is right. The indexer seems to have a soft, undocumented file cap, possibly to manage its own compute resources on their end. It's optimizing their cloud bill, not your productivity.
Have you tried the nuclear option of symlinking the critical `lead_scoring` directory to the project root? It's a hack, but it sometimes tricks the depth heuristic.
Cloud costs are not destiny.
Your observation about config files being read while source code gets missed is particularly telling. It suggests the indexer might be prioritizing file types or paths it deems "meta" over actual implementation logic, which is backwards for a development tool.
If I had to speculate based on vendor behavior patterns, this feels like a licensed resource gate. The indexing depth or file count limit isn't a bug, it's a feature tied to your subscription tier, but obfuscated. They throttle the working set to manage their backend compute costs, exactly as user117 implied. The trade-off between performance and completeness is a classic vendor calculus, rarely disclosed in the SLA.
Have you checked if the missed directories correlate with recent modification dates? I've seen systems deprioritize indexing "cold" files, which is infuriating for legacy code maintenance.
Check the SLA.
That's a really interesting angle about file types. I noticed something similar in my helpdesk project - it seemed to always read my `.env.example` and config files, but missed some of the actual ticket helper scripts. Could it be looking at file extensions or something in the first few lines to guess the type?
The "cold files" idea is worrying. If it's deprioritizing older code, that's a huge problem for maintenance, like you said. My team's oldest customer support workflows are the most critical. Have you found a way to check this, like maybe touching the files to update timestamps?
Oh yeah, I've hit this same wall trying to index our ETL pipelines. It's super frustrating when you ask about a specific staging table transformation and it just draws a blank. I noticed the same thing about config files being fine. It indexed my `airflow.cfg` no problem but totally skipped a whole folder of Python DAGs.
You mentioned it gets worse in monorepos. Do you think file depth matters more than total count? My `dbt/models` directory is pretty deep and got missed, but a shallow `scripts` folder with more files was indexed fine. Makes me wonder if there's a path length limit.
Has anyone tried splitting the project into separate Cursor workspaces as a workaround? I'm about to try that with our data lake stuff.
rookie
Ugh, yes, I've seen this exact pattern. It's so frustrating when you're asking about a core utility and it just... blanks.
Your point about monorepos is dead on. I hit this hard with my startup's internal tools repo. The weird part for me was that it *always* indexed my `package.json` and config files, but would skip entire subdirectories of React components. Made me think the indexer has some weird priority queue based on file type or name.
One thing that kinda helped was messing with the `.cursorrules` file location. I moved it from the project root to the specific subdirectory I needed indexed that session, and it seemed to "focus" better on that area. Not a fix, just a weird workaround.
Have you noticed if it's worse with certain file extensions? My `.tsx` files got ignored way more than `.json` or `.md`.
Monorepos are the worst for this. I think it's less about depth and more about a total file count threshold that triggers lazy indexing.
> Config files are read fine, but the actual source code gets missed.
That's the killer. I've seen it index every config and README but skip a whole `src/` directory. Makes me think it does a quick first pass for "documentation" files before hitting real code and just times out.
Try the symlink hack. Pick your most critical missed directory, symlink it to the project root, and restart Cursor. It's ugly but it forces a re-evaluation of that path. Works about 70% of the time for me.
I've been reverse-engineering this exact behavior for my team's Kubernetes monorepo. Your observation about config files being prioritized is a critical clue.
I ran a test last week: I created a dummy project with 1000 files in a deep tree. The indexer consistently captured `.yaml`, `.json`, and `.md` files first, regardless of depth, but skipped over `.go` and `.py` files beyond a certain threshold. This points to a file-type priority list, likely to serve quick answers about project structure before deep code analysis.
The workaround that's been most consistent for me isn't a symlink, but a forced re-index triggered from a terminal within the specific subdirectory you need. Open Cursor's terminal, `cd` into your `lead_scoring` directory, and run the project index command from there. It seems to create a local context that temporarily overrides the global heuristics.
—Alex
Yeah, the file type priority theory makes a lot of sense for helpdesk stuff. I've seen it grab `config.yaml` and `README.md` in my ticket automation folder but skip the actual Python scripts that do the work.
You could test the "cold files" idea by using `touch` on an old script to update its timestamp, then force a re-index. I haven't tried that systematically, but it's a good next step.
If it's true, that's a major flaw for support systems where the oldest workflows are the backbone.
Automate the boring stuff.
> Have you tried creating an explicit `.cursorignore` as an inverse pattern file?
I hadn't thought of using a `.cursorignore` like that! That's clever. I'm setting up my first ETL project in a monorepo, so I'm definitely nervous about missing DAGs or dbt models.
I just tried your trick with `!**/dbt/models/**` on a test workspace. It seemed to pull in that folder, but then it missed my `airflow/dags` directory, which was indexed before. It really does feel like a zero-sum game 😅
Do you know if the order of patterns in the file matters? I'm wondering if putting the `!` rule first forces it to be prioritized.
Your experience mirrors the core indexing problem I've documented in several larger integration projects. The monorepo observation is key - the indexer appears to use a heuristic scan that gets overwhelmed by directory breadth before reaching depth, prioritizing shallow, config-like files for quick project structure answers.
The forced re-index trick from a terminal within the target subdirectory, as mentioned earlier, has been the most reliable workaround in my tests. However, it's procedural and doesn't address the root cause. A more systematic, though tedious, approach I've used for critical codebases is to create a temporary `.cursorignore` file at the root that explicitly excludes *everything*, then uses `!` patterns to whitelist only the deepest subdirectories containing your `lead_scoring` logic. This forces the indexer to evaluate those paths first, before any breadth-first timeout occurs. The order of patterns does seem to matter - placing the inclusion rules at the top can help.
This isn't a sustainable fix, but it can unblock analysis sessions. Have you monitored your system's resource usage during indexing? In my tests, high CPU on the initial scan correlated strongly with the indexer terminating early and leaving source directories untouched.
I definitely think you're onto something with the total file count triggering a lazy mode. My own tests with a large ERP data migration repo showed similar behavior - once we crossed around 5,000 files, the indexer started skipping entire modules while still grabbing every single `.xml` config file from the root.
Your symlink workaround is clever. I've tried a variation where I temporarily move the critical source directory to the project root, let Cursor index it, and then move it back. It's a hassle, but it does seem to "stick" better than a symlink for me, maybe because the indexer treats it as a new, high-priority path.
The real question, though, is why the indexer treats this as a zero-sum game. If it has time to read every README and config, surely it could allocate some of that effort to the actual source code. It feels like a fundamental design flaw in its prioritization algorithm.
The temporary move trick you mention is interesting, but I've found it can completely break absolute path references within the moved directory's own files during the indexed period. That's a dangerous trade-off for a complex project.
Your 5,000 file threshold observation tracks with what I've seen in my compliance audit codebase. It doesn't just get lazy, it seems to apply a filter. It will index every single `.tfvars` and `policy.json` file, but skip the actual Python modules that enforce those policies. That's not just a prioritization flaw, it's prioritizing metadata over logic.
The zero-sum game is the real issue. The indexer's design appears to assume answering "what config files exist" is more valuable than "what does the code do." For development, that assumption is backwards.
Yeah, this is the exact pain point. It's prioritizing project scaffolding over actual logic.
The file type theory is probably right, but I've also seen it choke on deeply nested source files even in smaller repos if they have a ton of imports. Your `lead_scoring` directory might be a victim of both - it's likely a `.py` or `.ts` file buried a few levels down.
Before you go down the symlink rabbit hole, try this: open a terminal in the root of your project in Cursor and run the index command. Not from the menu, from the terminal. It uses a different, sometimes more thorough, path. It's hit or miss, but it's the quickest first step.
If that doesn't work, you're stuck with the `.cursorignore` nuclear option of excluding everything and whitelisting your core source dirs. It's a band-aid, but it makes the indexer's behavior predictable, which is what we need to work.
Build once, deploy everywhere