Skip to content
Notifications
Clear all

Anyone else having weird issues with Cursor's codebase indexing? It misses files.

42 Posts
41 Users
0 Reactions
64 Views
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Oh, the order absolutely matters! It reads the file top-down, so an earlier, broader exclude pattern can block everything before your specific whitelist rule even runs. It's a filter chain.

For your case, you might need to whitelist both directories explicitly. Something like:

*
!**/dbt/models/**
!**/airflow/dags/**

That 'star' at the top excludes everything first, then the 'bang' lines re-include your critical paths. Otherwise, if it's just a list of includes, the default indexing logic might still be fighting you.

It's a bit manual, but for an ETL project, locking down those two key directories should cover the core.



   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Oh man, that's exactly where it falls apart. You're using it to understand dependencies and it's missing the very files that define them.

The "config files read fine, actual source code gets missed" part is the dead giveaway. I see the same thing in SaaS onboarding projects. The indexer grabs the easy, shallow stuff first - your `.json` configs, your `.cursorrules` - and then seems to think its job is done before it even touches the deeper `/handlers` or `/utils` directories.

One thing that's helped me beyond the whitelist hacks: try asking about a file you know is missing *before* you need it in a real task. Sometimes just forcing a query that mentions a deep path seems to nudge it into scanning that area. Not ideal, but it's kept me from having to restart constantly.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Yeah, the pre-query trick is interesting. I've noticed that too - asking a direct question like "what's in `src/utils/lead_scoring.py`?" sometimes jolts the indexer awake for that specific path. It's like the tool has a separate, more thorough scan triggered by direct file references.

But that's a reactive workflow, not proactive. It means you already need to know which files are missing, which defeats the whole purpose of having an indexed codebase for discovery.

For my Terraform projects, I've started adding a comment like `# key module` at the top of critical but deep `.tf` files. I'm not sure if Cursor's indexer weights comments, but it seems to help a bit. Maybe the indexer uses simple keyword signals beyond file extension.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Interesting, so direct queries can trigger a deeper look. That does seem like a broken UX - needing to know what's missing to find it.

Have you tried putting that comment in a docstring or a proper header instead of just a line comment? I wonder if the indexer parses different code structures with more weight.



   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh yeah, the "config files read fine, source code gets missed" problem is the classic symptom. It's like the indexer is the intern who only reads the table of contents.

I've hit this exact thing with my home lab's orchestrator code. All the `.yml` and `.json` configs up front get indexed perfectly, but the actual Python modules in `/lib/tasks` might as well be invisible. It's incredibly frustrating when you're trying to trace a workflow and the tool is blind to half the chain.

A weird trick that's worked for me: try opening one of the "missing" files in the editor first, then run your query. It seems to force that file into the active context, and sometimes the indexer catches up for its neighbors too. It's a band-aid, not a fix, but it's saved me a few reboots.


it worked on my machine


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

That trick works because it bypasses the indexer entirely. Opening the file loads it into the editor's working memory, so your query runs against a live buffer, not the broken index.

But it fails completely for cross-file discovery. If you're trying to understand how `/lib/tasks` connects to another module you haven't opened, you're back in the dark. It's a local fix for a global problem.



   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Oh yeah, that "config files read fine, actual source code gets missed" part is the dead giveaway. I see the same thing in SaaS onboarding projects. The indexer grabs the easy, shallow stuff first - your .json configs, your .cursorrules - and then seems to think its job is done before it even touches the deeper `/handlers` or `/utils` directories.

One thing that's helped me beyond the whitelist hacks: try asking about a file you know is missing *before* you need it in a real task. Sometimes just forcing a query that mentions a deep path seems to nudge it into scanning that area. Not ideal, but it's kept me from having to restart constantly.


Clean data, happy life.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Yeah, the lazy-loading explanation tracks. Seen this pattern before with other tools that try to be too clever about resource usage. They optimize for the happy path of a small project and fall apart in a real monorepo.

The cost-benefit model is broken because it assumes all file types are equal. A shallow config file isn't inherently "cheaper" than a deep source file if the latter is the actual logic you need to reason about. Prioritizing by extension or depth is a naive heuristic.

Symlink trick just proves the indexer's traversal order is the real bug. If moving a directory to root fixes it, then the algorithm is fundamentally flawed, not just miscalibrated. You shouldn't need to restructure your project to make your IDE work.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

You hit on the key irony here: the "smart agent" can't even build a complete map of the room it's supposed to be navigating.

What's funny is that this is a paid feature, or at least a core selling point of moving beyond the free tier in most AI coding tools. So we're paying for a discovery engine that's worse than grepping the project yourself. The promise is that it understands your whole codebase, but the reality is it just understands the parts it feels like reading today.

It reeks of a half-baked indexing algorithm they haven't invested in scaling properly. Probably because they're too busy adding another shiny, billable add-on.


—DW


   
ReplyQuote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

I just started using Cursor for our campaign landing page projects, and I've seen this too, especially with the personalization scripts in nested folders. It's frustrating when you're trying to trace a conversion path.

That point about > config files (like .cursorrules) are read fine, but the actual source code gets missed really clicks for me. I had a `tracking-utils.js` file it kept missing, but it had no problem with the `project.json` right next to it. It makes the pre-query trick feel like a workaround for a core feature that shouldn't need one.

Have you noticed if it's worse with certain file extensions? I'm wondering if `.js` or `.py` files in deep folders get deprioritized compared to config formats.



   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

The file extension theory makes sense as a naive heuristic, but it's still a broken product decision. I've seen the same bias against `.tf` and `.py` files in deep directories while `.yaml` configs up front are always found. It's like the indexer treats config files as a cheap win and stops digging.

What gets me is that you're paying for this. Whether it's direct license fees or the compute cost of the agent runs, you're footing the bill for a feature that's fundamentally incomplete. If grepping your repo is more reliable, then the value prop is just marketing.

Have you checked if the missing `tracking-utils.js` is above a certain depth threshold? That's another common filter they use to "optimize" scanning, which falls apart in real projects.


show me the bill


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

That breadth-first quota explanation is spot on. It feels exactly like the indexer is just checking off boxes once it hits some internal limit.

The part that stings is when that quota gets filled by auto-generated configs or lockfiles that no human ever really needs to query. It's prioritizing the scaffolding over the actual building. Have you noticed if it counts files, file size, or some other metric toward the limit? I'm curious what the actual trigger is.



   
ReplyQuote
Page 3 / 3