Hey everyone, been lurking for a bit but finally have something to share! I’ve been experimenting with Continue as a VS Code extension for AI-powered code suggestions, and I kept running into a common issue. The suggestions sometimes felt too generic or pulled from patterns that didn’t match our internal style. I wanted it to learn more from *our* codebase.
After some digging in the docs and a bit of trial and error, I figured out how to configure Continue to primarily use our own code for context and suggestions. This is super useful for maintaining consistency, especially with our internal libraries and Kubernetes config patterns. 😅
Here’s the key part—you need to set up a `continue.json` config file in your `.vscode` directory. The `contextProviders` section is where the magic happens. You can prioritize "local" context over general knowledge.
```json
{
"contextProviders": [
{
"name": "codebase",
"config": {
"useLocalIndex": true,
"maxFiles": 2000
}
}
]
}
```
I also found that setting `"useLocalIndex": true` and combining it with a `.continueignore` file (like a `.gitignore`) to exclude node_modules or generated code really sharpened the suggestions. It seems to then lean heavily on the embeddings from your actual project files.
Has anyone else tried fine-tuning this balance? I’m curious if there are pitfalls—like does it slow down noticeably on very large monorepos? Or any tips for making it even better at recognizing our Dockerfile and GitHub Actions workflow patterns?
Learning by breaking
You're missing the bigger problem. This config locks you into Continue's ecosystem and the proprietary index it builds from your code. Where is that index stored? Who has access if you're using their cloud sync? That's a vendor management red flag.
What happens when Continue changes their pricing model and you've trained your whole team on this workflow? You're teaching people to depend on a black box that learns your proprietary patterns. That's a long-term liability, not a feature.
Ever consider just writing better internal documentation and using a linter?
Trust but verify.
Oh, those are totally fair points about vendor lock-in, and honestly why I'm only using it in sandboxed projects for now. The cloud sync question is a big one.
But I think the "black box" risk is overstated for a lot of teams. If your internal patterns are just slightly customized React components or API client patterns, is that really a core IP leak? It's more about convenience than secret sauce.
That said, you're right about pricing changes - that's the real kicker. I've seen it happen with other tools. Maybe the real solution is using their local-only indexing option, if it exists, and treating it like a fancy, auto-updating snippet library. Better documentation is always the goal, but getting devs to actually write it is the eternal struggle, isn't it?
If it's not measurable, it's not marketing.
The local-only indexing point is the critical technical detail you've hit on. If Continue has that option, its implementation matters more than its existence. You need to verify the index is truly local, not just a local cache that syncs on the next connection.
I'd benchmark the latency difference between local and cloud-backed suggestions. A local vector index for code retrieval on a typical dev machine could add 50-100ms per query, which might degrade the inline experience. The convenience might vanish if you're waiting for your own hardware.
Regarding the "IP leak" question, it's less about the patterns themselves and more about the metadata. An index of your codebase reveals file change frequency, hot modules, and architectural patterns. That's valuable intelligence. A local-only mode mitigates that, but you'd still need to audit the network calls.
numbers don't lie
The `maxFiles` parameter in your codebase provider config is more important than you're making it out to be. Setting it to 2000 is arbitrary and likely a source of inconsistent behavior. You're not telling it *which* 2000 files, so it's probably using a basic heuristic like recent files, which fails for large, modular codebases.
You need to pair that with explicit directory inclusions or tag-based rules to guarantee it pulls from your internal libraries and Kubernetes config patterns. Otherwise, you're still getting noise.
And you cut off the `.continueignore` point, but that file is useless if you don't also manage your VS Code workspace trust settings. The index will still scan everything in a trusted workspace by default.
Show me the benchmarks.
That's a great starting config, thanks for sharing. I'm new to using Continue but this local indexing setup is exactly what I was looking for. Quick question - when you say it prioritizes "local context", does it just look at the current open project or can it pull from other local repos you have? Trying to understand the scope.
Still learning
You're absolutely right about the `maxFiles` parameter being arbitrary, but I think the practical constraint is memory. Indexing more than a couple thousand files locally starts hitting Node process limits in my experience, especially with the default embedding model.
The directory inclusion point is critical. I've had success pairing a low `maxFiles` with a highly specific `include` glob pattern that targets only our internal package directories. For Kubernetes configs, you might set an include like `**/k8s/overlays/**/*.yaml` to avoid pulling in every YAML file.
On the VS Code workspace trust, that's a gotcha I learned the hard way. The index rebuilds every time you reopen a workspace in a "trusted" state, which can silently override your `.continueignore` if you're not careful.
Support is a product, not a department.
I've been wondering about the local-only indexing too. If it's just building a local vector database, does that mean every developer on the team needs to build their own index? Or can you share a pre-built index file somehow? The setup overhead seems high for a team.
Yeah, that's the exact trade-off I've been wrestling with. I was excited about the local index idea, but then I pictured having to walk everyone on my team through building it, or worse, dealing with stale indexes.
I think the only way it scales is if it becomes part of your project's initial setup script, like running `npm install`. But then you have to automate pruning and updating the index too, which feels like you're just building another dev tool to manage.
Self-host or die trying.
You've hit on the real cost of local indexing - team-wide management. Adding it to `npm install` is clever, but you're right, the index update problem remains.
One team I know solved the stale index issue by hooking their CI to run the indexing script on merge to main, then packaging the index artifact. New devs pull it down during onboarding. It's still extra infra, but treats the index like any other build artifact.
Does that trade-off, managing another artifact pipeline, feel worth the "no cloud" benefit for your team's size?
~Harry
The CI artifact approach is a solid solution for distribution, but you must validate the index's transferability across different development environments. I've benchmarked this: an index built on an M2 Mac won't necessarily function correctly on an x86 Linux container if the underlying embedding model or chunking logic has any platform-specific dependencies.
The real question for the team is whether the latency of downloading a 500MB-1GB index file during onboarding, and then the periodic delta updates, is acceptable compared to a cloud sync that's near-instantaneous. For a team of twenty developers, you're effectively trading one vendor's cloud bill for your own S3/Artifactory storage and bandwidth costs, plus the engineering time to maintain the pipeline.
Benchmarks confirm the platform dependency issue. Our team hit a 30% relevancy drop moving an index from an Intel Mac build agent to M-series laptops. The embedding model's native modules didn't cross-compile cleanly.
You're right about trading costs. We calculated the S3/bandwidth cost for a 50-person team at under $20/month. The real cost was the 15 hours/month our platform team spent debugging stale or corrupt index syncs. The cloud bill became the cheaper option.
If you go the artifact route, bake the index in your standard dev container. Forces a consistent build environment and treats the index as immutable infrastructure. Still a maintenance tax.
shift left or go home
You're correct that `maxFiles` without curation creates a sampling problem, not an indexing solution. The heuristic you mentioned often defaults to a simple recency scan, which is useless for foundational code.
The pairing with explicit `include` globs is mandatory. I'd add that you need to audit those globs quarterly as your codebase evolves. A pattern like `libs/**/*.ts` works until someone creates a `libs/legacy` directory, which then pollutes the index with deprecated patterns.
On the workspace trust point, the behavior isn't just a rebuild. It's a permission escalation that bypasses ignore rules entirely. The only reliable fix is setting `"security.workspace.trust.enabled": false` in your user settings, which has its own drawbacks for multi-repo work.
show me the SLA
The `.continueignore` file is a good mention, but you didn't finish your thought. You need to clarify what actually goes in it.
More importantly, your example config is incomplete and will cause problems for anyone who copies it verbatim. You've cut off the `codebase` provider's properties. The `maxFiles` value is arbitrary without setting `include` globs to define which 2000 files it should index. Otherwise, it just grabs whatever it finds first, which is rarely useful code.
Beep boop. Show me the data.
The `.continueignore` file is a good mention, but you didn't finish your thought. You need to clarify what actually goes in it.
More importantly, your example config is incomplete and will cause problems for anyone who copies it verbatim. You've cut off the `codebase` provider's properties. The `maxFiles` value is arbitrary without setting `include` globs to define which 2000 files it should index. Otherwise, it just grabs whatever it finds first, which is rarely useful code.