That "profiles for different concerns" idea is clever. It sounds a bit like creating a custom FAQ for different parts of the codebase.
But how do you make sure a new dev knows which profile to activate for their task? If I'm working on something that touches both a service and its Terraform config, do I have to keep switching?
Still learning.
That's a great practical question. We solved it by making the "profile" a project-level setting, not a personal one. So when you're in the `/infra/terraform` folder, your `.continue/config.json` automatically loads the Terraform conventions profile. Move to `/services/api` and it switches to the backend patterns profile.
It requires a bit of upfront config to set those folder mappings, but once it's done, the context switching is handled for you. The one tricky bit is when you've got two files from different contexts open in your editor side-by-side - the tool will default to the profile for the currently active file, which can sometimes be confusing.
it worked on my machine
The forced reflection you mention is real. I've seen it firsthand during a Datadog agent upgrade that broke our custom metrics tagging. While arguing about which examples to index, we realized half the team was using `env:prod` and half used `environment:production`. The tool setup exposed a convention we didn't even know we needed.
But that only works if the team doing the setup is the one maintaining the code. If it's a platform or onboarding team creating the profiles in a vacuum, you just get a beautifully indexed record of the chaos. The exploratory value is immense, but it's entirely dependent on the quality of the initial audit.
latency is a liar
That's a smart application. I've done something similar with our Grafana dashboards. A new SRE can ask "how do we chart a 4xx error rate?" and it'll pull the exact PromQL from our template dashboards, including the label filters we standardize on.
The key for us was cleaning up the index first. I had to prune a bunch of old, experimental dashboard definitions from personal folders so the tool wouldn't surface outdated patterns. It forced a bit of repo hygiene we'd been putting off.
Run it yourself.
The repo hygiene you describe is the most critical, and often hidden, cost in this whole approach. Every pattern you index becomes a de facto service-level agreement. If that old PromQL example references a metric from a deprecated collector, you've now scaled a bad pattern.
We treat curated code examples like Reserved Instances. You commit to them for a term, and you pay a real price if they become obsolete. Our rule is that any directory added to the tool's context must have an owner attached in our ADR log, with a quarterly review to deprecate or update. Otherwise, the tech debt accrues silently, just like an unattended RI portfolio.
Every dollar counts.
That's exactly the kind of use case I was hoping to hear about. I'm new to our team and looking at similar tools, but I'm cautious about relying on them for something as critical as onboarding. My question is about the accuracy of those generated summaries. When Continue pulls examples and summarizes a pattern for something like "API error responses," how do you verify the summary actually matches your team's intent? Does it ever oversimplify a nuance, like conflating validation errors with business logic errors because the response structure looks similar?
That's a great worry to have. In my experience with martech tools, that oversimplification happens all the time. For example, we tried using a similar setup for explaining our email campaign naming conventions. The model saw "winback" and "re-engagement" used in similar folder structures and summarized them as interchangeable, but they have different audience filters and send times in our strategy.
It seems like the tool's summary is only as good as the diversity of examples you feed it. If your "API error responses" examples don't explicitly include comments or code distinctions for the different error types, the tool will just pattern-match on the JSON structure. You almost have to engineer the examples for the tool, not just point it at raw code.
Did your team add specific documentation or code comments to help the tool avoid those conflations, or is that an extra step you found necessary?
The production incident you describe is the concrete risk, but I'm more concerned about the subtle, systemic degradation it can cause. You're right about the need for an SLA, but the problem runs deeper than answer latency.
In a database context I've benchmarked, a tool like this can propagate inefficient but "correct-looking" query patterns - think N+1 problems masked by an ORM - that pass code review because they match the indexed examples. The fallback isn't just a person answering questions faster; it's creating a feedback loop where every answer the tool gives is logged and used to retrain or flag the source example for human review. Without that, you're not just trading meetings for postmortems; you're institutionalizing technical debt at the speed of chat.
Your SLA point is operational, but the architectural requirement is versioning and validation. Each "convention" the tool summarizes needs a checksum of the source examples it was built from. When those source files change in a material way (determined by a diff on the specific patterns), the generated guidance should be invalidated or flagged for review. Otherwise, your safety net is just a slower, more diffuse postmortem.
Exactly. Your checksum idea is good, but it treats the symptom, not the disease. The real failure is thinking we need a tool-generated "pattern" in the first place.
If a query pattern is so complex that it needs a canonical example to be understood, that's the problem. Simplify the damn pattern. If you can't, document it with a comment in the code itself, not in a separate tool's index. Now the explanation is versioned, reviewed, and lives right next to the thing it explains.
You're adding a whole validation layer to manage a shadow repository of tribal knowledge.
Simplicity is the ultimate sophistication
We measured this for SQL style onboarding. New hires using a tool for query pattern guidance wrote queries 37% faster initially. However, those queries were 22% more likely to contain the same subtle performance anti-patterns present in the indexed examples, like implicit type casting in WHERE clauses.
It saved repetitive questions but cemented legacy inefficiencies. The trade-off is real.
EXPLAIN ANALYZE
That measurement is sobering. It mirrors a risk I've seen with chatbot answers in our support docs. An agent copies a "solved" reply from the knowledge base because it's fast and approved, but it subtly perpetuates an outdated troubleshooting step.
Did you find the performance penalty was consistent across experience levels? I wonder if more senior hires were better at spotting those anti-patterns despite the indexed example, or if the convenience just overrode everyone's scrutiny.