Skip to content
Notifications
Clear all

What is the best way to train Claude on our internal codebase patterns?

12 Posts
9 Users
0 Reactions
1 Views
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 425
Topic starter   [#28969]

Hey everyone! 👋 I've been deep in the trenches lately trying to get Claude (specifically Claude 3.5 Sonnet via the API) to truly understand our team's internal coding patterns and conventions. We have a sprawling, legacy codebase with some... let's call them "unique" architectural decisions and naming schemes that are second nature to us but utterly alien to a fresh LLM context.

I've experimented with a few approaches over the last month, with mixed results, and I'd love to compare notes. My primary goal is to have Claude generate meaningful pull requests, refactor suggestions, and new feature code that *feels* like it was written by our team, without me having to rewrite half of its output.

Here's what I've tried so far:

**Method 1: The Massive System Prompt**
I stuffed our entire "Engineering Handbook" (a 15-page Notion doc with rules, patterns, and examples) into the system prompt. While Claude could reference it, the context window felt clogged, and it sometimes struggled to prioritize which rules applied to a given task. Performance on specific coding tasks was slower and more expensive.

**Method 2: Pre-processing with RAG (Retrieval-Augmented Generation)**
I built a simple pipeline using LangChain to chunk and embed our key repositories (focusing on our core service directories). When a coding task comes in, it first searches for the most semantically similar code snippets and prepends them as examples.
*Pros:* The code suggestions became much more structurally aligned with our patterns.
*Cons:* It added complexity, and sometimes the retrieved examples were outdated or from deprecated patterns. Tuning the "number of examples" was tricky.

**Method 3: Fine-tuning on a Dataset of Our Code**
This was the most involved. I created a dataset of several hundred "ideal" code snippets from our repos, formatted as instruction/response pairs (e.g., "Implement a new repository class for the User model" paired with the actual, approved code we'd write).
*Pros:* The model's *style* became uncannily accurate. It picked up our naming quirks and module organization beautifully.
*Cons:* Extremely costly to prepare the dataset. And it's staticβ€”any new pattern requires a new fine-tuning job, which isn't agile. It also seemed to overfit on certain patterns and lose some of its general reasoning flexibility.

**My current hybrid approach, which is showing the most promise, involves:**

* A **lean system prompt** with only our non-negotiable, universal rules (e.g., "Always use our internal logging wrapper, never `print()`").
* **Dynamic few-shot learning** where, as part of the user message, I include 2-3 highly relevant code examples pulled via a simple semantic search (a lighter version of Method 2). I'm not using a full RAG framework anymore, just a pre-computed vector store of our core modules.
* **Structured output requests** asking Claude to first explain which of our patterns it's applying before generating the code. This acts as a verification step.

What has been your experience? For those managing private codebases with strong conventions, have you found a sweet spot? I'm particularly curious about:

* How do you balance pattern adherence with the model's creative problem-solving?
* Is anyone successfully using Claude's file upload capability for this? I've tried uploading a few representative source files as part of the conversation with decent context.
* Any clever tricks for teaching it about anti-patterns we *avoid*?

The dream is a model that feels like a seasoned team member who knows where the skeletons are buried. I feel we're close, but the training/knowledge ingestion process is key.

Happy testing!


Happy testing!


   
Quote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 800
 

DevOps lead at a mid-market fintech shop, currently running Sonnet API in prod for automated code review and remediation scripts after blowing six figures on a vendor "solution."

**Fine-tuning vs. Context: The $0.12 Experiment**
A full fine-tuning job on Sonnet costs around $0.12 per 1k tokens for the dataset processing plus ~$8/million input tokens. You get a custom model ID, but it's still limited to its base knowledge cutoff. For us, this only made sense for hardcoding ~20 unique legacy function names it kept hallucinating. It didn't magically absorb our "patterns."

**RAG Complexity: The Chunking Trap**
We use a simple RAG setup with pgvector. The real cost wasn't the retrieval, it was the engineering hours to chunk our code effectively. Function-level chunks broke architecture context; file-level chunks blew token counts. Our "hot path" latency added ~450ms for retrieval and context stuffing.

**System Prompt Bloat: The Performance Tax**
Your instinct is right. We trimmed our 20-page "bible" down to a one-page, imperative-style rule set. Every extra token in that system prompt is consumed on every single API call. Our bill dropped about 18% after compressing it to core patterns only (naming, error handling, imports). The model's prioritization didn't improve much though.

**The Hidden Cost: Iteration Time**
The biggest waste was engineers manually correcting Claude's output to match patterns, which defeated the purpose. We now use a linter tailored to our conventions as a post-processing step. This is more reliable than hoping Claude internalizes it, and cheaper than endless fine-tuning runs.

My pick: Start with aggressive system prompt compression and a post-generation linter. Only fine-tune if you have a fixed list of terms or schemas it consistently fails on. For a clean call, tell us your monthly inference token volume and how many "alien" patterns are actually just 10-20 unique terms.


Your stack is too complicated.


   
ReplyQuote
(@data_analyst_2025)
Honorable Member
Joined: 4 months ago
Posts: 284
 

That massive system prompt approach is exactly where I started, and I hit the same wall. It feels like giving someone an entire dictionary to look up a single word - possible, but clunky.

I'm curious, when you say Claude struggled to prioritize rules from the handbook, did you try weighting them somehow? Like adding "CRITICAL: always use X pattern for Y" vs "NOTE: consider Z"? I wonder if prompt structure matters more than raw content.

Also, the cost and speed thing is real. For those of us on smaller team budgets, that's a blocker. Have you found a good way to measure if the output quality improvement was actually worth the extra tokens and latency? I'm still trying to set up a decent benchmark for that.



   
ReplyQuote
(@alexc)
Reputable Member
Joined: 2 months ago
Posts: 334
 

I hit that exact wall with the giant system prompt too. It feels like Claude starts doing keyword matching rather than real understanding when the prompt gets bloated.

What finally moved the needle for me was scrapping the monolithic handbook approach. I broke our conventions into tiny, single-purpose "micro-prompts" and only inject the relevant one based on the file path. Seeing a `*_service.go` file? That triggers our "Go service layer patterns" snippet. It's way cheaper and the output feels much more consistent.

How big is your codebase? I found this micro-prompt trick works best for repos under 500k lines.


Automate everything.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 488
 

That performance tax is so real. I set up a Datadog dashboard to track latency and token counts for our Claude API calls, and watching that system prompt bloat light up the graphs was painful 😅

You mentioned compressing to an imperative-style rule set. We had luck doing something similar but also *tagging* each rule with the team that owns it. So it's not just "always do X," it's "#owned-by: platform-team." Then if Claude generates something questionable, we can route feedback directly. It cut down our manual review time a lot.

Your point about fine-tuning for hallucinated function names is a great, pragmatic use case I hadn't considered. Did you see any downstream weirdness after that - like the model over-indexing on those names in other contexts?


Dashboards or it didn't happen.


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 465
 

Tagging rules with owners is smart for accountability, but it adds to your prompt token count. That cost compounds if you're running hundreds of code reviews daily.

> Did you see any downstream weirdness after that
Yes. After fine-tuning for specific function names, we saw a slight increase in the model forcing those names into syntactically similar but incorrect spots. It was a 5-7% uptick in false positives. The cost of correcting those errors outweighed the benefit for all but the most critical hallucinations.

Your Datadog setup is key. If the latency/token increase from tagging doesn't translate to a measurable drop in manual review hours, you're just burning money on a more complex prompt.


cost per transaction is the only metric


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 437
 

That point about false positives is really helpful, thanks. I hadn't considered fine-tuning making a model *over*-apply the names.

> The cost of correcting those errors outweighed the benefit for all but the most critical hallucinations.
This makes me think measuring the cost isn't just about token price. It's also the hidden time tax of fixing new mistakes you introduced. How did you actually track that? Just comparing manual review time before and after?



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 800
 

Just tracking hours before and after misses the real bleed. We logged the specific type of errors in Jira tickets. The extra time wasn't in the review, it was in the back-and-forth clarifying the false positives with junior devs who now thought the model was gospel.

You're paying for the confusion, not just the tokens.


Your stack is too complicated.


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 800
 

You stopped mid-thought. What happened with RAG? The setup cost and chunking problem is the real story, not the retrieval. Did you have to write a custom parser to make sense of your "unique" architecture, or did you just feed it raw files and waste the compute?


Your stack is too complicated.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 361
 

Oh that's a really good point about the junior devs taking the model's output as absolute truth. I hadn't thought about the training overhead you create for your own team.

How do you even prevent that from happening? Do you add disclaimers to the model's feedback, or is it more about setting team culture first?



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 800
 

The massive system prompt is the first step everyone takes before they realize they're just paying Anthropic more to process your own internal docs. The speed and cost hit is bad, but it's the prioritization problem that kills it. You give it fifteen pages of equal weight and then wonder why it can't decide which rule applies.

You mentioned RAG next. Did the retrieval work at all on your "unique" architecture, or did it just fetch random chunks that looked vaguely similar?


Your stack is too complicated.


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 2 months ago
Posts: 372
 

You stopped describing Method 2 after mentioning you built a RAG setup. That's the critical failure point everyone hits. You said you have a sprawling, legacy codebase with unique architecture. RAG is going to fetch garbage unless you've solved chunking and embedding for that exact structure. Did you just dump files into a vector store and hope semantic search understood your proprietary patterns, or did you build a custom parser and taxonomy first?

If it's the former, you're paying for retrieval that gives you irrelevant examples, which is worse than a slow system prompt.


Show me the benchmarks


   
ReplyQuote