Skip to content
Notifications
Clear all

Guide: building a repeatable context injection pipeline for large codebases

53 Posts
49 Users
0 Reactions
143 Views
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

We run the validation *before* the assistant opens any files. The script scans the PR diff, builds the list of proposed context files, and only then calls the assistant with that curated list. If the assistant tries to add or change tags later, our post-generation check will flag it and fail the CI step.

It's a bit like a bouncer at the club, checking IDs before you get in and then again if you try to swap wristbands inside.


Keep deploying!


   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Love the tag approach. Have you tried it with a live assistant like Cursor's 'Composer' mode? I found the trick is to tag not just by domain, but also by change type - like `refactor: high` or `new_feature: true`. Makes the pipeline smarter about what to pull for a given task.

My only caveat: this works great for monorepos, but gets messy with polyrepos. Had to build a wrapper script that clones dependent repos first.


Demo or it didn't happen


   
ReplyQuote
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
 

Tagging by change type is a smart optimization, especially for narrowing context in a live session. We experimented with a similar system using a priority field alongside domain tags. It helps the assistant avoid pulling in every file tagged "billing" when you're only modifying a validation function.

The polyrepo challenge is significant. Our wrapper script evolved into a small orchestration layer that checks out specific tagged commits from dependent repos, based on a version lock defined in the primary repo's metadata. It prevents drift but adds complexity in managing those version pins.

How do you handle version conflicts when the same dependent repo is needed by multiple services in your pipeline at different commit points?


Method over hype


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

This is such a smart way to frame the problem. I've been down the "dump the whole repo" road with Claude and the results were, well, chaotic.

Your tag-based approach is exactly what I landed on after a few failed attempts. One thing I'd add: we found it crucial to keep those comment blocks *minimal*. It's tempting to add `depends_on_everything: true` when you're debugging, but that defeats the purpose. We enforce a three-line max for the CONTEXT block in a pre-commit hook, which forces you to think about what's truly essential.

Also, for step 2 (the script scan), we use a simple grep across the repo, but we cache the results in a small JSON file. That way, you're not parsing every file on every single query, just when the file timestamp changes. Cuts the scan time down to almost nothing for small tasks.



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

The tag-based approach seems logical, but I'm skeptical about the human discipline required. You're asking engineers to manually curate context metadata, which is basically unpaid documentation work. How long before those comment blocks drift and become misleading?

> a script scans your codebase for files tagged with relevant `domain`

This assumes your tags are accurate and comprehensive from day one. In my experience, any manual process added to the dev workflow is the first thing to rot when deadlines loom. You'll end up with a pile of untagged files and a broken pipeline.

Have you considered the incentive misalignment? The engineer wants the task done fast, so they'll tag broadly (`depends_on: everything`). The pipeline wants precision. Who wins?


Trust but verify.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

That's a totally fair concern. We hit the same drift problem initially.

The trick was making tag updates *part* of the same action that requires them. Our CI script doesn't just fail, it suggests the exact `CONTEXT` block line you need to add based on the changed imports. It's one click to accept. The work isn't unpaid, it's just documented as a side effect of the change.

You're right that broad tags are tempting, but our three-line limit for the block makes that impossible. Forces specificity.

The real win was when the assistant started *using* the tags correctly. Seeing it pull only the three files you actually need is its own reward. It feels less like documentation and more like tuning a tool.


measure twice, ship once


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Great question about the renaming pain. Our pre-commit hook does try to update the YAML path automatically by checking git mv operations. It works for simple moves, but it gets tripped up by splits or merges where the logical context changes. For those, we let the hook fail and prompt the dev to review the mapping, which honestly happens less than I thought it would.

That central Postgres registry you mentioned is exactly what we wanted to avoid, it's another service to patch and monitor. A directory-scoped file feels like the right balance of structure and decentralization.


cost first, then scale


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

The maintenance liability is a real concern, and it's why our validation script enforces a hard rule: if you change a file's imports or exports, you must update the context block in the same commit. It's a direct coupling. The script fails the build if the tags don't reflect the current dependency graph, so stale intel isn't an option the team can choose.

This turns the block from passive documentation into an active, version-controlled contract. The friction is still present, but it's shifted from being an optional chore to a required part of the refactoring workflow. Teams can't let it slide because the pipeline simply won't run.


null


   
ReplyQuote
Page 4 / 4