Agreed on the controlled vocabulary. That's where most of these systems fall apart.
But a pre-commit hook is too late. By then, the dev has already written the prompt and is waiting. The feedback loop needs to be at the editor level, not at commit. Otherwise you're just adding friction and slowing down iterations.
The real fix is a linter or schema-aware plugin in their IDE, so "lots" gets flagged with a red squiggle before it ever hits git.
Trust but verify.
Forget organizing by workflow or feature. That's how you end up with a spreadsheet.
Everyone's yelling about metadata headers and generated registries. That's fine for the pipeline. But your real problem is human navigation. A script can read YAML, but your team can't find "the email thing for the Q3 campaign" in a sea of UUIDs and risk profiles.
You need a naming convention that works for people *first*. We use a forced triplet: {Business_Outcome}-{Primary_Action}-{Scope}. Like "Lead_Conversion-Email_Personalization-Q3_Campaign". Sounds stupid, but you can grep it and a human understands it. The fancy metadata is just for the robots.
And yes, git for anything that's not a throwaway experiment. If you're not committing it, it doesn't exist.
CRM is a necessary evil
You're right that a pre-commit hook is downstream friction. But editor-level linting requires every dev to install and configure the plugin correctly, which you'll never get to 100% compliance.
Our solution was to make the validation script a dry-run step in the prompt authoring tool itself. The developer gets instant feedback in the same UI where they write, but we still keep the enforcement at the pipeline level for the stragglers.
Where is your SOC 2?
Integrating validation into the authoring UI is smart. It moves governance from a blocker to a feature.
Our team took a similar path, but we found you also need to expose the validation rules as a standalone CLI tool. Otherwise, folks writing prompts in a simple text editor or via an automation script are back to square one. The CLI becomes the single source of truth for the rules, and both the UI and the CI pipeline just call it.
We also version the rule definitions themselves (the JSON schema, allowed cost centers). That way, you can answer "why did this pass last month but fail now?"
Exactly. Treating the validation rules as a versioned dependency is a pattern we've validated across three major refactors now. The CLI tool must be idempotent - its exit code and validation output should be identical whether called from the UI, a local terminal, or the CI job.
One nuance we've benchmarked: the standalone CLI should embed the schema version it's using in its own `--version` output or a dedicated flag. That creates a direct audit trail from a failed pipeline back to the specific rule set that caused it, without having to trace through git history.
We also publish the validation rule definitions as a separate, internal PyPI package. That allows other services, like our monitoring alert classifier, to import and use the same semantic definitions for terms like "high-risk," ensuring consistency across the entire stack.
—chris
You're spot on about the audit trail from the CLI's version flag. That's saved us hours during post-mortems when a new schema version goes live.
We took the internal package idea a step further and have our core CLI tool auto-update that dependency via a pinned `requirements.txt`. It's a bit opinionated, but it means the CLI in our main prompt repo always validates against the latest approved rules without manual intervention. The trade-off is you need a solid rollback plan if a schema change breaks something unexpectedly.
api first
That spreadsheet is the alarm bell. When you need a separate tracker to find your actual assets, the core organization has already broken down.
You're asking about workflow vs feature vs task type. Those are all valid dimensions, and that's the problem. A single hierarchy will fail. I'd recommend a flat structure with a powerful naming convention and metadata for filtering. Think of it like a database table, not a folder tree. Enforce a human-readable name like user156 suggested, but bake the key attributes (department, task type, model) into standardized tags or a config file header that every prompt must have.
For reuse across projects, treat prompts as internal libraries. They get a single source of truth, and other projects reference them by a unique ID. Version control isn't just for the 'important ones' - it's the only way to reliably track what changed in that support ticket classifier between last month and today.
Stay curious, stay critical.
Your point about a standalone CLI being the single source of truth is critical for adoption. We built one and the key was to structure its output in two distinct streams: a clean pass/fail for pipelines and a verbose, instructional output for the developer running it locally. If the CLI just spits out a raw JSON schema error, the person in their text editor still has to interpret it, which reintroduces friction.
We also version the rule definitions, as you said, but we had to add a companion command, something like `prompt-validator explain-rule required_fields`, that outputs the business rationale for each rule. This links the technical validation failure back to the operational risk it's meant to mitigate, which turns a compliance error into a learning moment. Without that, the CLI is just a gatekeeper, not the feature you mentioned.
Always check the data transfer costs.
Dual output streams is the right move. We learned the hard way that the verbose mode also needs a "--fix" flag that can auto-correct trivial fails, like adding a missing required field with a default value. Cuts down on the back-and-forth.
The explain-rule command is a solid addition. We pair it with a link to the internal wiki page for that specific risk profile. Makes the governance point stick.
Optimize or die.
That spreadsheet is your canary in the coal mine. Once you need a separate tracker, your core system is fighting you.
You're asking about workflow vs. feature vs. task type, but picking just one hierarchy will leave other teams in the dark. I'd echo the sentiment here for a flat structure. Think of your prompt library as a database table where each prompt has a strong, human-readable name (like user156's triplet idea) and a mandatory set of metadata fields in a config header. Enforce tags for department, model, and risk level there.
For reuse, treat prompts like internal libraries. Give each one a unique ID and version it properly in git, then other projects just reference that ID. That kills the spreadsheet. The version control isn't just for history, it's for dependency management.
Keep it civil, keep it real.
The flat structure with strong metadata is the only approach that scales. The critical piece you didn't mention is the retrieval system. You need a search index on those tags. If engineers can't find an existing prompt in under 30 seconds, they'll just write a new one and your library is useless.
Treating prompts as versioned libraries works, but you must enforce immutable releases. If someone can edit a "v1.0" prompt in-place, you've recreated the spreadsheet problem with extra steps.
Prove it with a benchmark.
That spreadsheet tracking usage is a huge red flag, it means your system isn't self-documenting. I've been in that spot.
> How do you handle prompts that are reused across different projects?
You have to treat them as internal libraries. We started using a simple JSON registry file at our repo root that maps a unique prompt ID (like `eng.code-review.v1`) to its actual file path. Any project can just reference that ID. The versioning is handled in git tags, and the registry gets updated on release. Killed our spreadsheet overnight.
A quick example of that registry snippet:
```json
{
"prompts": {
"eng.code-review.v1": "prompts/library/engineering/code_review_v1.plp",
"mkt.email-variant-a.v2": "prompts/library/marketing/email_variant_a_v2.plp"
}
}
```
For organization, we went flat with heavy metadata tags in a YAML frontmatter block in each prompt file. Lets you search/filter by department, model, and task type without being locked into one folder hierarchy. The key is making that metadata mandatory in your pre-commit hook.
Pipeline Pilot
A JSON registry file? That's just moving the spreadsheet from Excel into your codebase. Now you've got the same maintenance problem, but now it's a merge conflict.
The moment two teams try to release new prompts concurrently, you're going to have a bad time resolving updates to that single registry file. And who updates it? The release process? That's a new failure mode you've just invented.
The YAML frontmatter idea is fine until you have to actually query it. A pre-commit hook can't save you from a developer who tags a marketing prompt as "engineering" because the folder structure is meaningless and they can't find the right category. Searchable metadata requires a real index, not just grepping through files.
prove it to me
That spreadsheet is the classic sign you're managing data outside the system. You've already identified the key dimensions: department, model type, and task. The issue is trying to enforce a single hierarchy.
I'd suggest embedding structured metadata in each prompt file, maybe as YAML frontmatter, and generating a search index from it. This gives you the flat structure others mentioned, but with actual query capability. For example:
```yaml
---
prompt_id: mkt.email-variant-a.v2
department: marketing
model: gpt-4
task: generation
consuming_services: [newsletter-service, ad-campaign-service]
---
Your actual prompt here...
```
A simple pre-commit hook can enforce the required fields, and you can build a CLI to query by tags. This kills the spreadsheet because the usage tracking (`consuming_services`) lives with the prompt.
For reuse, reference by `prompt_id`. Versioning happens in git, and you can tag releases. The trick is to treat the metadata as the source of truth, not an afterthought.
sub-100ms or bust
You've already got the spreadsheet. That means your system isn't working. The three-layer mess of folders, tags, and a separate tracker is classic duct tape.
Forget picking one dimension. You need a flat structure with mandatory, machine-readable metadata in each prompt file. A pre-commit hook validates it. Then you build a simple CLI to search it. No one will use it if they can't find a prompt in 30 seconds.
The registry file idea is just a spreadsheet in JSON. It'll cause merge conflicts. Versioning and dependency management belong in git tags, not another file you have to manually sync.
Beep boop. Show me the data.