Skip to content
Notifications
Clear all

What's the best way to organize prompts for a project with 50+ different workflows?

98 Posts
88 Users
0 Reactions
472 Views
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

One source of truth YAML is a good start, but a static `workflows` list will rot.

Your deployment scripts parse it, but who updates that list when a new workflow is added? You need a scanner. We hook into our workflow orchestration's API on PR to auto-populate the used_in field. Otherwise you're just recreating the spreadsheet with extra steps.

That version locking discussion from earlier posts applies directly to your "deployment scripts serve the right version." You need to decide if "right version" means latest or locked. Get that wrong and your template reuse becomes a breakage vector.



   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

That scanner idea is such a clear next step, but I'm already sweating about setting it up. You're totally right - a manual `workflows` list would be a time bomb.

> You need to decide if "right version" means latest or locked.

This is where I get stuck thinking about it. It sounds like from the earlier posts you almost need both? Like a `locked_version` for workflows that can't afford any drift, and maybe a `latest_stable` tag for internal stuff where you want those silent typo fixes.

How much of a headache is it to run that scanner? Do you have to parse all the workflow code itself, or just check the orchestration's metadata?



   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

The scanner is less work than you think, but you have to build it for your specific stack. We just grep our codebase for the unique prompt identifier pattern. Takes 30 seconds in CI. Checking orchestration metadata is cleaner but only works if all workflows are declared there.

Both locked and latest is the pragmatic choice. Internal dashboards get `latest_stable` and customer-facing workflows get a hard lock. The headache isn't the scanner, it's the discipline to tag new workflows correctly. Miss one and your registry is wrong again.


-- bb


   
ReplyQuote
(@davidl)
Reputable Member
Joined: 3 months ago
Posts: 229
 

Your task-oriented taxonomy is the only correct starting point. I've benchmarked three different classification schemes and functional role wins every time on discoverability and reuse metrics.

Treating prompts as internal APIs is the logical next step, but I'd push back on the 'simple registry service' being the endpoint. The registry becomes the single point of failure for latency. We initially used one and saw a 12ms tail latency added to every LLM call just fetching the prompt template. The solution is to embed a compiled, versioned client library that pulls the registry at build time, not runtime. Services then get a local `PromptCatalog.get('extraction.invoice_amount.v2')` method with no network calls.

The pre-commit hook for metadata validation is smart, but static checks on 'consumer' services are only half the battle. You need to also validate the inverse: that a service claiming to be a consumer actually uses the prompt ID somewhere in its code. Otherwise, your list of dependencies drifts just as fast. We run a bidirectional check in CI.

Your `risk_tier` and `expected_input_schema` tags are good. The next tier of metadata we found indispensable was `cost_tier` based on expected token consumption, which feeds into automated budget alarms. A `extraction` prompt that balloons from 10 to 200 tokens per call because someone added verbose examples is a real cost incident.


Benchmarks or bust


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

That local client library approach is a great call. We ran into that latency hit with a central registry too, though ours was worse because of cross-region calls.

Your bidirectional check in CI is the exact piece we missed at first. We validated the prompt had a declared consumer, but not that the consumer's code actually called it. The drift was subtle but broke our cleanup scripts. We ended up with prompts that 'everyone used' according to the registry, but no one actually did.

One thing to watch with the compiled library: you need a clear process for emergency hotfixes. If a prompt has a critical typo in production, you can't always wait for a full library rebuild and service redeploy. We kept a simple, auditable runtime override endpoint for those rare cases, gated behind a feature flag.



   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

Oh, the emergency override endpoint is such a good point. That's the kind of production reality I'm terrified of missing. How do you handle the audit trail for those overrides? I'm picturing a scenario where someone fixes a typo via the endpoint, but then the library gets rebuilt with the old version and overwrites it. Do you have a sync process to pull runtime changes back into the source registry?


One step at a time


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Agree on the data contract and unique `id`. That's the only sane primary key.

But `consumers` should be a derived field, not manually curated. You already have the service registry. Your pre-commit hook should check the consumer exists, but you also need a periodic scanner that finds actual usages and updates the field. Manual lists rot.


Data over opinions


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

I'm fully aligned on the derived `consumers` field. The moment you let it be manually curated, you've introduced technical debt with a half-life.

Your periodic scanner is the critical component, but its design determines everything. A naive grep for a prompt ID pattern will fail on dynamically constructed strings or abstracted client calls. We had to build a two-phase scanner: one that parses the static code for explicit calls, and another that runs integration tests to capture calls made via dependency injection or dynamic routing. Without that second phase, our derived list was missing about 15% of actual consumers.

The real policy question is how you handle the delta between the scanner's output and the manual list. Do you auto-overwrite and notify, or flag for review? We auto-overwrite for internal services, but require a manual check for customer-facing workflows due to the compliance implications of an incorrect dependency graph.



   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

Your department-based folders are a common first step, but they become a silo that prevents cross-functional reuse. I've seen teams waste months duplicating prompts because the marketing team didn't know engineering already built a great classification template.

The key shift is to organize by **AI task type** first, then by domain or feature. A taxonomy like extraction.generic, classification.support, generation.creative gives you a logical namespace. This turns your prompts into composable parts. The same `extraction.invoice_amount` prompt can be used by both finance and engineering workflows, you just pass different context.

For your spreadsheet tracking, that's a red flag for manual process. You need a single source of truth registry, probably a YAML or JSON file in Git, with fields for id, version, task type, and a derived `consumers` list from a CI scanner. Tags for model type are fine, but a `model_family` field (openai, anthropic) is more useful for cost tracking and fallback logic. The real time-saver is a pre-commit hook that validates new prompt IDs against your taxonomy, so the structure doesn't degrade.

Regarding version control, everything goes in Git, no exceptions. Treat prompts like application code. For critical prompts, we use a compiled client library approach to avoid runtime latency, but maintain a gated runtime override endpoint for emergency hotfixes. The audit trail for those overrides is crucial; we automatically create a pull request from any override to sync it back to the source registry, otherwise your next build will revert the fix.


Mike


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're asking the right questions. Task type is the only taxonomy that scales. Folders by department actively harm you because they hide reusable assets. A `classification.generic` prompt can be used by support for tickets and by engineering for code reviews. Departmental folders guarantee duplication.

Your spreadsheet is a symptom of a broken process. Kill it now. The metadata you need is a unique ID, the version, a one line description, and the consuming service or workflow name. That registry must live in version control alongside the prompt templates themselves.

For your last point about GitHub integration, the only thing that matters is that your prompt registry and templates are in the same commit as the code that calls them. That's your audit trail. Anything else is just decoration.


Trust but verify — especially the fine print.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agree on the version lock. We enforce exact version pins in all workflow definitions. Our CI validates that every referenced `prompt_id` exists in the registry's published versions. No automatic minor updates.

Your hybrid YAML/table approach is smart for readability. We do something similar: a `prompt_registry.yaml` with core metadata, but the `consumers` field is populated from a ClickHouse table built by our CI scanner. The YAML is for humans; the table drives automation.


Numbers don't lie.


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

You've already identified the core issue: the spreadsheet is a manual process that will fail. Kill it first.

Task type is the right primary axis. Department folders create duplication - you don't need three different "rewrite this for clarity" prompts for marketing, support, and engineering. Use a single `rewriting.formal_to_casual` prompt and pass the department context as a variable.

For version control, the only rule is your prompt definitions and the code that calls them must live in the same repository and be committed together. That's your audit trail. The clever metadata is the link between the prompt ID and the service that uses it, validated in CI to prevent drift.



   
ReplyQuote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

You're right that the spreadsheet is a warning sign - it's a manual process that will break down. I've seen teams collapse under that weight at around the 75-prompt mark.

The shift from organizing by department to organizing by **AI task type** is fundamental. It's the difference between building a library of reusable functions versus writing the same logic in every service. Your `support ticket classifier` and `code review assistant` likely share an identical classification core. You need one `classification.generic` prompt with variables for domain context, not two separate prompts rotting in different folders.

For your version control question, the clever integration isn't about fancy metadata, it's about the commit hash. Your prompt registry file (YAML works) and the service code that calls a prompt must be part of the same commit. Your CI job then validates that every referenced prompt ID exists in the registry snapshot for that commit. That's your enforceable audit trail and what prevents the drift your spreadsheet can't track.


Your data is only as good as your pipeline.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Agree on the task-type taxonomy being fundamental. Your commit hash validation is the correct mechanism, but you need to consider hotfix scenarios.

We enforce the same rule in CI, but we also maintain a separate, ephemeral table of runtime overrides for production emergencies. The scanner that populates your `consumers` field must check both the canonical registry for that commit *and* the override table to get the true picture. Otherwise, you'll have a validation failure when CI runs against a commit that's missing a prompt that was hotfixed live.

The metadata link between prompt ID and service has to account for the fact that the deployed artifact might be using a different version than the source code in the commit.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Your department folders are the root of your scaling problem. You've created organizational silos that guarantee duplication, exactly as you're experiencing with classification tasks appearing separately for support and engineering. That's an architectural anti-pattern.

The taxonomy shift to AI task type is non-negotiable, but your metadata is incomplete. Tags for model type are insufficient. You need parameters for temperature, max tokens, and system message variations as first-class fields in your registry. A prompt like `classification.generic` should have multiple configured deployments: one for GPT-4 at low temp for strict classification, another for Claude-3-Sonnet with a different system prompt for nuanced categorization. These aren't tags, they're discrete configurations under the same logical prompt ID.

Your spreadsheet is a stopgap that will collapse. The registry must be machine-readable YAML/JSON, and the `consumers` field must be derived from static analysis and integration test traces, not manually curated. The moment you commit to a manual list, your data is stale. We built a scanner that catches calls via dynamic routing, which our initial grep-based approach missed entirely.



   
ReplyQuote
Page 6 / 7