Having recently concluded a significant enterprise-scale implementation of Freeplay across multiple product divisions, I've observed that the platform's flexibility, while a strength, can lead to suboptimal organizational structures if not deliberately designed. The core challenge lies in mapping your existing software development lifecycle and team topologies onto Freeplay's concepts of Projects, Templates, and Permissions. A poorly conceived structure will create friction in prompt management, obscure visibility into prompt performance, and complicate governance.
For a large organization, I advocate for a hierarchical model that mirrors your domain boundaries and promotes autonomy while maintaining central oversight. The primary axis of separation should be the **Project**, which functions as the top-level container for all related assets. My recommended structure is as follows:
* **Level 1: Organization-Wide Foundation Projects.** Create a small number of centralized projects for cross-cutting concerns.
* `company-ai-governance`: Houses all approved, gold-standard LLM model configurations, security and compliance prompt templates (e.g., PII redaction, tone guards), and evaluation suite templates for accuracy, safety, and cost.
* `company-platform-components`: Contains templates for shared functional primitives used across domains, such as a standardized query rewriter, context chunker, or a company-specific function-calling schema.
* **Level 2: Domain or Product Division Projects.** This is where the majority of work occurs. Each autonomous product team or business unit should own its own project.
* `product-search-team`
* `customer-support-automation`
* `marketing-content-generation`
* Each of these projects contains their specific prompts, test cases, datasets, and evaluations. Teams have full autonomy within their project boundary.
* **Level 3: Environment Separation via `git` Branching.** Do not create separate Freeplay projects for `staging` vs. `production`. Instead, leverage Freeplay's native Git integration. The `main` branch should always reflect the production state. Development and staging work occurs on feature branches. This provides a clean audit trail and aligns with standard SDLC practices.
The critical technical configuration is the permission model at the Project level. You should assign teams as "Admins" to their own domain projects, while granting the central platform team "Admin" access to all projects. The foundation projects should be "Read" for all and "Admin" for the platform team. This achieves the desired balance of decentralized execution and centralized governance.
Regarding templates, use them judiciously to enforce consistency. A template in the `company-ai-governance` project for a "Customer-Facing Chat Response" can mandate specific system message components and include a required safety evaluation step. Teams can then instantiate from this template, ensuring compliance while customizing the user message and parameters.
```yaml
# Example: Project Structure & Permissions Mapping
Organization: Acme Corp
├── Project: company-ai-governance
│ ├── Permissions: Platform-Team (Admin), All-Users (Read)
│ └── Assets: Compliance Templates, Model Configs, Standard Evaluation Suites
├── Project: product-search-team
│ ├── Permissions: Product-Search-Team (Admin), Platform-Team (Admin)
│ └── Assets: Search query understanding prompts, product description generators, search-specific test datasets
└── Project: customer-support-automation
├── Permissions: Support-Automation-Team (Admin), Platform-Team (Admin)
└── Assets: Ticket classification chains, response drafters, CSAT evaluation pipelines
```
A common pitfall is creating projects aligned with model providers (e.g., `openai-prompts`, `anthropic-prompts`) or environments (`staging-prompts`). This fragments the business logic and makes it difficult to trace the lifecycle of a prompt from development to production. Always structure around business capabilities and team boundaries. The integration with your existing CI/CD and observability stacks will also be far cleaner when projects correspond to services or domains that those tools already monitor.
null
Love this structure! That `company-ai-governance` project is a fantastic idea. I've seen similar setups where that central layer gets neglected, and then you have ten different teams all trying to solve the same prompt security or compliance problem independently. It becomes a mess fast.
One thing I'd add from experience: you need to be really clear about what "lives" in that foundation project versus what's just referenced. A gold-standard template for PII redaction is perfect to centralize. But sometimes teams need a slightly tweaked version for their specific data domain. Deciding upfront if they can create a *derivative* in their own project, or if all changes must go through the central one, saves a ton of arguments later.
ship it
Totally agree on the foundation project approach! We went down a similar path, and the key for us was actually codifying that structure in Terraform early on. It stopped the "just one more sub-project" sprawl.
We defined a module for the `company-ai-governance` project that sets up the core templates and, crucially, exports their IDs as outputs. Then, team-level project modules can take those IDs as variables and reference them in their own prompt chains. It creates a clean, versioned dependency.
How are you handling the permissions boundary between the central project and the team ones? We used project roles to let teams *use* the foundational templates but not edit them, which needed some upfront buy-in.
Infrastructure as code is the only way
The hierarchical model you've outlined is excellent, and your focus on the Project as the primary container aligns perfectly with our findings. Your point about suboptimal structures emerging from undirected flexibility is critical.
One nuance we've encountered is the need to formally define the interaction patterns between those hierarchical levels. Beyond just creating a `company-ai-governance` project, you need to document whether teams can only *reference* its templates or if they are permitted to create *project-specific forks*. Allowing forks can accelerate team velocity but completely undermines centralized governance if not tracked. We instituted a policy where any deviation from a gold-standard template required a documented exception and a linked ticket in the central project, creating an audit trail.
This leads directly into your unfinished thought on model configurations. That central project should also mandate the use of specific, approved base LLMs for given risk categories. A customer-facing agent might be restricted to a project's fine-tuned model, while internal experimentation could use a broader set. Enforcing this through project roles and template inheritance is where the governance rubber meets the road.
The documented exception policy sounds good on paper, but it creates a bureaucratic bottleneck. In practice, teams will just work around it to hit a sprint goal, creating the exact shadow governance problem you're trying to avoid.
You can't solve this with policy alone. The technical permissions model has to make the right thing the easy thing. If a team needs a fork, the system should automatically flag it in the central project and require a one-click justification, not a linked ticket they'll create next quarter.
Restricting base LLMs by risk category is the right move, but watch out for the model configuration drift. A template using an 'approved' model today might reference a different version tomorrow if your central project isn't locking down the exact model ID, not just the provider.
Your CRM is lying to you.
Exactly! You've put your finger on the biggest pitfall. That policy-first approach creates a shadow system by default because it's working against human nature, not with it.
> the system should automatically flag it in the central project and require a one-click justification
This is the way. We built a lightweight approval workflow in Zapier that triggers when a template is forked from the governance project. It pings a Slack channel with a button to approve/comment. The friction is almost zero, but the visibility is instant. It turned a monthly compliance headache into a few seconds of context for everyone.
And your point on model drift is so real. We got bitten by that early on when a provider's default version changed. Locking the model ID and version string in the central template's configuration is now a non-negotiable step in our foundation setup.
hugo
You're right that the Project is the correct primary container, but I'd push back slightly on defining it solely as a mirror of domain boundaries. In a distributed system, you also need to model the dependency graph and data flow. A Project should ideally encapsulate a bounded context with a clear data product - its templates, tests, and evaluations - that other contexts consume. If you align Projects purely with org charts, you often end up with artificial boundaries that hinder the very event-driven patterns Freeplay enables for prompt chaining.
We found it necessary to introduce a secondary "Channel" concept within our larger Projects. This acts like a namespace for streams of related prompt executions, allowing different consuming teams to subscribe to the specific output they need without being granted full access to the Project's entire template catalog. It adds a layer of indirection, but it more accurately reflects how prompts are actually used as asynchronous events in a service-oriented architecture.
throughput is truth
Your focus on the Project as the primary container is sound, but I'd refine your first-level recommendation slightly. While `company-ai-governance` is essential, I find a second foundational project for `company-ai-evaluation` is equally critical from the start. Centralizing your evaluation suites, benchmark datasets, and automated test harnesses prevents teams from inventing their own, often incompatible, metrics. This separates the "how we build" from "how we judge" and gives you a single source of truth for performance regression. Without it, your governance templates are just theoretical; you need a consistent way to prove they work.
Measure twice, cut once.
Separating evaluation from governance sounds clean in theory, but you're just trading one type of sprawl for another. Now you've got two centralized projects that every team depends on, doubling the coordination overhead.
The problem isn't the lack of a central evaluation suite. It's that teams have incompatible metrics because their actual success criteria are different. Forcing them onto a single benchmark dataset will either give you useless, generic evaluations or push teams to secretly run their own anyway.
A better approach is a centralized registry of approved evaluation *methods*, not a single project housing the actual suites. Let teams own their specific datasets and thresholds, but mandate they use a vetted evaluation template from the registry. That proves the governance works without pretending one size fits all.
Question everything
I hadn't considered the coordination overhead angle, but you're right. Two central dependencies sound messy. Your idea of a registry for approved methods instead of a single suite is interesting. How do you handle versioning for those evaluation templates? If a method gets updated in the registry, do all the team projects using it have to manually update their references, or does it propagate automatically? That seems like a potential friction point.
Thanks, this is super helpful for me as we're just starting out. The idea of a `company-ai-governance` project as the single source for approved models is exactly what we need.
Do you find teams push back on using the central templates because they feel it's too restrictive? I worry about getting adoption right from the start.
Great question. The pushback is real, especially at first. The trick isn't to make the central templates less restrictive, but to make them so useful that teams *want* to use them.
We saw success by shipping a few "golden" templates that were demonstrably better than a blank slate, like a customer support summarizer with built-in PII redaction and a pre-configured testing suite. Teams adopted it because it saved them a week of work. The restriction becomes a feature when it's paired with obvious velocity.
That said, we also included a "sandbox" project with relaxed rules for pure R&D, which stopped teams from fighting the governance system for experimental work. The key is showing the value before enforcing the rule.
✌️
Love the focus on a central governance project as the top-level foundation. That's been critical for us too.
But I'd add a caveat to > a hierarchical model that mirrors your domain boundaries. If the org chart changes (and it always does!), you've now got a misaligned permission structure that's a pain to migrate. We've had more success defining Level 2 projects around enduring product capabilities or user journeys, which tend to be more stable than reporting lines.
Also, for that `company-ai-governance` project, we made sure to include not just the approved model configs, but also a set of starter "template of templates" for common patterns. It gives teams a running start and subtly enforces structure. Did you do something similar?
Beta tester at heart
Your caveat on domain boundaries is precisely why our Level 2 projects are organized around data products, not departments. The "Customer Sentiment Analysis" project outlasted three reorgs because the output, a cleaned and scored sentiment dataset, remained a constant need for consuming teams.
We also included "template of templates," but we found they needed to be extremely minimal. Ours are essentially parameterized skeletons with pre-populated sections for model configuration, grounding data reference, and a placeholder for the chain-of-thought block. Any more structure than that, and teams started forking just to strip out the parts they felt were in their way.
How do you handle inheritance when a team's use case diverges significantly from a starter template? Do you allow a full fork, or do you mandate a pull request back to the governance project to extend the base template?
Data doesn't lie, but folks sometimes do.
Agreed on starting with a central governance project. In our smaller rollout, we also made that Level 1 project the only one with direct billing integration. It forces cost visibility from day one, because all model usage flows through those approved configs.
But I'm curious about your point on mirroring domain boundaries. Our product teams are aligned by domain, but their AI needs are all over the place. How do you handle a team that needs both a customer-facing chatbot and an internal analytics agent under one domain? Would that be one project or two?
Still learning