Pushback is a given. The trick is making the pain of *not* using the central templates greater than the pain of using them.
We did two things: first, made billing flow through the governance project. Teams could build their own rogue model, but they couldn't get it approved for a budget without a config from the central library. Money talks louder than policy docs.
Second, as someone mentioned above, we baked in the evaluations. A template wasn't just a model config; it included a pre-loaded test suite. Using it meant a team's model was instantly ready for benchmark comparisons against the org standard. That made their lives easier, not harder.
- elle
You're spot on about using the Project as the top-level container for autonomy. I'd strongly echo the need for that `company-ai-governance` project at Level 1, but I'd add it's critical to implement strict, programmatic tagging conventions from the start.
Our governance project defines not just the model configs, but also a mandatory set of tags for cost center, data classification, and intended user (internal/customer). This is enforced via the project's template inheritance. Without it, you get a flat list of hundreds of prompts later with no way to slice usage data or audit effectively. The container gives autonomy, but the metadata inside is what provides the central oversight you mentioned.
Mirroring domain boundaries is a neat theory that falls apart the moment someone moves desks. Projects structured around shifting org charts become technical debt. I've spent more hours than I'd like untangling permissions for a project whose "owning" team was disbanded six months prior.
Your Level 1 governance project is necessary, but calling it a "foundation" undersells the trap. It's also a single point of contention. Every team's edge case becomes a ticket for your central platform group. Good luck scaling that support model when fifty teams all want a "small" exception to the gold-standard config.
And what's the actual ROI on that friction versus just letting a team use a slightly different temperature setting?
trust but verify
You've put a finger directly on the operational cost that governance can create. That central point of contention is real. The ROI question is crucial; you can't justify the friction for a temperature tweak.
Our solution was to formalize the exception process with clear tiers. A "small" deviation like a temperature setting gets an automated, self-service approval path using tags that flag it for later review. Only changes touching data handling or cost thresholds require a ticket. This keeps the platform team focused on high-impact governance, not configuration policing.
The harder ROI to measure, but the one we defend, is in audit and vendor negotiation. Without that central config library, you have no leverage when the vendor's terms change or during a security review. The friction upfront buys a clean contract position later.
Check the SLA.
Your point about modeling the dependency graph is critical, and I think your "Channel" concept addresses a real gap. We attempted a similar abstraction but called it a "Deployment," which represented a specific, versioned instance of a template that was wired into a particular event stream or API endpoint.
However, we ran into a significant issue with observability. By adding that layer of indirection, we fragmented the telemetry. Tracing a single execution from the triggering event, through the prompt template, to the model call and final output became a manual correlation exercise across systems. Freeplay's native metrics were attached to the template, but the "Channel" or "Deployment" abstraction lived outside it. Did you instrument your Channels with a consistent tagging strategy to maintain traceability, or did you accept that split?
Trust but verify.
That Zapier-for-Slack flow sounds like the perfect kind of low-friction gate. We tried something similar but used the Freeplay API to post a formatted message with the template diff pre-attached. It cut down the approval time even more because the approver didn't have to click through.
> Locking the model ID and version string
This is so critical. We even lock the underlying provider API key to the governance project. It prevents teams from accidentally (or "accidentally") using a deprecated or personal key, which keeps our security team happy and our cost tracking clean.
Did you extend that locking to things like the max tokens parameter? We've debated if that's a cost guardrail or an unnecessary constraint on a team's prompt design.
Spreadsheets > marketing slides.
Making the non-governed path painful is such a great point. The billing flow is clever.
I'm curious about how you built that benchmark test suite - did you create a standard set of evaluation prompts and expected outputs that everyone used? And was that managed in the governance project too? I'm trying to figure out how to set that up without it feeling like a huge extra task for teams already building their thing.