Okay, I have to jump in on this. I've been trying to get a handle on prompt management for our team, and it's been... messy. Spreadsheets, random docs, no version control.
The claim in the title makes sense to me now. Once you have more than a handful of prompts live, the chaos gets real. Which version is running in the A/B test? Why did performance drop last week? Did someone change the staging prompt but not production?
Freeplay feels like it's built for that exact tipping point. It's not just a vault to store prompts. For me, the killer feature is being able to trace a specific API call back to the exact prompt version and test config that generated it. It stopped three "it worked in staging!" fires last month alone.
Is that the main benefit for others, or are you using it differently? I'm still exploring the evaluation features. Would love to hear how teams are setting up their testing workflows.
Thanks!