I've been using Freeplay for a few months now, primarily through the UI to run evaluations on our LLM prompts. It's been solid for tracking performance drift and testing new templates.
While poking around their documentation for something else, I spotted a note about their GraphQL API. Turns out, you can trigger a new evaluation run programmatically. This is a game-changer for our CI/CD pipeline. We can now automatically run a suite of evals whenever a prompt change is merged to our main branch, instead of relying on manual clicks.
The setup was pretty straightforward. You need:
* Your project's API key (found under Project Settings)
* The specific `testSuiteId` for the evaluation suite you want to run
* A simple POST request to their GraphQL endpoint
I wired it into a GitHub Action. Now, every merge to our main branch fires off our critical "smoke test" evals and posts the results back to a dedicated Slack channel. It saves our team from remembering to run them manually and gives us immediate feedback if a change degrades performance.
Has anyone else automated their eval workflows? I'm curious about other use cases—maybe triggering based on production metrics or scheduling daily regression tests.
Straightforward until your API key rotates or they change their GraphQL schema without a major version bump. How are you handling version pinning in that GitHub Action?
Also, "immediate feedback" sounds good, but what's the actual SLA on that API? If your pipeline is blocked waiting for a result and their eval queue is backed up, you're stuck.
read the fine print
Your concerns are valid, but they represent a failure in procurement and integration design, not necessarily a flaw in the API itself.
On key rotation, you shouldn't be embedding a static API key in a pipeline. The key should be sourced from a secrets manager, which your automation can check at runtime. If rotation breaks your process, your secret management is the bottleneck. For schema changes, any competent vendor provides a deprecation timeline; you should be monitoring their changelog channel as part of your operational hygiene.
Regarding the SLA and queue, that's a critical contractual point. You don't rely on implied performance. This needs to be specified in your agreement: a maximum runtime or an asynchronous callback pattern. If they can't commit to that, then you're right, it's unsuitable for a blocking CI/CD step. Have you reviewed their terms for execution time guarantees?
That procurement line is a bit idealistic. Most of us are stuck integrating with tools our sales team already bought. The contractual SLA review is the right step, but good luck getting engineering time allocated to renegotiate terms after the fact.
My bigger issue is the async callback pattern you mentioned. If we're talking CI/CD, you can't just fire and forget. The pipeline needs a pass/fail signal. That means either blocking on the result with a timeout, or building a secondary notification system that can update the commit status. Both add complexity.
I ended up with a hybrid approach: trigger the eval async, but then poll for completion with a max wait time. If it times out, the step fails and we get a notification, but the eval keeps running in Freeplay for later review. It's not perfect, but it decouples the pipeline from their queue depth.
Automate everything. Twice.
That's a really cool idea for CI/CD! I never thought about hooking it into Slack like that.
How are you handling the results when they come back to the channel? Are you just getting a simple pass/fail or are you piping the full report? I'd worry about it being too noisy if the suite is large.
I'm still getting my head around the basics, so this might be a dumb question, but how do you find the exact `testSuiteId` you need? Is it just in the URL when you're looking at a suite in the UI?
Oh that's super interesting! I never thought about automating evals like that. My team just uses the UI, but we keep forgetting to run them after we tweak a prompt.
> how do you find the exact `testSuiteId` you need? Is it just in the URL when you're looking at a suite in the UI?
I had the same question! I was poking around the UI and couldn't find it easily. Did you just look at the URL? Also, how long does the API call usually take to kick things off? I'm wondering if there's a delay that would slow down our merge process.
Ah, the joy of adding yet another API key and external dependency to your CI/CD pipeline. The immediate feedback loop sounds nice until you realize you've just tied your deployment velocity to a third-party's uptime and queue depth.
Have you calculated what happens to your team's throughput if Freeplay has a bad day? That Slack channel will get very noisy, very quickly.
Beware of free tiers
The API trigger is useful, but you should separate the latency-sensitive CI gate from the full evaluation. I run a smaller, faster subset of tests via the API to fail the build quickly for regressions. The comprehensive suite runs on a schedule and results go to a dashboard, not Slack.
You're right that adding dependencies always carries a risk. That's a valid concern for any critical pipeline.
The key is to architect around the dependency, not just hope it never fails. Building in a sensible timeout and a clear fallback path means a Freeplay outage doesn't have to halt all deployments. It just means that specific quality gate is temporarily blind, which you'd then handle with a manual check or a rollback safety net.
Keep it constructive.
Automating your evals into CI/CD is the right move, but embedding it directly into your main branch merge is reckless.
You're creating a hard dependency on an external service for deployment. If that API call hangs or fails, your pipeline is broken. That's a single point of failure you just volunteered for.
And you're posting the results to a Slack channel? That's just alert fatigue waiting to happen. What's your criteria for a pass? If it's anything more nuanced than a binary, you've now dumped a complex report into a chatops pipeline that's meant for signals, not analysis.
— geo
I generally agree with your risk assessment, but framing it as "reckless" oversimplifies the engineering choices involved.
You're correct that a direct dependency is a potential single point of failure, but that's true for any external service in CI/CD, like your artifact registry or cloud provider. The mitigation is standard: implement aggressive timeouts and circuit breakers. If the Freeplay API call fails, the pipeline step can be configured to fail open with a warning, not block deployment, logging the incident for follow-up. You're not "volunteering" for a SPOF; you're making a calculated trade-off for automated quality checks, which can be designed defensively.
On the Slack noise, that's a pure implementation detail. The pass/fail criteria should be a binary decision computed *before* notification. Our process runs the eval, the pipeline script parses the JSON response for key metrics against thresholds, and only posts a failure summary to a dedicated alerts channel. A full report goes to a data warehouse table for later analysis. The Slack message is just the signal.
Garbage in, garbage out.
The "fail open with a warning" pattern is critical. My threshold for blocking a deploy is high, so if the eval service flakes, the step logs a severe warning but the pipeline continues. The next stage, a canary deployment with real user monitoring, becomes the actual gate.
Parsing the JSON for a binary decision is the only sane way. Our rule: if the pipeline script can't cleanly parse the result, it's a fail. That catches API changes or malformed responses.
The testSuiteId is indeed in the URL when you view a suite in the Freeplay UI. The pattern is typically something like ` https://app.freeplay.ai/projects/{project_id}/test-suites/{test_suite_id}`. You can extract it from there.
On handling results in Slack, piping the full report is a recipe for notification fatigue. You need a parsing layer in your CI script that distills the JSON response into a strict pass/fail. The script should post only the conclusion and, perhaps, a link to the detailed run in Freeplay for anyone who needs to investigate.
Regarding noise with a large suite, that's why you should consider a tiered approach. Run a critical subset of your evals as the CI gate for speed and binary feedback, and schedule the full suite separately, sending those more comprehensive results to a dashboard or a dedicated channel.
numbers don't lie
This is exactly how we use it too. We actually trigger evals from a webhook whenever a new prompt version is tagged in our internal prompt management system, not just from a branch merge.
One important caveat we ran into: the GraphQL mutation returns quickly, but the evaluation run itself is async. If you need to block the pipeline on the *results*, you'll have to poll the API for that specific run's completion. That added complexity is why we went the "fail open" route others mentioned.
Ship fast, measure faster.
Absolutely, the async nature is the real kicker. We handle that polling with an exponential backoff script, but it's still extra moving parts.
Your webhook trigger idea is smart, it decouples the eval from the code deployment cycle entirely. That way, prompt changes can be validated as soon as they're ready, not just when the app ships.
Show me the accuracy numbers.