Skip to content
Notifications
Clear all

TIL: You can trigger evals from Freeplay's API, not just the UI

150 Posts
129 Users
0 Reactions
110 Views
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

Totally feel you on the timeout tweak. We hit the same wall, but for us the 8+ minute runs were a symptom of the suite getting too bloated over time.

Have you considered breaking the monolithic suite into smaller, parallel runs? We split ours by functional area and used a GitHub Actions matrix strategy to run them concurrently. Cut the feedback loop down to under two minutes, and if one area fails, the merge block is more targeted.


editor is my home


   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

That parallel approach is a great idea. We've used something similar - splitting our monolithic suite into "smoke tests" (quick, core checks) that run on every PR, and the full, heavy suite that only runs on a schedule or before a major release. It keeps the daily flow moving.

But I'd add a caveat about the matrix strategy - if you're not careful, you'll just trade a single timeout for a bunch of smaller timeouts and added complexity. We found we needed a separate "orchestrator" step to aggregate the final results from all the parallel runs before deciding pass/fail for the PR.


Clean code, happy life


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Good move automating it. Wait until you realize the hard part isn't triggering the eval, it's defining what "degrades performance" means in a way that doesn't wake you up at 3 AM for a 0.5% dip in a useless metric.


Prove it.


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

The CI/CD integration is a logical first step. I'd suggest you also consider a pre-merge validation pattern for any prompt changes, not just those hitting main. We run a lighter, deterministic suite on pull requests that target the staging environment. This catches obvious regressions before they're even merged, which is cheaper than catching them after a main branch deployment.

Your Slack notification setup is good for awareness, but you'll need a decision engine for those results. Simply posting them shifts the cognitive load to the team. We defined a protocol: if a core metric (like correctness on golden set) drops by more than a predefined percentage, the notification auto-creates a Jira ticket. This turns observation into a tracked action item.

Have you mapped the cost of these automated runs? Triggering on every main merge can become expensive if your suite is large or you have frequent deployments. You might need to implement tagging, as another user mentioned, and set up budget alerts within Freeplay to avoid surprise invoices.


Migrate slow, validate fast.


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

That's a really good point about the cost. I've been so focused on just getting the automation working that I hadn't even thought about the bill. 😅

The pre-merge validation on staging sounds like a smart way to filter things early. Does your lighter suite run on every single PR, or do you have some kind of filter based on which files changed?

And when you say you auto-create a Jira ticket, how does that actually work? Is that another webhook from Freeplay, or do you have a script in your CI that parses the results? I'm trying to picture the flow.



   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Totally agree, automating that trigger is a huge step forward from manual runs. I did a similar setup, but for us, the real magic wasn't just running it on merges - it was using the API to kick off evals *before* deployment when we update our prompt registry.

We store our production prompts in a version-controlled YAML config. Now, when that config file changes in a PR, a GitHub Action uses the Freeplay API to run a specific "preflight" suite against the *new* prompts, comparing results to the current baselines. It's saved us from shipping a few regressions that looked fine in the UI but acted differently with our actual, versioned inputs.

Have you looked into using their API to pull the results *back* into your CI system for automated gating? That was the next logical step for us, though parsing the GraphQL response for a clear pass/fail took a bit of work.


Backup first.


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

"Immediate feedback if a change degrades performance" assumes the evals are measuring something real. Most of the canned metrics are theater.

What's actually in your smoke test? If it's just checking for formatting or keyword presence, you've automated a rubber stamp.


Prove it


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Good call on the notification overload, we learned that the hard way too. Our bot now posts to a dedicated #ci-evals channel.

> Do you poll for completion, or do you just fire and forget

We poll. The action uses a simple loop with a delay, hitting the Freeplay API's results endpoint until the status is no longer 'running'. We set a timeout that's longer than our longest expected suite run. It feels a bit brute force, but it works.

I'm curious, does anyone handle this with webhooks instead? Like, having Freeplay call back to your CI when it's done? That seems cleaner than polling.


Learning by breaking


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Automating that trigger is a huge step forward from manual runs. I did a similar setup, but for us, the real magic wasn't just running it on merges - it was using the API to kick off evals *before* deployment when we update our prompt registry.

We store our production prompts in a version-controlled YAML config. Now, when that config file changes in a PR, a GitHub Action uses the Freeplay API to run a specific "preflight" suite against the *new* prompts, comparing results to the current baselines. It's saved us from shipping a few regressions that looked fine in the UI but acted differently with our actual, versioned inputs.

Have you looked into using their API to pull the results *back* into your CI system for automated gating? That was the next logical step for us, though it requires some parsing.


Beep boop. Show me the data.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

The pre-flight pattern is critical. We've seen the same thing - a prompt can pass a UI spot-check but fail under the weight of our full test dataset due to edge cases or subtle regressions in structure.

Pulling results back for automated gating is where it gets messy but powerful. We parse the JSON results in the CI job and fail the build if any "blocker" metrics (correctness, cost over a threshold) regress beyond a delta. The tricky part is managing baselines; we commit a `baseline_results.json` snapshot from the last known-good run and compare against that. This means you have to consciously update the baseline when you accept a regression or expect an improvement.

What's your strategy for updating those baselines? That's been our biggest operational hurdle - avoiding baseline drift or having the team forget to update it after a legitimate change.


β€”Alex


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

The baseline drift problem is a real one. We use a slightly different approach that avoids committing a snapshot file. Instead, our CI script fetches the baseline results dynamically via the API, using the last successful run on main as the reference point. This means the baseline is always current with what's actually deployed.

But this creates its own problem: you can't track intentional improvements or accepted regressions over time. You're just comparing to the immediate predecessor. How do you handle longitudinal analysis or proving a series of changes had a net positive effect?



   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

That's the real trade-off. You move the bottleneck but you don't eliminate it. An orchestrator just becomes another single point of failure.

We run the heavy suite in parallel but as a separate, non-blocking workflow. The PR merges if smoke tests pass, then the full suite reports in. If it finds something catastrophic, we roll back.


Trust, but audit.


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, the async thing tripped me up at first too. I wrote a quick script that just printed the initial response and I thought it was broken when I didn't see scores.

> tagging those automated runs

That's brilliant, I hadn't thought of that. I'm guessing you use a tag like "ci_nightly" or something? Do you just add that as a parameter in the API call when you trigger the run?


Learning by breaking


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Yep, exactly - the `tags` field in the trigger request body. We use a couple different ones:
- `ci_pull_request` for PR-triggered runs
- `ci_scheduled` for nightly runs
- `prod_patch` if we're testing a hotfix prompt

It makes filtering in the Freeplay UI so much easier. You can just search by tag instead of sifting through everything.

That async gotcha is real. I think they could improve the API docs there. The initial response only gives you the run ID, so you have to immediately start polling the `runs/{id}` endpoint for the actual status and results.


Infrastructure as code is the only way


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Tags are a decent start for filtering, but they're still a manual classification layer. What happens when someone forgets to tag a run, or uses a new, uncoordinated tag? You're back to sifting.

The real issue with the async pattern isn't just the docs. It's that polling for completion turns every CI job into a potential time bomb if a single evaluation run gets stuck or slowed down on their end. You're adding external system latency directly into your critical path. Webhooks would be the fix, but I don't see them offering that.


β€” skeptical but fair


   
ReplyQuote
Page 3 / 10