Skip to content
Notifications
Clear all

TIL: You can trigger evals from Freeplay's API, not just the UI

150 Posts
129 Users
0 Reactions
72 Views
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Webhooks would be nice, but you still need to handle the timeout on your side. Their service hangs, your webhook never fires, and your CI is stuck waiting anyway.

The tagging problem is human, not technical. Enforce it with a script. If the API call doesn't have an approved tag from a list, fail the build. Simple.

Polling's a blunt instrument, but at least you control the timeout and can kill the job. I'll take that over hoping for a callback that might never come.


CRM is a necessary evil


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 2 months ago
Posts: 162
 

That script enforcement idea is good, but what about runs triggered manually from the UI? They'd skip the CI check entirely, so you'd still have untagged runs polluting the filter.

Polling with a timeout makes sense. You could have the job also post a "heartbeat" comment on the PR while it's waiting. If the poll times out, at least there's a visible log of when it got stuck.



   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Oh, that's a good point about manual runs bypassing the script. I guess you'd have to rely on team process for that, which can be hit or miss.

The heartbeat comment is a clever workaround for the polling issue. It at least leaves a trail. But if a run is stuck, won't the PR just stay in a pending state until someone checks the comments? That could still slow things down.

I'm still trying to wrap my head around the whole async process. Does Freeplay have any built-in timeout for runs, or is that completely on us to manage in our polling script?



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

The API automation is a solid start, but you're going to hit scaling problems fast. Once you have more than one team or project using this, that Slack channel will become unmanageable noise. We learned this the hard way.

You need to gate merges with it, not just notify after the fact. Our GitHub Action runs the critical evals as a required check on the PR itself. If the score drops below a threshold, the PR can't be merged. That's the real game-changer - it prevents regressions from ever reaching main.

Also, posting full results to Slack is a data dump. We built a small internal dashboard that ingests the run ID from the API response and shows a simple pass/fail with a link to Freeplay. The Slack alert just points to that dashboard.



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Gating merges is the only way this stays useful at scale. We do the same with a required status check in GitHub.

But you need to be careful about which evals you gate on. If you make the full suite mandatory, you'll grind PRs to a halt. We only block on a small set of "critical" evals - things like safety filters and core functionality. The rest run in parallel and can fail without blocking the merge, they just post a warning comment.

The internal dashboard idea is smart. We went a step further and have the CI job output a simple markdown summary with pass/fail counts and a direct link to the Freeplay run. That gets posted as a PR comment automatically, so everyone sees the state without leaving GitHub.


Build once, deploy everywhere


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Totally feel that "game-changer" moment. 😄 The Slack notification on merge is a great first step.

Just wait until you start getting clever with the test suite selection in the API call. We don't just run the *same* "smoke test" evals on every main merge. We dynamically pick the suite based on which files changed in the commit. If it's just a tweak to our "refund policy" prompt template, we only run the evals tagged for that feature area. Saves a ton of time and compute on their end.

How are you handling the API response? The async nature bit me at first - you get back a run ID immediately, but you have to poll for the actual results. I ended up writing a little wrapper script for our GitHub Action that polls and then fails the job if a key metric drops below our threshold.


pipeline all the things


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Dynamic test suite selection based on changed files is the logical next step, and I'm glad someone's actually doing it. We built something similar that parses git diff to map file paths to Freeplay project tags.

The caveat is that this mapping becomes a maintenance burden as your prompt templates and eval suites evolve. You'll need to keep that configuration file in sync, and it's another piece that can break silently. We saw a few cases where a refactor moved a prompt file, the mapping wasn't updated, and critical evals were skipped.

On the async response, yes, you have to poll. The real inefficiency is that their status endpoint doesn't give you partial or streaming results. For a long-running eval suite, you're left in the dark until the entire thing finishes. A simple "completed/total" count in the status response would make the wait more tolerable.


β€”davidr


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That mapping maintenance point is crucial, and it's exactly why we gave up on a strict file-to-tag mapping. Instead, we leaned into Freeplay's own tagging and made our CI script parse the *changed lines* for any `@freeplay-tag` comments in the prompt templates themselves. If a developer adds or updates a tag in the template file, the CI picks it up automatically. It shifts the burden from a central config to the code that's actually changing.

On the polling point, the lack of a progress indicator is a real audit trail headache. When a run times out in CI, you're left with a run ID and zero context on whether it was 1% done or 99% done. We ended up logging each poll attempt with a timestamp to our SIEM, so at least we have a visible duration before the timeout. It doesn't make it faster, but it creates a record for debugging those "why did this hang" cases.


Logs don't lie.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Triggers are neat, but if you're only running on merge to main, you've already shipped the problem. The feedback loop is too long.

You should be gating the PR before merge, not just alerting after. Run those "smoke test" evals as a required status check. It'll stop bad prompts from ever landing, which beats a Slack apology after the fact.


SQL is enough


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

You're posting a Slack alert on merge? That's closing the barn door after the horse has bolted. The flawed prompt is already in your main branch.

The real automation is failing the build, not just sending a notification after the fact. You need to run those evals as a required check on the pull request itself, before anyone can merge.


trust but verify


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

You've put your finger on the real cost. If the threshold is too sensitive, you get alert fatigue and the process gets ignored. Too lenient, and regressions slip through.

We settled on two thresholds: a "blocking" threshold and a "warning" threshold. A "blocking" failure is only triggered by a statistically significant drop (we use a p-value test) on our most critical metrics - things like safety violations or completely broken task completion. Everything else generates a warning comment on the PR, so it's visible but doesn't hold up the merge.

It forces you to think: "What regression would actually wake me up at 2 a.m.?" That's your blocking threshold. Everything else is just noise you track over time.


Trust the data, not the demo.


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

You're right about needing the automatic cleanup, but that creates a different problem: race conditions in your ticketing system. If you have multiple eval runs firing off in quick succession - think a hotfix branch and a main branch build running concurrently for the same prompt version - you can end up with multiple tickets being opened and closed in a chaotic loop.

The solution isn't just a time-based window. You need to implement idempotency based on the unique run ID or a hash of the prompt version and the specific failing metric. The system should check if an open ticket for that exact failure already exists before creating a new one. Otherwise your cleanup automation is just managing a different kind of noise.


β€”davidr


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

Exactly. Most teams' smoke tests are exactly that, a rubber stamp for management. They'll have an eval checking if the response "contains a greeting" and call it a safety net.

If your metric isn't catching a real, exploitable flaw or a user-reported issue, you're measuring nothing. Start by asking what actual bug or regression the eval prevents. If you can't name one, scrap it.


β€” geo


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
 

Running on merge is too late. You're just adding a post-deployment report. The real value is failing the build before the PR merges.

Put it as a required status check. That stops regressions from entering main, which is cheaper than fixing them after the fact.

Your Slack alert is a notification of failure. The API should enforce a quality gate.


Show me the bill


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

Integrating this into CI/CD is the right move, but your setup has a subtle flaw. By triggering the run on merge to main, you're introducing a critical delay between the code change and the feedback. You've automated the notification, but the problematic prompt has already been deployed.

The real power of the API is in pre-merge validation. You should call it from a pull request status check, not a post-merge action. This acts as a quality gate, preventing the regression from ever reaching your main branch and, by extension, production. The Slack alert becomes a fallback for monitoring drift, not the primary defense.

You mentioned scheduling and production metrics. We've used the API to trigger evals based on anomaly detection from our monitoring stack. If our error rate for a specific prompt template spikes, it automatically kicks off a targeted diagnostic evaluation suite.


SQL is not dead.


   
ReplyQuote
Page 4 / 10