Skip to content
Notifications
Clear all

TIL: You can trigger evals from Freeplay's API, not just the UI

150 Posts
129 Users
0 Reactions
74 Views
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

That split is sensible. The risk is letting the "fast subset" drift from what the full suite actually catches. If your quick gate passes but the scheduled run finds a regression, you've already moved on and now you're debugging a change from two days ago.

You've got to treat that dashboard like a real alert, and budget the time to triage it. Most teams I've seen set this up end up ignoring the scheduled results because they aren't blocking anything.


Show me the bill


   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

The programmatic trigger is indeed useful, but I'd caution against posting the results directly to Slack without significant aggregation. We tried that initially and the channel became unusable due to notification spam from every merge.

Instead, we have the Action store the run ID and a link to the Freeplay UI in our deployment tracker. A separate, scheduled job collates all runs from the past 12 hours, calculates pass/fail rates and any performance metric deltas, and posts a single daily summary. This provides oversight without the noise.

Also, be mindful of the cost scaling if you run the full suite on every merge. We found it necessary to segment our evals into a fast "validation" suite (checks for correctness, no LLM-as-judge) that runs on PR, and the heavier "performance" suite that runs nightly. This keeps CI feedback under two minutes.



   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

That daily summary idea is a lifesaver. I can see how Slack would get flooded fast.

> be mindful of the cost scaling

This is my biggest worry. Your split into validation and performance suites sounds perfect. Do you think the validation suite alone is enough to catch most regressions, or are you mostly relying on the nightly run for the real safety check?



   
ReplyQuote
(@emilyc)
Reputable Member
Joined: 2 months ago
Posts: 161
 

Oh wow, I didn't know you could do that from the API! That sounds amazing for catching mistakes early.

> posts the results back to a dedicated Slack channel

I'm curious, do you get a lot of notifications? I'm worried my team would just mute the channel if it's too noisy 😅

Also, what happens if the API call fails? Does your whole pipeline stop?



   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Good question. It can get expensive if you run the full suite on every commit.

We split evals: a cheap, fast set (no LLM judge, just structure/parsing) runs on PRs. The expensive, comprehensive suite runs on a schedule. The PR check catches immediate breakage; the nightly run tracks drift.

Cost is all about what you put in that "fast" bucket. If you rely on LLM-as-judge for your gate, your bill will spike.


metrics not myths


   
ReplyQuote
(@emilyc)
Reputable Member
Joined: 2 months ago
Posts: 161
 

That makes so much sense. I was just thinking about all those LLM-as-judge calls adding up. Quick question, how do you decide what goes in the "parsing" bucket? Like, are you just checking that the output is valid JSON, or are you doing something more?



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Great question! We started with just JSON validation and regex checks for structured outputs, but that missed a lot. Now our parsing bucket includes things like:
- Required keywords or phrases in the response (for specific instructions like "include the word 'summary'")
- Length boundaries (is the response too short or suspiciously long?)
- Basic sentiment polarity using a simple lexicon, not an LLM
- Check for the presence of a numeric value if one was asked for

It's not as nuanced as an LLM judge, but it catches a ton of the "oops, we broke the prompt" scenarios without costing a penny in judge calls. The trick is to make these checks as semantic as possible for your use case.

What kind of outputs are you trying to validate? Maybe I can think of a specific check that'd work.


test everything twice


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

That's exactly the workflow we set up a few months back, and it's been a total lifesaver for preventing bad prompt changes from sneaking into prod. The immediate Slack feedback loop is fantastic.

One hiccup we hit was API timeouts on larger test suites. Freeplay's API accepts the trigger request quickly, but the actual evaluation run can take a while. Our CI/CD step would finish, thinking everything was fine, before the run failed. We had to adjust our action to poll for the run status for a few minutes before marking the step as successful. Just something to watch for if your suites get bigger.

Have you looked into triggering evals based on a drop in production metrics, like a sudden dip in user satisfaction scores? I'm curious if anyone's connected that data source.



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

That retry with backoff is a solid move for network flakes. The logging-to-observability piece is key, but I'd be careful about the 'exit with success' pattern. It's good for noise, but you need a hard failure mode too.

If we see the same eval fail consistently due to infrastructure timeouts, even after retries, we treat that as a blocker. It usually points to an underlying issue, like a dataset that's grown too large for the eval runtime, and you don't want that to silently degrade your reliability. So we have a separate threshold: two consecutive infrastructure failures on the same gate triggers a pipeline failure. That forces us to look at the root cause.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Absolutely agree with that separate threshold. We landed on something similar: three infra failures on the same pipeline step in a 24-hour window flips a circuit breaker and sends a PagerDuty alert. It's saved us a few times from a slowly degrading dataset that was timing out our semantic checks.

> It usually points to an underlying issue

Spot on. In our case, it's almost always a data issue - a new, huge log entry got added to the test set, or a dependent API we mock started returning massive payloads. Treating it as a blocker forces the cleanup.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That's a fantastic use of the API! Automating the smoke test evals on merge is exactly the kind of thing I love to see. It turns quality from a manual chore into a core part of the workflow.

We took a similar approach but with a slightly different trigger: we run our fast validation suite on every PR, not just the merge to main. It gives developers immediate feedback in their PR status checks before anything gets merged. The CI failure blocks the merge if the prompt change breaks a required check, like missing a key data point.

Your Slack notification is a great touch. We found that for PRs, posting directly to the PR comment thread worked better to keep the context tight. For the main branch runs, Slack is perfect for the broader team.


Cheers, Henry


   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

That's a great use case. I'm just starting with automated expense report analysis and hadn't considered running evals from CI/CD.

How do you handle test data? Do you have a static set of example customer queries you run for the smoke test, or does it pull from something live? I'm worried our test data would get stale.



   
ReplyQuote
(@emmab3)
Reputable Member
Joined: 2 months ago
Posts: 271
 

> how do you find the exact `testSuiteId` you need?

Yes, it's in the URL in the UI, but you can also get it programmatically. The `listTestSuites` endpoint returns everything, including IDs and names. I keep a mapping of suite names to IDs in our config so we're not hardcoding brittle IDs. The UI URL is fine for a one-off script, but for CI/CD you want something maintainable.

On the Slack noise: we post a summary, not the full report. The message includes the suite name, pass/fail, total duration, and a link to the detailed run in Freeplay. If it fails, we also include the top three failing test case names. The full JSON output goes to S3 for historical analysis. Piping everything into Slack would be unreadable.

The CI/CD integration is solid, but you need to build that polling logic for completion, as others mentioned. Don't just fire the trigger and assume success.


FinOps first, hype last


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

> Polling's a blunt instrument, but at least you control the timeout

Precisely. I've seen pipelines hang for an hour waiting on a webhook that died quietly. With polling, you set a thirty minute hard timeout and move on. It's ugly, but it fails closed.

The real trick is building a sane status check loop. Don't just sleep for sixty seconds between calls. Use an exponential backoff and quit early if the run status is clearly terminal. That at least gives you a chance to fail fast.


null


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

That Slack integration is such a clever idea for the main branch! We're doing something similar, but I'm curious about the mechanics.

When you say it posts the results back to Slack, what's the shape of that data? Are you parsing the full Freeplay response in the Action, or do you use their webhooks? I've been using their webhook to notify a channel when a run completes, but then we have to click through. Embedding a quick summary directly in the CI notification would be smoother.

Also, have you considered making the suite that runs on-merge different from your daily monitoring suite? Our on-merge tests are super fast, just checking for regressions in key outputs. The deeper, slower evals we run on a schedule.



   
ReplyQuote
Page 7 / 10