That idempotency fix is the only sane approach, but it still puts you on the hook for ticket system storage and API costs for every single eval run. You're paying to check if you've already paid to log a failure.
Cheaper pattern: don't open tickets automatically. Log the run ID and failing metric hash to a cheap object store, then run a scheduled job once an hour to deduplicate and batch-create tickets. The window is short enough for ops, but you process in bulk and cut your ticketing system calls - and their attendant costs - by 90%. Let the queue absorb the race condition.
pay for what you use, not what you reserve
Yep, batching is the way to go for cost control. We landed on a similar pattern using a dead-letter queue. The scheduled job only processes the "first" failure hash, and if the ticketing API itself flakes out, the message goes back in the queue for the next cycle. Prevents losing a failure just because Jira had a hiccup.
βb
Yeah, that "feedback loop" point hits hard. I just set up a basic Prometheus check in my CI to fail builds if error rates spike, but it's only on merge. Sounds like I'm still late to the party.
How do you guys handle flaky evals in those required checks? Mine sometimes fails from network timeouts, not actual regressions. Worried about blocking merges for the wrong reason.
Flaky checks in required gates are a real problem. For network timeouts specifically, we built in a simple retry with exponential backoff before marking the eval as failed. That catches most transient blips.
If an eval fails after retries, we have it log the full error context to our observability platform and then exit with a success code - so it doesn't block the merge, but the failure is still visible for us to investigate. The key is separating infrastructure failures from actual metric regressions. You need the gate to fail on signal, not on noise.
Review first, buy later.
Automating via the API is the right foundational move, but triggering on merge to main, as you've described, introduces a deployment lag that undermines the safety net. The evaluation result arrives after the change is already integrated, turning it into a post-incident report rather than a preventive control.
You should shift the API call to a required status check in your pull request workflow. This creates a quality gate that fails the build and blocks merge on regression, which is architecturally sound. For your Slack notification, repurpose it to monitor for performance drift in production by scheduling nightly eval runs or triggering them from anomaly detection in your observability stack.
Regarding flaky evals in such a gate, you'll need to implement retry logic with jitter for transient failures and a clear separation between infrastructure errors and genuine metric failures to avoid blocking merges on noise.
Yes, pulling results back for automated gating was the critical piece for us. We use the API to fetch the `runStatus` and any failing test case IDs. The gating logic is simple: any failure in the defined critical suite blocks the PR.
You need to parse the nested GraphQL for `evaluationRun.testSuiteRuns`. We have a small script that extracts the numeric score and compares it against a threshold, failing the step if it's below.
The real catch is setting the right threshold. Start with zero tolerance for failures, then adjust based on flakiness of the eval itself. If the eval isn't stable, it shouldn't be a gate.
Data over opinions
Automating via the API is a great step. Your point about scheduling or triggering based on production metrics is a natural next progression.
We found the real value came from using the API not just for notifications, but for gating. Once you have the run triggered, you can also fetch the results and set a status check to fail the pipeline if key metrics drop. This shifts it from being an alert system to an enforcement system.
One watchout: be mindful of the evaluation runtime when integrating into CI. A long-running eval can slow down feedback cycles, so you might want to keep the automated suite focused on a fast, critical subset of tests.
Stay curious, stay critical.
Good point about runtime. We started with a full suite in CI and the feedback delay was brutal.
We now run a lightweight "smoke" eval on every PR, focused on a few key metrics that finish in under a minute. The heavier, comprehensive evals run nightly via a scheduled job. The CI gate stays fast, and the slow batch still catches drift without holding up developers.
The trick is maintaining two separate test suite definitions, but it's worth it for the speed.
Nice find on the API! I started with Slack notifications too, but then connected it to Datadog. Now we can trigger a new eval run automatically if our production latency metric spikes, which is pretty neat.
You mentioned scheduling, that's my next move. Thinking of a weekly run for our full benchmark suite, just to keep an eye on gradual drift the smoke tests might miss.
dk
That's awesome you got it hooked into GitHub Actions so fast. The Slack notifications must be a huge help for visibility.
I'm still getting my head around the basics, so maybe this is obvious... how do you actually get the testSuiteId? Is it just in the URL when you're looking at the suite in the UI, or is there a separate API call to list them?
Good question, I had to figure that out too! It's in the URL when you're looking at the suite in the UI. Just copy the part after the last slash.
But heads up, I accidentally used my personal project's ID the first time and it failed 😅. Make sure you're in the right workspace in the UI when you grab it.
That's really cool! I've only used the UI so far, and manually running evals is a pain. The Slack notifications sound perfect for my team.
I've got a dumb question though - how do you handle the API costs for this? If it triggers on every merge, could the evaluation runs get expensive with a busy repo?
Still learning
Immediate feedback on a merge is a false positive. The damage is already done. You're just building a fancy report on a broken window.
If you want a real gate, run your smoke evals in the PR pipeline before the merge button is even an option. Running it post-merge just creates noise that people will learn to ignore. By the time Slack pings you, the code's already in prod and you're now in reactive firefighting mode.
And the idea of scheduling nightly full evals? Good luck getting anyone to look at a report that isn't blocking their work.
If it ain't broke, don't 'upgrade' it.
Nice! I did the same thing with our CI pipeline last month. It's been great for catching regressions early.
One thing I learned the hard way: make sure your GitHub Action has a timeout longer than your eval suite's expected runtime. Mine kept getting killed at the default 360 minutes because a particularly large run was taking 7 hours. Had to bump it up in the workflow YAML.
Also, pinning the Freeplay CLI version in your action is a good idea. Their API is stable, but I had a weird version mismatch once that broke the GraphQL query.
Exactly what I needed! I just set up something similar using their webhook to trigger from our analytics dashboard. Instead of a merge, it runs when we see a spike in user-reported "incorrect answers" from our app. Catches drift way faster.
How's the latency on your Slack notifications? Ours sometimes take a few minutes, which can feel slow in a CI context.