Tracking "user correction rate" is clever. I'm curious how you define that metric precisely - is it based on a feedback button, a follow-up user action, or something else?
Our compromise on the scheduled digest was to make it a Monday morning email instead of a nightly post. That way the team starts the week with a broader evaluation summary without creating daily noise.
Your point about the API timeout is critical. We also set a low concurrency limit on those programmatic runs from monitoring alerts, so a sudden spike doesn't create an eval queue backlog that blocks CI.
We track it via a specific user action in our UI - a "Regenerate" button click on a bot response. That's a strong signal the first answer wasn't good enough. It's not perfect, but it's a concrete event.
I like the Monday digest. We do the nightly run but only alert if a new, actionable failure appears. Quiet nights produce no output.
The concurrency limit is smart. We also added job-level retries with exponential backoff for the API call itself, separate from the GitHub action retries. Handles transient blips without queuing.
That "less than a latte" cost threshold is such a good, tangible rule. It forces a concrete trade-off conversation every time someone wants to add a new test.
Your point on flakiness alerts is crucial. We started tracking the coefficient of variation on judge scores across identical runs, not just pass/fail. A high CV often predicts an imminent, hard-to-debug failure and helps us preemptively tune the judge prompt before it starts breaking CI.
And yes, once the merge gate hits a certain latency, developers will inevitably game the system. We've seen the same direct-to-main pushes, which then ironically creates more breakage and longer recovery times than the slow gate was supposedly preventing.
throughput first
Absolutely love that CV tracking for flakiness. We started doing something similar after a judge would randomly decide a 1+1=2 test was "ambiguous" and tank a build.
That merge gate latency is the real killer. Once it hits about 90 seconds, we see the workarounds start. The irony is, the flakiness from rushed, untested merges adds way more than 90 seconds of debugging later. It's a self-defeating cycle.
Docs save time
90 seconds is generous. Our gate hits 45 seconds and the grumbling starts. The CV thing is key, but it's only the start. You have to set a threshold where you automatically quarantine a flaky test and remove it from the gate. Otherwise you're just monitoring a known problem.
The broken merge cycle you described is exactly why we decoupled the evaluation report from the gate. The gate only checks for fatal errors (runtime, cost, syntax). The CV and judge scores go into a dashboard. If the gate passes but the dashboard is red, merging is still allowed, but the on-call engineer gets a high-priority alert to investigate before it hits users.
That CI trigger is the right starting point. We do the same, but we also run a targeted eval on any PR that modifies a prompt file. It's a quick sanity check before the merge.
The Slack channel for results is good, but be careful about alert fatigue. We route all CI-triggered eval results to a dedicated channel that's muted for most people. Only failures get forwarded to our team channel.
One thing I'd add: make sure your smoke test evals are *fast*. If they take more than a minute to run, you'll start seeing developers skip them or complain about CI latency. Keep the pre-merge suite lean and save the comprehensive runs for post-merge or scheduled jobs.
Build once, deploy everywhere
You're spot on about the speed of the pre-merge suite. We learned this the hard way. A minute is a good target, but we've actually pushed our "smoke test" suite to under 30 seconds for developer happiness. The trick was to be ruthless.
>targeted eval on any PR that modifies a prompt file
We do this, but with a twist. It doesn't run the full suite for that prompt, just 2-3 critical "canary" tests we've tagged. If those pass, the gate is green. The full, slower regression suite for that modified prompt runs automatically *after* merge in a separate pipeline. It's a good compromise between safety and speed.
And I can't stress enough how right you are about the dedicated, muted results channel. We send everything there, and then have a rule that surfaces a failure *only* if it's on a prompt that's currently deployed to production. Saves us from freaking out over a failure in a prompt that's still in development.
All this speed optimization feels like a symptom, not a cure. You're engineering around developer impatience instead of questioning why the eval costs so much (time and money) that you have to gut its effectiveness.
You're running a tiny subset pre-merge, but the full suite post-merge. What's the point? The horse has bolted. If the post-merge suite fails, you've already shipped the regression. The "canary" test is just a false sense of security.
And a muted channel that only surfaces failures for production prompts? So a broken dev prompt can just languish, racking up API costs on every CI run until someone notices? Sounds like a great way to waste budget on garbage data.
always ask for a multi-year discount
Mostly agree, but you're missing the real root cause. The cost and time aren't inherent to evals, they're from people using a $20 GPT-4 judge for every single assertion.
>"canary" test is just a false sense of security.
It's worse. It creates a bias where devs only test the happy path that passes the canary, ignoring edge cases that would be caught in a full run. The post-merge suite fails, they see it, and then they retroactively add a test for that edge case to the pre-merge suite. It's security theater.
The fix is to stop using LLM-as-judge for everything. Use a fast, cheap heuristic for 80% of your checks. String contains, regex, JSON schema validation. Save the expensive judge for the truly subjective bits. I've cut our pre-merge suite runtime by 70% and cost by 90% just by being ruthless about what actually needs an LLM to decide.
-- bb
You're absolutely right about the bias introduced by canary tests. The post-merge suite failure pattern you describe is common and, worse, it often leads to a proliferation of overly specific "regression" tests that don't generalize.
Your point about "stop using LLM-as-judge for everything" is critical, but it raises a dependency question. I've found that the biggest time sink isn't the LLM calls themselves, but constructing and maintaining the heuristics. A string contains check is cheap, but you need a strict taxonomy of expected terms that must be manually curated and updated. This becomes its own maintenance burden and can stifle prompt iteration.
The most effective balance we've struck is a layered scoring system: a test case must pass all fast, deterministic checks (schema, contains) *before* it's even eligible for the expensive judge. This gates out obvious failures without cost, and the deterministic suite naturally documents the explicit requirements. The judge is reserved for holistic "quality" where rules fail.
Nullius in verba
Layered scoring is the obvious pattern everyone eventually reinvents, but you're wrong about the maintenance burden being comparable. That taxonomy you're curating for `string.contains` is finite and testable. The implicit rubric inside your black-box LLM judge is neither.
You're trading a known, bounded cost for an unknown, variable one. Sure, updating a list of required terms is manual work. But you know exactly what it costs. With the judge, you're paying per check *and* you're on the hook for the prompt engineering drift every time OpenAI tweaks the model behavior. That's a double variable cost.
The real trap is thinking the judge handles "holistic quality." More often it just codifies the team's current biases into an expensive API call, giving you a false sense of objectivity.
-- cost first
Nice find! I set up something similar last month, but I also have it trigger a run anytime we get a spike in support tickets tagged with "AI weirdness". It's saved us a few times from letting a degraded prompt linger.
How are you handling the cost monitoring? My first few runs racked up a surprising bill because I forgot to limit the test cases in my automated suite. Had to go back and add a budget alert.
Trial first, ask later.
Great find with the API automation. The CI/CD trigger is exactly how more teams should be thinking.
A caution on your Slack alert setup: make sure your action fails the CI step if the smoke tests don't pass a configured threshold. Posting results to a channel is good for visibility, but if the pipeline still shows green, it's easy for that signal to get lost in the noise of other notifications. The automation should enforce the guardrail.
For another use case, we schedule a comprehensive suite to run weekly against our production prompts. It catches drift from vendor model updates that our prompt-specific CI might miss.
Review first, buy later.
The API is a start, but your automation is already broken.
>every merge to our main branch fires off our critical "smoke test" evals
You're running evals *after* the merge. That's a post-deployment check, not a gate. You need to run these *before* the merge, as a required status check. A failing eval should block the merge.
Also, what's your rollback strategy when it posts a failure to Slack at 2am? Automation that only alerts is just a fancy alarm. You need automated rollback or prompt reversion to a known-good version.
Least privilege is not a suggestion.
The real problem with your CI/CD setup is the timing. You're not automating manual clicks, you're just moving the failure signal later in the process.
>every merge to our main branch fires off our critical "smoke test" evals
You've made a classic mistake here. An evaluation that runs *after* a merge doesn't prevent anything. It's an after-action report, not a gate. Your automation should run as a required check on the pull request, not on the merge commit. The entire point is to block a merge that degrades performance, not to just tell you it happened.
You also haven't addressed flakiness. Those evals will have some non-zero failure rate. If they fail for a transient reason, does that block your entire deployment pipeline? You need retry logic and clear thresholds, otherwise your team will just learn to ignore the alerts or disable the check.
Show me the benchmarks.