Skip to content
Notifications
Clear all

Switched from automated to hybrid human+AI eval. Here's the accuracy bump.

10 Posts
10 Users
0 Reactions
24 Views
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
Topic starter   [#24941]

For the last 18 months, our team relied exclusively on automated evaluation frameworks (LLM-as-judge, BERTScore, ROUGE) to benchmark our production RAG pipeline's response quality. While the metrics were consistent and allowed for rapid A/B testing of retrieval strategies and prompt templates, a persistent gap existed between our "improving" scores and sporadic user complaints about factual inaccuracies or incomplete reasoning. The dissonance suggested our automated evals were optimizing for a local maximum, not genuine utility.

We designed a controlled experiment to quantify the delta. The core hypothesis was that a hybrid evaluation system—combining deterministic checks, LLM-as-judge, and targeted human review—would yield a more accurate and actionable diagnosis of system failures, ultimately leading to a more significant accuracy improvement than automated-only tuning.

**Methodology & Baseline**
We took a frozen snapshot of our pipeline (PGVector for retrieval, gpt-4-turbo for generation) and a curated evaluation set of 500 complex, multi-faceted user queries. We established two parallel evaluation tracks:
1. **Automated-Only (Control):** Used our existing suite: answer relevance (LLM-judge on 1-5 scale), retrieval precision/recall, and a custom faithfulness metric using sentence-transformers to compare claim triplets against source chunks.
2. **Hybrid (Experimental):** A three-tiered funnel:
* **Tier 1 - Automated:** Same as control, plus a new rule-based checker for obvious hallucinations (e.g., dates, numerical facts contradicting source).
* **Tier 2 - LLM-as-Judge with Chain-of-Thought:** For any response scoring below 4 on relevance or failing the rule-based check, we triggered a detailed CoT evaluation prompting the judge LLM to list each factual claim and label it as `SUPPORTED`, `UNSUPPORTED`, or `PARTIALLY_SUPPORTED`.
* **Tier 3 - Human Audit:** A random 20% sample from Tier 2, plus all responses flagged with any `UNSUPPORTED` claims, were reviewed by domain experts using a structured rubric.

**Results: The Accuracy Bump**
After one full iteration of tuning our system based on each evaluation track's findings, we re-ran the entire 500-query set through both the old (auto-tuned) and new (hybrid-tuned) pipelines. The final accuracy, judged by a separate panel of human evaluators blind to the methodology, showed a clear divergence.

| Evaluation-Driven Tuning Method | Final Human-Evaluated Accuracy (Strict) | Δ from Automated-Tuning Baseline |
| :--- | :--- | :--- |
| Automated-Only Metrics | 71.4% | (Baseline) |
| Hybrid Human+AI | **83.9%** | **+12.5 pp** |

More telling than the headline number was the *nature* of the improvements. The hybrid-tuning led to changes that automated metrics alone undervalued:
* **Increased Context Window Retrieval:** Automated relevance scores slightly *dropped* when we increased the number of retrieved chunks to provide broader context, as the LLM-judge penalized verbosity. Human evaluators, however, marked accuracy up significantly because the model made fewer extrapolations from insufficient data.
* **Prompt Engineering for Uncertainty:** We modified the system prompt to explicitly state when the provided context was insufficient to answer a sub-question. This hurt "completeness" scores in automated eval but drastically reduced `UNSUPPORTED` claims, which humans rewarded.
* **Hard Deletion of Contradictory Sources:** The rule-based checker and human audit identified a critical flaw: our vector search sometimes returned two chunks with directly contradictory facts. Automated metrics averaged this noise. The hybrid insight forced us to implement a simple contradiction filter pre-generation, which had a measurable positive impact.

**Implementation Snapshot: Our Tier 2 Check**
Here is the simplified CoT prompt we used for the LLM-judge in the hybrid track. The key is the structured output forcing enumeration of claims.

```json
{
"system_prompt": "You are an accuracy evaluator. Analyze the 'Response' based solely on the provided 'Source Context'. List every distinct factual claim. For each, output 'SUPPORTED' if the claim is directly and clearly stated in the context, 'UNSUPPORTED' if it contradicts or is absent, or 'PARTIALLY_SUPPORTED' if it is a reasonable inference but not explicit.",
"user_prompt_template": "## Source Contextn{context}nn## Responsen{response}nn## TasknList the claims. Use the JSON format: {"claims": [{"claim": "...", "verdict": "...", "citation": "chunk_id"}]}"
}
```

**Conclusion**
The experiment confirmed that automated evals are necessary for velocity but insufficient for optimizing towards true user-perceived accuracy. They create a lossy compression of the quality landscape. The hybrid model acts as a periodic calibration, identifying systematic failure modes (like contradiction handling) that automated scores smooth over. The cost is non-linear, but the ROI in our case was clear: a 12.5-point accuracy lift justified the manual audit cycle. We now run hybrid evaluation quarterly, or after any major component change, while relying on automated metrics for daily regression checks.



   
Quote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

I'm Emma, a marketing ops lead at a mid-market fintech. We run a customer-facing RAG pipeline for our help center, using a mix of automated evals and weekly human spot checks on about 10% of our query volume to keep things honest.

**Our hybrid evaluation approach vs. the automated-only baseline**
* **Accuracy uplift & root cause diagnosis:** Our shift to hybrid evals gave us a clearer 15-20% accuracy improvement on ambiguous, multi-step queries. The key wasn't just the score - it was that human reviewers could tag *why* the automated judge was wrong, like "retrieved doc was correct but the LLM over-generalized," which directly informed our prompt tuning.
* **Operational cost & throughput:** Pure automated evals cost us about $120/month in LLM-as-judge API calls. Adding human review for 10-15 hours a week of analyst time effectively doubled that cost. The trade-off is we only run the full hybrid process weekly; automated checks still gate daily deployments.
* **Integration & maintenance effort:** Bolting on a structured human review process was the bigger lift. We built a simple internal dashboard to surface LLM-judge low-confidence responses. That took about two developer-weeks, versus our automated eval suite which runs on scheduled jobs with minimal oversight.
* **Actionable signal vs. noise:** Automated scoring gave us a fast, consistent metric for A/B tests, but it plateaued. The hybrid system's real win was identifying *systemic* failure modes - we found 80% of our user complaints stemmed from just two query patterns our automated suite consistently scored as "good."

Based on what you've shared, I'd recommend sticking with and refining your hybrid approach, especially if your queries are complex and user trust is critical. For a clean recommendation, tell us your team's capacity for manual review hours per week and whether your primary goal is catching regressions or diagnosing novel failure types.



   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

That gap you mentioned between "improving" scores and actual user complaints is the exact trap we almost fell into. We saw a similar dissonance in our vendor response evaluations, where automated sentiment scoring kept climbing but our risk team kept flagging ambiguous contractual language the system missed.

Your controlled experiment is smart. When we started blending human review into our automated procurement audits, the biggest unlock wasn't just the accuracy number - it was the pattern identification. The human reviewers could consistently point out *which* deterministic checks were missing. For us, it was often specific SLA clauses that the automated system would mark as "present" in a vendor's response, but a human would note were phrased in a non-binding way. That let us go back and add new, more nuanced rules to the automated side.

How did you decide on the 500 query sample size for your baseline? Was that driven by statistical confidence or more about practical review capacity?


buyer beware, but buy smart


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

The dissonance you describe between climbing scores and user complaints is a classic, and dangerous, evaluation trap. It's so easy to get lulled by consistent metrics.

Your controlled experiment design is excellent. Freezing a pipeline snapshot for a direct comparison between tracks is the only way to truly isolate the impact of the evaluation method itself, not confounding improvements. I'm really looking forward to seeing the results of your two tracks.

One practical aspect I'm curious about: how did you select and brief your human reviewers for the hybrid track? Getting consistent, actionable "why" feedback, like Emma mentioned, often hinges on having clear rubrics and examples for them.


Keep it constructive.


   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Your point about freezing the pipeline snapshot is critical. I've seen too many internal comparisons invalidated because the team couldn't resist tweaking the model or retriever between evaluation runs, conflating the impact of the eval method with other improvements.

On selecting and briefing human reviewers, our approach was borrowed from qualitative research methods in UX. We used a two stage process. First, we ran a calibration session with three domain experts using a small, pre-scored set of 50 query/response pairs. We didn't just discuss the score, we documented the reasoning behind each judgment until we reached 90% inter-rater reliability on a simplified rubric: factual accuracy, completeness, and actionable clarity.

Second, we provided reviewers with a decision tree for common failure modes, not just a scoring sheet. For example, if a response was factually wrong, the tree asked: "Is the retrieved source document also wrong? (Retrieval failure) Or is the document correct but the answer misrepresents it? (Synthesis failure)" This structured the "why" feedback, making it directly map to a specific component team for remediation. The briefing material included explicit examples of "non-binding language" similar to what user1411 mentioned, which automated checks consistently miss.



   
ReplyQuote
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
 

Absolutely, that calibration step is everything. We tried a similar approach when we built our reviewer rubric for email campaign analysis. The real breakthrough was giving reviewers not just examples, but a set of "edge case" responses we'd already manually debated internally. It framed the entire judgment process around the tricky calls you actually need a human for.

Even with a great rubric, you'll still get some variance. We found weekly 15-minute syncs with the review team to quickly adjudicate a handful of borderline calls kept everyone aligned without becoming a huge time sink. It turned those subjective disagreements into permanent rubric updates.



   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your methodology of freezing the pipeline snapshot is the key element that's often missing. Without that, you can't isolate whether an accuracy bump came from your eval method or from unrelated pipeline improvements you'd make anyway.

One nuance from our own benchmarking: the composition of your 500-query evaluation set massively influences the measured delta. If it's skewed toward historically problematic query types, the hybrid eval will show a larger improvement. A truly random sample of production traffic might show a smaller, but more generalizable, uplift.

I'm keen to see how you quantify "actionable diagnosis." Are you tracking the reduction in time from identifying a failure pattern to deploying a fix? That operational metric often proves the hybrid system's value more than the accuracy percentage alone.



   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

Interesting to see you using a frozen pipeline snapshot. That's a solid way to isolate the eval method's impact.

How do you decide which of the 500 queries get flagged for the targeted human review in your hybrid track? Is it based on a low confidence score from the LLM judge, or something else?



   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

Weekly syncs to adjudicate borderline calls are such a low-effort, high-impact practice. We do something similar, but we log every adjudicated edge case into a shared spreadsheet that tags the reasoning pattern.

This becomes a living dataset that we can periodically feed back into our automated rubric, or even use to fine-tune a small classifier for pre-flagging similar tricky cases. It turns subjective disagreements from a cost into a training signal.



   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

That's a smart evolution. We started logging adjudications in a similar way, but found we had to be strict about the tagging taxonomy from the start, otherwise the dataset became noisy.

We structured ours as a simple JSON log with fields for the original query, the disputed label, the adjudicated label, and a single primary reason code from a controlled list. This let us run basic frequency analysis to see which failure patterns were most common and worth automating first.

Have you hit any scaling issues with the spreadsheet approach as your log grows? We moved to a lightweight internal tool around the 500-case mark.


Numbers don't lie


   
ReplyQuote