Skip to content
Notifications
Clear all

Troubleshooting: My LLM judge agrees with humans only 60% of the time. Is that bad?

3 Posts
3 Users
0 Reactions
20 Views
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
Topic starter   [#24436]

I'm setting up an evaluation pipeline for our help desk's internal LLM. I'm using a fine-tuned GPT-4 as an automated judge to score responses against our ITIL-based criteria (accuracy, clarity, actionability).

My initial results show the LLM judge's scores only agree with human expert ratings about 60% of the time (looking at exact score matches). The correlation isn't much higher.

I'm new to this. Is 60% agreement considered acceptable for moving forward, or is it a sign my judge prompt or criteria need a major overhaul? Should I be aiming for a specific benchmark before trusting the automated scores?

What are others seeing for human-judge alignment in similar service desk contexts?



   
Quote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

60% exact match is a solid starting point, especially with multi-criteria scoring. I'd be more concerned if your correlation was low too.

Consider looking at "agreement within one point" or binary "pass/fail" alignment rather than exact matches. Human raters often don't perfectly align either. The key is whether the judge reliably flags the truly bad or excellent responses for human review.

For a service desk, actionability is often the most critical and surprisingly subjective. You might need to calibrate that rubric with more concrete examples before your judge or your humans can score it consistently.


Keep it constructive.


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Agree on looking at binary pass/fail. For operational use, that's often the only decision you need to automate: "Does this require a human to step in?"

You mentioned actionability being subjective. That's exactly where prompt engineering can fall short without a solid retrieval layer. If the judge doesn't have explicit, canonical examples of "actionable" vs. "not actionable" steps to reference from your knowledge base, it's just guessing based on its training. Consider embedding your rubric examples and using RAG to inject them into the judge's context.

Human inter-rater reliability for something like ITIL scoring might not be much better than 60-70% exact match. The real test is if the judge's "fail" bucket contains most of the human-flagged fails, even if the specific score is off.


sub-100ms or bust


   
ReplyQuote