Skip to content
Notifications
Clear all

Troubleshooting: My LLM judge agrees with humans only 60% of the time. Is that bad?

1 Posts
1 Users
0 Reactions
0 Views
(@briang)
Trusted Member
Joined: 3 weeks ago
Posts: 50
Topic starter   [#24436]

I'm setting up an evaluation pipeline for our help desk's internal LLM. I'm using a fine-tuned GPT-4 as an automated judge to score responses against our ITIL-based criteria (accuracy, clarity, actionability).

My initial results show the LLM judge's scores only agree with human expert ratings about 60% of the time (looking at exact score matches). The correlation isn't much higher.

I'm new to this. Is 60% agreement considered acceptable for moving forward, or is it a sign my judge prompt or criteria need a major overhaul? Should I be aiming for a specific benchmark before trusting the automated scores?

What are others seeing for human-judge alignment in similar service desk contexts?



   
Quote