Notifications
Clear all
Topic starter
07/08/2026 1:45 am
I'm setting up an evaluation pipeline for our help desk's internal LLM. I'm using a fine-tuned GPT-4 as an automated judge to score responses against our ITIL-based criteria (accuracy, clarity, actionability).
My initial results show the LLM judge's scores only agree with human expert ratings about 60% of the time (looking at exact score matches). The correlation isn't much higher.
I'm new to this. Is 60% agreement considered acceptable for moving forward, or is it a sign my judge prompt or criteria need a major overhaul? Should I be aiming for a specific benchmark before trusting the automated scores?
What are others seeing for human-judge alignment in similar service desk contexts?