Using LLMs to score other LLM outputs is common. The bias towards verbose, hedging responses is a real problem. It skews benchmark results.
You need to force the judge model to pick a single, concrete answer. Don't ask "which is better?". Structure the prompt to eliminate waffling.
Example scoring prompt for a QA task:
```
You are a strict evaluator. Assess the two answers to the question based solely on factual accuracy and conciseness.
Question: {question}
Answer A: {answer_a}
Answer B: {answer_b}
You MUST output ONLY a single JSON object with the following keys:
- "winner": either "A" or "B"
- "reason": one-sentence justification based on the criteria.
No other text.
```
Key controls:
* Explicit output format constraints (JSON).
* Criteria that penalize unnecessary elaboration.
* Instructions to ignore writing style unless specified.
* Temperature set to 0.
Without this, the judge defaults to favoring the longer, more "thorough" response, even if it's less correct. It's a measurement flaw.
-dk
Trust but verify, then don't trust.
Totally agree about the JSON structure forcing a concrete choice - that's been a game-changer for my eval scripts. But I've found even with temperature 0, you sometimes get subtle bias creep. If one answer happens to mirror GPT-4's own phrasing patterns, it seems to get an unconscious boost. I now add a line like "Disregard any stylistic similarity to your own writing" as an extra guardrail.
The verbosity penalty is tricky though - what if the truly correct answer just requires more detail? I've started splitting my criteria: one score for factual completeness, another for conciseness, then weighting them. Sometimes the "waffling" is actually important nuance.
Ever try using a smaller model as judge, like Claude Haiku? Less likely to overthink, but you trade off reasoning depth.
You're right about the stylistic bias. I've seen GPT-4 consistently favor answers that use its signature transitional phrases, even when the content is equivalent.
Splitting criteria into separate scores is the right approach. It mirrors how we evaluate cloud spend: you don't just look at total cost, you separate compute, storage, and data transfer. For LLM evals, I use a weighted rubric with points for accuracy, conciseness, and source citation. It forces the judge to consider each axis independently.
I've used Haiku for simple classification tasks where speed is critical. But for any judgment requiring nuance, like assessing if a technical explanation contains a critical oversight, the depth trade-off isn't worth it. The savings in judge cost gets offset by needing manual review of its questionable calls.
Right-size or die
That JSON constraint is a solid starting point, but it's not enough on its own. I've had the judge return a perfectly formatted JSON that still contained a contradictory or nonsensical reason, like "Winner: B, Reason: Answer A was more accurate and concise."
You need to validate the output schema and the logical consistency of the reason against the declared winner. I pipe the judge's response through a second, much stricter validation step in the pipeline that checks for this. If it fails, the eval run gets flagged for manual review.
Also, for technical answers, "conciseness" can backfire. I once had it penalize a correct Helm command for including a necessary `--dry-run` flag because the other, shorter answer omitted it and was actually wrong for production use. The criteria have to be domain-specific.
Automate everything. Twice.
That's a really good point about the stylistic bias creeping in even with temperature set to zero. I've seen the same thing with answers that use certain structural phrases, like starting with "Certainly!" or framing things as a list of key points.
Splitting the criteria is smart, and it gets to the heart of what we're actually trying to measure. I think the weighting you mention is crucial - for a technical support answer, factual completeness should weigh much more than conciseness. The nuance is often the valuable part.
I haven't tried Haiku as a judge, honestly. I worry the loss in reasoning depth would mean more flagged evaluations for manual review, which defeats the purpose of automation for me. But for high-volume, low-stakes classification, I can see the appeal.
Weighting criteria sounds good in theory, but who's setting the weights? That's just another bias vector. You're trading verbosity bias for subjective rubric bias.
And if the "nuance is often the valuable part," then why are we penalizing the model for including it? You can't have it both ways. Either you value completeness and accept some fluff, or you prioritize conciseness and risk missing critical context. The weighting becomes a manual override, which defeats the point of automated judging.
Stick to judging based on a verified source of truth. For cost reports, I need a screenshot of the bill, not a model's weighted opinion on what the bill should say.
show me the bill
Exactly. The JSON constraint is a basic but essential control, similar to schema validation in a data pipeline. It forces a discrete choice, which is the first step in getting a measurable signal from an inherently fuzzy evaluation.
But you're right about it being a measurement flaw. I'd add that even with temperature zero and a structured prompt, the verbosity bias can persist if the criteria aren't operationalized. "Conciseness" is too vague. I specify a maximum sentence count or character limit in the prompt itself, relative to the complexity of the question, to make the penalty objective. Otherwise, the judge's internal definition of "unnecessary elaboration" will drift.
The real parallel here is with monitoring systems. You don't just alert on "high latency". You define the threshold, the duration, and the specific metric. Treating LLM judges the same way is what moves it from a neat trick to a reliable, if limited, tool.
Measure twice, cut once.
Great point about the output format constraint forcing a decision, that's absolutely step one. But I've run into a practical snag with the JSON instruction - sometimes GPT-4 gets a bit too creative with the `reason` field, even when the `winner` is correct.
For a recent comparison of Kubernetes ingress controllers, the judge returned:
```json
{
"winner": "A",
"reason": "B provided more specific configuration examples and correctly mentioned the newer API version."
}
```
The logic literally contradicts the declared winner! It's like it autocompletes a plausible-sounding reason based on the answer content, decoupled from the actual choice. So I've started adding a simple script to my pipeline that checks for keyword alignment in the reason. If the reason mentions "more" or "better" but the winner is A, it scans the reason for any mention of B. It's an extra validation layer, but it's caught enough of these flips to save me time.
The verbosity bias is real, but as others noted, sometimes the detail is necessary. For infrastructure-as-code answers, a concise one-liner that omits a critical `depends_on` clause is worse than a longer, correct one. Maybe the fix is to be hyper-specific in the criteria: "Penalize only elaboration that repeats existing facts or offers unrelated background."
Automate all the things.
That's a great practical point about the validation step. It's not enough to just enforce a JSON structure, you need to check the actual logic inside the box. I've seen similar contradictions.
Your example about the Helm command is spot on, and gets to a bigger issue: generic criteria like "conciseness" often break down with real-world technical content. The necessary `--dry-run` flag isn't fluff, it's a critical safety step. This suggests the judging criteria themselves need to be as specific as the answers we're evaluating, maybe even referencing specific domain requirements.
Keep it constructive.
Absolutely, and this ties directly back to the earlier point about needing a source of truth. Generic criteria fail because they lack context. If your judging prompt just says "be concise," it can't differentiate between a redundant phrase and a required flag.
The solution I've landed on is embedding domain-specific guardrails into the criteria themselves. For a security question, one criterion might be "Does the answer include the principle of least privilege?" For a cost analysis, "Does the answer specify the pricing model (e.g., on-demand vs reserved)?" This moves the judgment from a stylistic preference to a verifiable, domain-relevant checklist. You're not just checking for fluff, you're checking for missing required components.
RTFM — then ask for the audit