Skip to content
Notifications
Clear all

TIL: You can use GPT-4 as a judge, but you need to control for its verbosity bias.

4 Posts
4 Users
0 Reactions
28 Views
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
Topic starter   [#19380]

I was reading a paper on using LLMs as judges for things like summarization or code generation, and a point really stuck with me. Everyone knows about bias towards a certain model's own outputs, but they specifically mentioned a **verbosity bias**. Basically, GPT-4 tends to favor longer, more detailed answers when asked to score them, even if the shorter one is more concise and correct.

This makes sense because in its training, thoroughness is often rewarded. But for a real evaluation, it's a problem. If you're comparing a model that generates one-sentence answers to one that writes three paragraphs, the judge might unfairly penalize the concise one.

So, how do you control for it? The paper suggested a couple of things you can bake into your judge prompt. I tried a simplified version for comparing two SQL query explanations.

Here's a basic judge prompt **without** controlling for verbosity:

```sql
-- System prompt for GPT-4 as judge
You are an expert data analyst evaluating SQL explanations. Which explanation is more accurate and helpful? Output only "A" or "B".

Explanation A: {explanation_a}
Explanation B: {explanation_b}
```

And here's one **with** a control:

```sql
-- System prompt for GPT-4 as judge
You are an expert data analyst evaluating SQL explanations. Consider accuracy and helpfulness. **Do not favor longer or shorter explanations.** A concise, correct explanation is better than a verbose one with fluff. Output only "A" or "B".

Explanation A: {explanation_a}
Explanation B: {explanation_b}
```

Just that explicit instruction seems to help steer it. They also mentioned more rigorous methods like asking the judge to rate answers on specific, separate criteria (accuracy, conciseness, completeness) on a scale, and then combining the scores. That way, verbosity only affects the "completeness" part, not the overall score.

Has anyone else tried implementing something like this? I'm thinking of using it to evaluate some chart captions generated by different models, but I'm worried about other hidden biases too. What other biases should we be controlling for in these LLM-as-judge setups?



   
Quote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

The verbosity bias is real, but frankly, the entire concept of using an LLM as a "judge" feels like adding another layer of unreliable abstraction on an already shaky stack. You're basically using a black box to evaluate other black boxes.

Even with a clever prompt, you're just trying to correct one bias you know about. What about the dozen you don't? It's like trying to calibrate a thermometer you found in a junk drawer.

For something like SQL explanations, you'd get more signal by just running the dang queries and comparing outputs. Or, you know, having an actual human glance at them for five seconds.


null


   
ReplyQuote
(@gracel)
Reputable Member
Joined: 3 months ago
Posts: 227
 

I get the black box worry, it's a fair point. But in my work, sometimes you have to evaluate hundreds of email variations or CRM reply templates for tone before they go out. Running an A/B test on all of them takes weeks and a real human glancing at each one for five seconds isn't scalable either.

Isn't the real issue whether the judge's bias is *consistent*? If you learn its quirks and it always applies them the same way, you can still get useful ranking data internally, even if the scores are slightly off.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

You're right that consistency matters for internal ranking. The bias becomes a manageable systematic error, like a scale that's always two pounds off. You can still sort items correctly even if the absolute scores are skewed.

However, this falls apart when you try to compare two separate evaluation runs, or if the judge model's behavior drifts between API updates. The "verbosity penalty" might be consistent today, but next month's GPT-4-Turbo could weight it differently. Your internal benchmark is now unanchored.

A practical mitigation is to include a fixed set of gold-standard reference examples in every judging batch. You calibrate against those known-good answers each time, checking if the judge's relative scoring holds. It adds overhead, but it catches drift.


benchmark or bust


   
ReplyQuote