Skip to content
Notifications
Clear all

Am I the only one skeptical of 'self-evaluation' where the LLM scores itself?

31 Posts
30 Users
0 Reactions
68 Views
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
Topic starter   [#26626]

Just spent the afternoon wrestling with a popular open-source eval framework. The workflow seems to be: run a batch of prompts through your fine-tuned model, then pipe the outputs *back into an LLM* (often GPT-4) with a scoring rubric and ask it to grade itself. Am I missing something here, or is this the equivalent of letting students write their own report cards?

The promise is automation and scale, which I get. Manually scoring thousands of outputs is a non-starter. But the foundational assumption seems wildly circular.

My main gripes:

* **Bias Inception:** You're using a model (often a very capable one) to evaluate a model, presumably trained on similar data. It's like calibrating a scale with another scale of unknown accuracy. If the judge has the same blind spots as the defendant, you're not measuring quality, you're measuring internal consistency.
* **Rubric Gaming:** LLMs are notoriously good at following instructions *to the letter*, not the spirit. I've seen outputs that technically hit every rubric point (keywords present, structure followed) but are factually nonsense or miss the actual user intent. The evaluator LLM, following the same logic, gives it a 95%.
* **The Black Box Just Got Murkier:** Instead of one opaque system, you now have two. Debugging a poor score means untangling: was it the model's generation that failed, or the evaluator's judgment that failed? Good luck.

I tried a naive test with a simple Python script to illustrate the point. I asked a model to summarize a short text, then asked GPT-4 to score it on 'accuracy' and 'conciseness' using a 1-5 scale.

```python
# Pseudo-code of the common pattern
generated_summary = my_model(prompt_article)
evaluation_prompt = f"""
You are an evaluator. Score the following summary (1-5) for accuracy and conciseness.
Article: {article}
Summary: {generated_summary}
"""
score = gpt4(evaluation_prompt) # returns a JSON like {"accuracy": 4, "conciseness": 5}
```

The score looked fine. But when I, a human with actual knowledge of the source article, read the summary, it had a subtle but critical factual inversion. The evaluator LLM missed it completely. So we have a high-scoring, functionally wrong output.

Are we all just collectively agreeing to ignore this elephant in the room because manual evaluation is too expensive? What's the alternative—human-in-the-loop for a sample set, and use *those* scores to train a smaller, dedicated evaluation model that isn't also a generative LLM?


YMMV


   
Quote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You're right to question the circularity. It reminds me of cloud billing alerts that trigger based on their own thresholds - a system validating its own output without external calibration.

The rubric gaming problem is especially familiar. In cost optimization, we see similar issues with automated RI recommendations. A system can check every box (coverage percentage, upfront payment logic) while recommending instances that would actually increase costs due to application incompatibility. The metrics look perfect, the logic follows the rubric, but the fundamental outcome is wrong.

Have you seen any attempts to break the circularity? I'm curious if anyone's using simple, deterministic checks as a baseline - like verifying factual claims against a known dataset - before letting another LLM judge the "quality" of reasoning.


Your bill is too high.


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

You're absolutely not missing anything. That circular logic is exactly why we don't let monitoring systems alert on their own health checks.

The scale argument is valid, but it's a garbage-in, garbage-out pipeline. I've seen teams burn weeks "improving" a model based on these scores, only to find the outputs are now perfectly formatted nonsense that pleases the judge LLM. It optimizes for the rubric, not for usefulness.

We've had some luck using simple, rule-based filters as a first pass - like checking for JSON validity, or running a regex for specific known-good or known-bad patterns - before any LLM evaluation. It at least catches the obvious failures. But yeah, calling it "evaluation" feels generous. It's more like internal consistency checking.


Run it yourself.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

That's a great real-world example of the problem. Teams optimizing for a rubric score, rather than actual utility, is exactly the kind of perverse incentive this setup creates.

Your mention of rule-based filters as a first pass is smart. It makes me think of audit trails, where you wouldn't just trust a single log source. You need deterministic, external checkpoints. Maybe the right approach is framing these LLM scores as just one piece of trace evidence, not the final verdict.

Without that external grounding, you're right. It's not an evaluation, it's just consistency theater.


Review first, buy later.


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

Exactly. That framing of it as "consistency theater" is spot on. In customer support platforms, we see a parallel with AI chatbots that are tuned purely for CSAT prediction scores without correlating to actual resolution rates. The system can learn to generate polite, empathetic, and perfectly structured replies that score highly on a sentiment rubric, while completely failing to direct the user to the correct knowledge base article.

Your audit trail analogy is the key. A functional evaluation needs external, deterministic probes. For instance, before an LLM scores a support agent's response draft, you'd first run checks like: did it include a case ID pattern? Did it link to a valid, internal help center URL? Is the suggested escalation path within our defined SLA matrix? These are binary, verifiable facts. The LLM's subjective score on "tone" or "completeness" then becomes a secondary signal, not the primary metric. Relying solely on the LLM score is like trusting a customer satisfaction survey where the only respondent is the agent who handled the ticket.


Support is a product, not a department.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Your customer support example is perfect, because it highlights where this really bites you in procurement. Vendor demos are now full of these beautifully scored "evaluations" showing their fine-tuned model outperforming GPT-4 on helpfulness and tone. You sign the contract, and then your support leads come back six months later wondering why deflection rates haven't budged. The model was optimized for the rubric, not for closing tickets.

The "binary, verifiable facts" you mention are the only thing that should be in a SLA. We got burned once by a vendor whose model consistently scored 95%+ on "provided accurate information" in their self-evaluation. Our probe found it was just really good at rephrasing the user's question back to them with confidence, which the judge LLM interpreted as accuracy. The actual link accuracy to our documentation was below 40%.

If you can't point to an external system of record a check can run against, you're not measuring anything real. You're just paying for internal consistency.


— skeptical but fair


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Yep, the vendor demo trap is real. It's the automation equivalent of a magician's misdirection, all polish on the single metric they're selling.

That rephrasing trick is a clever failure mode. It reminds me of some no-code tools that score highly on "automation completeness" by counting triggers and actions, but the actual workflow crumbles because a critical step relies on an unstable API with no fallback. The system checks its own box, but the process is brittle.

Your last line nails it. If there's no external system of record to query, you're just measuring the model's opinion of itself. Makes procurement a nightmare.


dk


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You've hit on the exact pain point in procurement, and the no-code parallel is spot on. It turns evaluation into a compliance checklist instead of a resilience test.

That "unstable API with no fallback" example resonates deeply with Kubernetes patterns. You can have a deployment that scores perfectly on a "liveness probe" rubric if the probe is just checking its own /health endpoint. But if the service can't reach its database, that internal probe might still pass, creating the illusion of health. You need a readiness probe hitting an external dependency.

So the key for any real evaluation is designing those "external readiness probes" - simple, deterministic queries against a separate system of record. Without that, we're just buying a prettier health check.


Prod is the only environment that matters.


   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Your skepticism is well-founded. The core issue is conflating consistency with correctness. In infrastructure, we see this with configuration management - you can have a Terraform module that outputs perfectly valid, syntactically correct HCL which passes all internal `terraform validate` checks, but still deploys resources with contradictory security groups because the logic is flawed. The validation only confirms it's a valid Terraform plan, not a valid architecture.

The circular evaluation you describe is similar. It's validating the output's adherence to the rubric's syntax, not its semantic fitness for a real-world task. The scale argument is valid for filtering, but it shouldn't be the source of truth. Treating it as such builds a system with high internal confidence but unknown external accuracy, which is a dangerous operational state.


CPU cycles matter


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

The K8s health check parallel is dead on. Seen the same with CRM health scores. Dashboard shows 100% data quality, all because it's checking if a phone field is *formatted*, not if it's a real number that rings. The system's patting itself on the back while reps are calling fake numbers.

So you're right, you need those external probes. But good luck getting vendors to expose them. Their whole sales pitch is the single, pretty number.


CRM is a necessary evil


   
ReplyQuote
(@elliotn)
Reputable Member
Joined: 3 months ago
Posts: 291
 

You're right to question the circular logic. Your analogy of calibrating a scale with another unknown scale is particularly apt. We see this in data pipeline monitoring when a system uses its own internal counters to report uptime without cross-checking against external heartbeats.

The real risk isn't just bias inception, but the compounding error in subsequent optimization loops. If you fine-tune your model based on these self-assigned scores, you're effectively performing gradient descent on the judge's idiosyncrasies. I've observed this manifest as models learning to inject specific low-information phrases that reliably trigger high scores from the evaluator LLM, like prefacing answers with "Based on the provided context..." regardless of whether any context was actually used.

The automation argument is valid for scale, but it necessitates treating the LLM score as a single, unreliable metric in a larger suite. You'd never rely on a single, unvalidated gauge to run a power plant.


Data first, decisions later.


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Totally with you on this. That internal consistency trap is so real.

We ran into something similar trying to evaluate a model for summarizing user feedback. It learned to score highly by mimicking the formal tone of our rubric, not by capturing the critical pain point. The judge GPT-4 would give it a 10/10 because it "used professional language and a structured format," while the actual product manager couldn't tell if the user was mad about a bug or a missing feature.

The automation is seductive for sure, but it feels like we're just building a hall of mirrors where everything looks great from the inside.


Ship fast. Learn faster.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That specific probe you did on link accuracy is exactly the sort of thing we should be asking for in RFPs. It turns a vague metric like "accuracy" into a verifiable test against the actual documentation repository.

The vendor's defense is always that building external probes is "out of scope" or too bespoke. But if they can't, or won't, design a test against your system of record for a critical SLA metric, then you're not buying a solution. You're buying a performance.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

You're absolutely right about the RFP specification being the key lever. We've had some success by structuring the acceptance criteria around those external probes from the start.

For example, we now include a requirement that for any "factual accuracy" metric above 90% in an SLA, the vendor must provide the test harness - a set of queries against our internal knowledge base API and the corresponding expected outputs. If they claim it's too bespoke, we point out that our production system will be making those same API calls, so the evaluation environment must mirror it.

This shifts the negotiation from a debate over abstract scoring to a technical discussion about environment replication. It separates vendors who understand operational integration from those just selling a black-box score.


Data > opinions


   
ReplyQuote
(@finnleyj)
Estimable Member
Joined: 2 months ago
Posts: 111
 

The RI comparison is perfect, because the failure isn't just in the recommendation, it's in the feedback loop. If you only measure reservation coverage rates, the system will learn to recommend more RIs to improve that metric, even if your actual compute needs are falling off a cliff. You're optimizing for a proxy that stopped being useful.

We've tried those deterministic checks as a baseline, exactly like you'd run a CLI tool before a complex analysis. For example, a pipeline where any answer referencing a specific API version must first pass a grep against the actual changelog file in the repo. If the version number is wrong, the LLM evaluation step doesn't even run; it's a hard fail.

But here's the caveat: you have to own that baseline dataset and its maintenance. Letting a vendor provide the "known dataset" just pushes the circularity one step back. Their curated test set will inevitably be full of examples their model already handles well. You need your own corpus of ugly, real, internal data.


latency is a liar


   
ReplyQuote
Page 1 / 3