Okay, I'll admit it: I spent way too long mixing these two up when reading eval reports. I kept seeing "groundedness" and "faithfulness" scores and thought, "Aren't they basically the same thing? Did the LLM tell the truth or not?"
Turns out, they're measuring two distinct—but related—failure modes. Here’s how I finally got it straight in my head.
**Faithfulness** is about the LLM *staying true to the source material you gave it*. It's an eval for RAG (Retrieval-Augmented Generation) systems. The question is: does every claim in the answer come directly from the provided context/retrieved documents? If the model adds extra "knowledge" from its own training (hallucinates), even if that info is factually correct *in the real world*, it fails the faithfulness check.
* **Example:** You provide a context that says "Company X's revenue in 2023 was $5M." The LLM answers "Company X's revenue grew to $5M in 2023, driven by their new product launch in Q2." The *growth* and *product launch* parts might be plausible, but they weren't in your source. That's **low faithfulness**.
**Groundedness** is about the *real-world factual accuracy* of the LLM's statements, regardless of the source. It asks: is this statement factually correct against a trusted knowledge base (like Wikipedia, a database, or the internet)? The model's answer can be 100% faithful to your provided source, but if your source document itself contained errors, the answer will have **low groundedness**.
* **Example:** Your provided (but flawed) source document says "The Eiffel Tower is located in Berlin." The LLM faithfully states "The Eiffel Tower is located in Berlin." It's perfectly **faithful**, but it's factually wrong. That's **low groundedness**.
The practical difference for us running evals?
- You measure **faithfulness** when you care about the system not making stuff up *beyond your supplied data*. Critical for internal knowledge bases.
- You measure **groundedness** when you need the final output to be factually true in the world. Critical for customer-facing chatbots or agents taking actions.
In an ideal world, you want both scores high. But I've seen trade-offs—like a super-conservative RAG system that's highly faithful but refuses to infer obvious, correct facts not explicitly spelled out in the text, which can hurt perceived quality.
Anyone else run into this distinction in their evals? What tools are you using to measure each one? I've been bouncing between TruLens and Phoenix, but always open to simpler, more cost-effective options.
You left out the part that matters: which one costs you money when it fails?
Groundedness failures mean you're shipping wrong facts. That gets expensive fast - support calls, reputation damage, maybe legal. Faithfulness failures mean your RAG pipeline is leaking. You're paying for a model that can't follow instructions.
Both are bad, but the business impact is different. One's a recall problem, the other's a precision problem. Which hurts more depends on your use case.
always ask for a multi-year discount
Great example, it really shows the nuance. I'd add that in a production RAG system, you often measure them at different stages.
Faithfulness is a check on the LLM's *generation* step. Did it stick to the provided context? You can evaluate this right after the response is generated.
Groundedness, on the other hand, ultimately requires an external truth source - a knowledge base, a human expert, or a verified dataset. You're checking the entire pipeline's output against the real world, which means a low groundedness score could be caused by a problem earlier in the chain, like your retrieval fetching outdated or incorrect documents, not just the LLM being wrong. So while they're separate concepts, a groundedness failure often prompts you to check your faithfulness metrics to see where the breakdown started.
yaml is my native language