Hey everyone, I'm new to building RAG pipelines with LlamaIndex and finally have a basic prototype working. It's pulling info from my docs and generating answers, which feels great!
But now I'm stuck on the next step. How do I actually know if the answers are *good*? I've heard terms like "faithfulness" and "relevancy" thrown around, but what metrics should a beginner look at first? Are there any simple, practical ways to test this without a huge evaluation framework? Any tools or libraries you'd recommend for getting started? 😅
For a first pass, focus on answer correctness against your source documents. Write 10-15 queries with known answers from your docs. Have your pipeline run them and manually check if the answers are factually correct and grounded in the retrieved text. This gives you a baseline.
If you want to automate, libraries like Ragas or LlamaIndex's own eval modules can compute metrics like faithfulness (answer grounded in context) and answer relevancy. Start with faithfulness, as hallucinations are a primary failure mode.
Don't get bogged down in a full framework yet. The manual check often reveals the biggest issues with retrieval or prompting.
The manual check is a solid starting point, I'll give you that. But "write 10-15 queries with known answers" is where this advice starts to unravel.
You're just testing what you already know. This doesn't capture the messy, ambiguous real-world questions you'll actually get. It's a performance test, not an evaluation of whether the system can handle novel inquiries faithfully. The moment a user asks something not in your curated set, your baseline is useless.
And starting with Ragas or any automated metric before you've even looked at retrieval failure rates is putting the cart before the horse. If your retriever pulls junk, all the "faithfulness" scoring in the world won't save you.
Show me the TCO.
I get your point about testing known queries being limited, but that's exactly where you start. You need a ground truth to even have a conversation about metrics. The curated set isn't meant to be the whole test suite, it's the control group.
You're right that retrieval failure is a huge issue, which is why those initial manual checks should include looking at the retrieved chunks. If the retriever is pulling junk, you'll see it immediately when you compare the answer to the provided context. That tells you to fix your embeddings or chunking before you even think about answer scoring.
The real gap is what happens after that baseline. You need a process to continuously gather those "messy, ambiguous" real user questions, log the pipeline's inputs and outputs, and manually review a sample. That's how you find the novel failure modes and expand your evaluation set. Without the baseline, you have no way to measure if those fixes actually improved anything.
Automate everything. Twice.