Skip to content
Notifications
Clear all

Has anyone benchmarked citation accuracy? I'm seeing some slippage.

4 Posts
4 Users
0 Reactions
22 Views
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
Topic starter   [#16797]

I've been using NotebookLM for a few months to parse vendor contracts and summarize terms. Lately, I've noticed some subtle but critical errors in the citations it's generating for my own uploaded source material.

For example, it cited a specific SLA clause from a PDF, but when I checked the source, the language was paraphrased and a key percentage was off. This is a big deal for my use case.

Has anyone else done a systematic accuracy check or benchmark? I'm trying to figure out:

* Is this a recent model regression?
* Does accuracy drop significantly with larger source corpuses?
* Any best practices to tighten it up?

Trying to gauge if I need to adjust my workflow or if this is a known issue. The TCO impact of a wrong clause is real!



   
Quote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Oof, that's worrying. I haven't benchmarked NotebookLM specifically, but I've seen similar citation drift in other tools when dealing with dense legal or financial docs. The paraphrasing is a major red flag; it shouldn't be re-interpreting key percentages.

For your question about larger source corpuses: in my experience, yes, accuracy often drops as you add more source material. The model seems to "average" concepts when the corpus gets big, leading to blended or misattributed clauses.

One thing that helped me was splitting my source material into smaller, hyper-focused notebooks by vendor or contract type, instead of one massive repository. It's a pain, but it reduced my error rate. Also, treating any cited number as "unverified until eyeballed" became my rule. 😅

Have you tried using the citation feature to jump back to the source for every single claim, even the seemingly minor ones? It's tedious, but it might show a pattern in where it's failing.


Cheers, Henry


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

I haven't benchmarked it, but you've hit on my biggest trust issue with these tools. That paraphrasing is a deal-breaker for anything contractual.

It makes me think we should treat these like a junior dev's first PR: you need a solid review process before merging any info into your decision-making. Maybe you could build a verification step into your workflow? Like, key findings from NotebookLM require a second, dumb lookup (ctrl+F in the source) before being accepted.

Sounds like a job for a CI check, but for facts. 😅 Is your material all text-searchable PDFs?


git push and pray


   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

I like that junior dev PR analogy, that's spot on. Makes me think the verification step needs to be as simple as a diff check.

> Is your material all text-searchable PDFs?

This is a good point. I've had issues even with text-searchable PDFs if the formatting is weird, like tables or footnotes. Sometimes the ctrl+F doesn't find the exact phrase because the PDF extraction messed it up. Have you seen that too?


CloudNewbie


   
ReplyQuote