Skip to content
Notifications
Clear all

Has anyone benchmarked citation accuracy? I'm seeing some slippage.

2 Posts
2 Users
0 Reactions
13 Views
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
Topic starter   [#26936]

Alright, who's been stress-testing NotebookLM's citations besides me? 🧐

I've been using it to summarize internal technical docs and RFCs, and I'm starting to notice a pattern. For straightforward Q&A, it's pretty solid—it'll point you to the right source. But the moment you ask it to synthesize information from multiple uploaded documents, or infer something that's *implied* but not explicitly stated, the citation accuracy gets... wobbly.

Here's a concrete example from my last experiment. I uploaded two Terraform module docs and a design spec. I asked: "Based on the network module's outputs and the design spec's security section, what would the ingress rule look like for the app module?"

The generated rule was *logically* correct and actually a decent suggestion. But the citations? It cited a paragraph about general security principles for the specific port number, and cited the network module's output block for the CIDR concept, but the actual *synthesis* part—combining those two into a firewall rule—wasn't directly cited anywhere. That's the "slippage." It's presenting a correct-ish conclusion but can't fully backtrace its logic to the source text.

It feels like the difference between a unit test pass ("citation found in source") and an integration test flake ("citation *context* is loosely coupled"). For light research, it's fine. But if you're using this for anything that needs audit trails or precision (compliance docs, code generation prompts), you've got to double-check its work.

Has anyone else pushed it on complex, multi-document reasoning and benchmarked the results? I'm curious if this is a known limitation or just my docs being messy.

- tm



   
Quote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Yeah, noticed that too. It's great for direct extracts but fails at cross-doc reasoning. I use it for K8s manifests and see the same thing. It'll pull a CPU limit from one doc and a memory request from another, then cite them correctly. But if you ask what the total resource footprint for a pod would be, the calculation is uncited. That's the real problem for technical work.



   
ReplyQuote