I've been stress-testing NotebookLM on a set of API documentation (150 pages) for a code generation benchmark. I'm getting unacceptable output variance on repeated queries.
Setup:
* Same source docs (3 PDFs, uploaded once).
* Same pinned sources for the session.
* Query: "List the required parameters for the create_session method and their data types."
Results from three consecutive runs in the same chat:
1. Listed 4 parameters with correct types.
2. Listed 5 parameters, one was hallucinated (not in doc).
3. Listed 3 parameters, missing one that was in the first answer.
No changes to sources or prompts between attempts. The temperature setting isn't exposed, so I can't lock it down.
Is this a known architecture issue? The whole point of grounding is consistency. Without it, you can't trust any single answer for verification.
- bench_beast
Benchmarks don't lie.
That's a frustrating situation, especially when you're trying to use it for verification. I've seen similar chatter about output variance, even with pinned sources. The grounding seems to improve relevance over a base model, but it doesn't appear to guarantee deterministic retrieval from the document set every single time.
A few of us have wondered if it's related to how the system chunks or indexes the docs on the fly, maybe pulling from different segments each time. Without a temperature control, you're sort of at its mercy for repeatability. Have you tried the exact same query in a brand new chat? Sometimes that resets the context window in a way that changes the result, for better or worse.
Keep it real, keep it kind.