Skip to content
Notifications
Clear all

Is anyone else's NotebookLM hallucinating more after the last update?

2 Posts
2 Users
0 Reactions
0 Views
 annt
(@annt)
Estimable Member
Joined: 7 days ago
Posts: 71
Topic starter   [#14344]

I've been conducting a systematic evaluation of NotebookLM's performance as part of my standard vendor security review workflow, specifically using it to cross-reference control frameworks (like mapping ISO 27001 Annex A to SOC 2 criteria). Following the most recent platform update, I've observed a marked and quantifiable increase in what we would classify in audit terminology as "factual misrepresentation" — colloquially, hallucinations.

Previously, the tool was reasonably reliable for synthesizing information from provided source materials, even if its inferences required careful validation. My current tests, however, show a significant regression. For instance, when I uploaded the official ISO 27001:2022 standard PDF and asked it to list the controls in Annex A that directly relate to cryptographic controls (A.8.24), it began inventing control identifiers and descriptions that do not exist in the source. It stated, with high confidence, that "A.8.24.5 mandates quarterly key rotation for all symmetric algorithms," which is a complete fabrication. The source material contains no such sub-control. This is not a minor inaccuracy; it fundamentally compromises the tool's utility for compliance analysis.

My testing methodology, which I recommend others replicate to validate these findings, involves:
* Selecting a well-defined, immutable source document (a published standard, a known-good compliance checklist, or a company security policy).
* Asking specific, extractive questions that require verbatim or precisely paraphrased responses from the source.
* Then asking synthesis or cross-mapping questions that require logical inference *within the boundaries of the source*.
* Documenting every instance where the output cannot be traced back to the provided source context.

The degradation appears most pronounced in scenarios involving:
* Synthesis across multiple, complex documents.
* Requests to extrapolate or predict based on provided guidelines.
* Any query where the source material might be slightly ambiguous; the model now seems to fill gaps with plausible but incorrect "best guesses."

This raises serious concerns for using NotebookLM in any controlled workflow, particularly for compliance evidence gathering, policy gap analysis, or drafting statements of applicability. The risk of incorporating a confidently-stated falsehood into a work product is now unacceptably high. I am curious if other members engaged in similar technical or procedural work have observed this trend, and if so, what mitigation strategies—beyond complete manual verification—you are employing. Have you adjusted your prompting techniques, or have you found certain document types or use cases to be more resilient than others?


—at


   
Quote
(@auditlog)
Estimable Member
Joined: 3 months ago
Posts: 130
 

That's a concrete and worrying example. I perform similar checks using CloudTrail logs and vendor audit reports, where accuracy is non-negotiable. Your test case is telling because it's a deterministic source: the standard is a static PDF. The tool isn't just misinterpreting nuance, it's generating entirely new normative statements.

I've noticed a related pattern when asking it to map between frameworks, like PCI DSS requirements to specific SOC 2 common criteria. Post-update, it's more likely to insert plausible-sounding but unsupported "bridges" that aren't in the provided guidance documents. It feels like the underlying model is now prioritizing linguistic coherence over strict source adherence, which is a major regression for audit evidence work. Have you tried replicating this with other static sources, like a company's own security policy document, to see if the behavior is consistent?


Logs don't lie.


   
ReplyQuote