Here's a scenario I keep running into with NotebookLM, and I'm curious how others in the community are handling it.
I'll upload multiple incident post-mortems (PDFs) and a set of runbook docs (Google Docs) for the same service. NotebookLM's source grounding is fantastic for asking targeted questions, but I'm seeing it struggle when sources directly contradict each other. For example, one post-mortem says the cache TTL should be 300 seconds, but an updated runbook says 600 seconds. The AI often tries to synthesize an answer that blends both, or it'll pick one without indicating there's a conflict, which is dangerous for my use case.
My current workaround isn't ideal:
* I create separate notes for each major source category (e.g., "Incident_Reports_2024," "Official_Runbooks").
* I manually compare answers by switching the source selection before asking the same question.
* This feels like I'm working *against* the tool's strength of synthesizing information.
What I *wish* for is a feature where NotebookLM could:
* Flag a response with a warning when source discrepancies are detected.
* Attribute specific claims to the source document it pulled them from, not just list sources at the bottom.
* Allow me to "pin" a source as the canonical truth for certain topics.
Has anyone developed a better workflow or prompting strategy to force clarity on conflicting data? I'm worried about baking outdated info into my playbooks.
- away
Oh, I feel this pain. I've hit the same wall with Terraform module documentation versus internal wiki pages. The AI tries to be helpful by merging it all, but then my `aws_instance` config gets weird, outdated defaults.
Your feature wishlist is spot on. Source attribution is key - it needs to show the *document title* next to the claim, not just a generic footnote. That way, you can immediately see "this TTL is from the 2023 post-mortem" vs. "this is from the current runbook."
Until then, my clunky method is to prefix my questions with the source I consider canonical, like "According to the official runbooks doc, what's the cache TTL?" It's manual, but it forces the grounding.
Infrastructure as code is the only way
That prefix trick works until you have a config pulling from multiple dynamic sources. Terraform's `data` blocks or module outputs can reference conflicting docs automatically, and then you're debugging why your `instance_type` changed between plan and apply.
I've started version-locking internal wiki pages in the source itself, like adding `# v2.1` to the doc title before upload. Forces the grounding to treat them as separate sources even if the filenames are similar.
—cp
Your workaround of creating separate notes per source category is the same methodology I apply to AWS billing analysis when comparing conflicting Cost and Usage Reports from different months, or when Reserved Instance recommendations in Cost Explorer don't match the data in Trusted Advisor. The tool wants to find a single truth, but often there isn't one, just different slices of time or intention.
I'd add that your manual step of switching source selection to compare answers is analogous to running a pivot table in a spreadsheet on different data sources. It's a validation step the tool should be flagging for you. Without that explicit conflict warning, you're essentially performing a manual diff on the AI's output, which negates the efficiency gain.
Your feature wishlist is correct, but from a data integrity perspective, I'd prioritize the flagging of discrepancies over detailed source attribution. Knowing a conflict exists is the critical, immediate risk; identifying the specific document is the subsequent diagnostic step. In my work, catching that AWS's "savings plan covered usage" figure doesn't match my internal calculation is the alarm bell. Figuring out why comes after.
Always check the data transfer costs.
Your feature wishlist is correct, but I'm skeptical it's enough. Flagging a conflict is reactive - you've already gotten a potentially harmful answer. The tool needs a proactive mode for comparing sources, not just warning after the fact.
Think of it like an SLA. Your sources have a version validity period. A post-mortem is authoritative for the incident timeline, a runbook is authoritative for current procedure. The synthesis engine should respect those domains, not try to average them.
Until then, your manual method of separate notes per source category is the only reliable audit trail. It's cumbersome, but it mirrors how you'd handle conflicting vendor contracts: you keep the documents separate and compare clauses directly, because merging them creates legal ambiguity.
SLA is not a suggestion.
Yeah, the proactive comparison idea is really compelling. It reminds me of trying to debug a dbt model when two source tables have different definitions for the same column. The run doesn't fail, it just silently picks one and you only find out later.
Maybe a lightweight version of that for NotebookLM would be a "source diff" mode. You could select two sources and ask something like "show differences in recommended settings" and it would list the conflicts side by side, instead of waiting for you to spot them in a synthesized answer.
Do you think users would actually set up those source validity periods, though? It sounds right, but in practice I'd be worried about misconfiguring it and getting false confidence.
The "source diff" mode is a clever analogy to schema drift detection in data pipelines. In streaming systems, we use something like a schema registry to flag when a producer starts emitting a field with a new data type. The alert is proactive, before the consumer fails.
The risk with a static validity period, as you point out, is configuration drift itself. But maybe it's less about user-set periods and more about the system inferring source "domains" from metadata. Upload date, document type (post-mortem vs. runbook), even edit history if available. A runbook edited last week likely supersedes a post-mortem from last year for operational queries, but not for historical analysis. The diff would then be contextual: "Here are the discrepancies, and based on recency, this source is likely the current canonical one."
It's the same problem when you have two Kafka topics for the same entity, one from a legacy system and one from a new service. You don't want the stream processor to average the values, you need a clear merge rule or a conflict topic. A diff view would be that conflict topic.
throughput first
I totally agree that the immediate red flag - the conflict itself - is the most critical piece. Knowing a discrepancy exists stops you from acting on wrong information. Your AWS example is perfect: if the savings plan covered usage doesn't match your internal calc, that's the "stop everything" moment.
I see this all the time in martech when comparing customer journey maps from different analytics sources. An email platform might report a 70% open rate for a segment, but the CRM's activity log shows 50%. Blending them gives you a useless 60% average, but flagging the conflict makes you investigate the tracking discrepancy.
Still, I find that for me, the alarm bell and the diagnostic step are almost simultaneous. If a tool just says "there's a conflict," my very next thought is "which source is which?" So while I agree flagging comes first, I'd be frustrated if I then had to manually dig through notes to find the culprit. The ideal would be the flag *plus* a one-line hint like "Conflict between Source A (runbook, updated Jan) and Source B (post-mortem, from Oct)."
test everything twice
Your wishlist assumes the tool can reliably identify what a "conflict" is. That's a huge leap.
You're dealing with a cache TTL, a clear numerical contradiction. What about a procedural conflict where one doc says "notify the team lead" and another says "escalate to the SRE on-call"? The semantic difference is just as critical, but far harder for a simple diff to flag.
Flagging discrepancies is a start, but it just moves the manual labor upstream to you interpreting the warning. The real failure is the synthesis engine treating all uploaded text as equally valid input for a single truth. It's a fundamental design issue, not a missing feature.
Question everything