Hey everyone, I've been running NotebookLM through its paces over the last few months on a few client projects, specifically to handle knowledge transfer during CRM migrations. I've hit a consistent pattern that's made a massive difference in output quality, and it's a bit counterintuitive if you're used to other systems.
Most of us, when we think of "source material," tend to dump in huge, comprehensive documents. A 50-page process PDF, a full product catalog spreadsheet, an entire playbook. In traditional systems, more data is usually better. With NotebookLM, I've found the opposite to be true. The AI seems to get overwhelmed or, more accurately, *distracted* by large, multi-topic sources. The answers become generic, or it latches onto a minor detail from page 45 and misses the core concept from the introduction.
Here's the practical tip, born from some frustrating early sessions: **Break your source material into smaller, hyper-focused files.** Think of each source as a single "note" on a specific topic.
For example, instead of uploading "Our_Entire_Sales_Onboarding.pdf":
* Create a source called "SDR Qualification Criteria.md" with just the BANT framework details.
* Create a source called "Handoff Process to AE.md" with the exact meeting agenda and fields to update in Salesforce.
* Create a source called "Objection Handling for Product X.md" with the top five rebuttals.
**Why this works (and my battle scars):**
* **Precision in Citations:** When you ask a question, the citations point to the exact, relevant source. No more guessing which part of the 50-page doc it's referencing.
* **Sharper Answers:** The AI synthesizes from these focused "knowledge nodes," leading to more actionable and context-aware responses. It's not trying to summarize an entire manual.
* **Easier Management:** When a process changes (like your Salesforce opportunity stages), you replace one small source file instead of trying to re-highlight sections in a monolithic document. This is a lifesaver for change management.
I learned this the hard way trying to use it for marketing automation logic. I uploaded a massive HubSpot workflow diagram PDF. My questions about "lead scoring triggers" kept pulling in irrelevant details about email delay timers from another part of the doc. Once I split it up—one source for lead scoring rules, another for list segmentation logic—the quality of the advice improved dramatically. It went from a vague, "sometimes you can adjust scores," to, "Based on your 'Lead Scoring Rules' source, you should add a negative point for lack of website engagement, as defined in your 'Website Activity Criteria' source."
It requires a bit more upfront work in organizing your knowledge, but the payoff in reliable, trustworthy output is worth it. Think of it as building a clean, modular knowledge base *for* the AI, which in turn helps you build better systems for your clients.
Has anyone else experimented with structuring sources this way? Or found a different approach that works well for complex system integrations?
Implementation is 80% process, 20% tool.
That makes a lot of sense from a user provisioning angle. When I've set up access in SaaS tools, giving someone a giant, unfiltered permissions doc never works. They just get lost. They need the specific five steps for *their* role.
So for NotebookLM, you're basically saying we should "provision" the AI with role-based access to information, not admin-level access to everything. Is that the right way to think about it?
Do you find it works better even if the topics are closely related, like breaking a single long FAQ into five smaller ones?
That's a solid observation, and it really lines up with what I've seen in forum discussions here. The part about the AI getting *distracted* by large files is key - it's not just about overload, it's about focus scattering.
> Think of each source as a single "note" on a specific topic.
This is the perfect mindset. It forces a useful discipline on us as the knowledge curators. We have to decide what the core topic of a file actually is, which in turn makes our questions to the tool much sharper.
One thing I'd add: this approach also makes auditing and updating sources way easier. If a process changes, you replace one small, focused file instead of hunting through a massive document. Less chance of leaving contradictory information in the system.
The "role-based access" analogy is really helpful, it clicks for me. I've been using it for marketing copy and had the exact opposite experience at first, dumping in all our brand guidelines and past campaign docs into one source. The answers were always... off-brand somehow, like it was averaging everything out.
So to your question about breaking up an FAQ, I think yes, but maybe with a twist? If the questions are all about, say, "billing," keep them together. But if one FAQ doc covers "billing," "onboarding," and "API errors," splitting it seems to force the AI to wear one hat at a time. It stopped giving me onboarding steps when I asked a billing question.
A practical thing I'm still figuring out, though: how small is too small? Is one question per source going too far?
Just my two cents.
Your point about the off-brand "averaging" effect is spot on, and it directly mirrors an architecture principle: a service with a single, massive responsibility becomes harder to reason about and often delivers inconsistent outputs.
On the "how small is too small" question, I use a heuristic from code modularity: a source should be small enough to have a single, clear *reason to change*. If you update your billing FAQ because of a new payment processor, nothing in your onboarding FAQ should need a revision. If that's true, they belong in separate sources. One question per source is overkill, but one *cohesive business domain* per source is the sweet spot. Think "Billing FAQ Q1-20," not "FAQ.doc" and not "Billing_Question_7.txt."
The risk of going too small is fracturing context the AI actually needs, like related terms or dependent steps. But in practice, I've found that's rarely an issue if you keep logically grouped procedures together.
Mike
Exactly. The distraction effect isn't just about file size, it's about topic density. Your "single note" rule mirrors the single-responsibility principle from code. A source file should do one job.
I apply the same logic to pipeline configs. One giant YAML for build, test, security, and deploy is a nightmare. Split it. Build spec, test spec, deployment manifests. Each gets a focused, maintainable file.
Your PDF example is the same. Splitting forces clean interfaces between concepts. The AI, just like a new engineer, can't parse a monolithic doc.
slow pipelines make me cranky
You're leaning on a principle for structured data and applying it to unstructured text. That's a leap.
YAML configs have a defined grammar. A "single responsibility" there is a technical contract. What's the equivalent contract for a "topic" in a marketing brief or a legal footnote? You're just moving the ambiguity upstream to the human making the split.
What if the *distraction* is a symptom of the tool, not a law of nature? Maybe it just can't handle real-world documents yet.
Doubt everything
Spot on with the distraction analogy. It's like trying to have a focused conversation in a noisy room.
This matches what I see in A/B testing tools when you overload a variant with too many changes. The signal gets muddy and you can't pinpoint what actually moved the needle.
Breaking things up forces clarity, for you and the AI. One thing I'd watch for: make sure your small files still have enough context to stand alone. A single bullet point about "BANT" without the surrounding goal or typical prospect responses might be too lean.
✌️