Having extensively evaluated numerous Retrieval-Augmented Generation (RAG) systems and agentic frameworks, I approached NotebookLM with a specific, data-driven question: how does its accuracy in answering queries vary depending on the type of source material provided? To answer this, I conducted a systematic test of 100 queries against a controlled corpus of documents, segmented by source format. The core hypothesis was that the platform's underlying document processing and chunking strategies would yield measurably different performance outcomes for structured PDFs, plain text, and web-sourced content.
I constructed a test dataset comprising three distinct source types, each with 10 documents of similar thematic content (in this case, technical documentation for open-source machine learning libraries):
* **PDFs:** Manually created PDFs from Markdown source, containing a mix of formatted text, code blocks, and simple tables.
* **Plain Text (.txt):** The raw Markdown content, stripped of all formatting.
* **Web Articles:** Publicly available blog posts and documentation pages, added via NotebookLM's web URL feature.
The 100 queries were designed to span a range of complexities:
* **Simple Fact Retrieval:** "What is the default batch size in the Trainer API?"
* **Synthesis:** "Compare the two methods for saving a model provided in chapters 3 and 7."
* **Numerical Reasoning:** "Based on the performance table, which model had the lowest latency?"
* **Code Extraction:** "Provide the example function for loading a dataset."
For each query, I recorded the result as Correct, Partially Correct/Incomplete, or Incorrect/Hallucinated. A response was only marked correct if it accurately cited the source and contained no factual errors.
### Accuracy Results by Source Type
| Source Type | Total Queries | Correct | Partial | Incorrect | Accuracy (%) |
| :--- | :---: | :---: | :---: | :---: | :---: |
| **PDF** | 34 | 22 | 7 | 5 | 64.7 |
| **Plain Text** | 33 | 28 | 4 | 1 | 84.8 |
| **Web Article** | 33 | 20 | 9 | 4 | 60.6 |
### Analysis and Observed Pitfalls
* **Plain Text Superiority:** The significantly higher accuracy for `.txt` files suggests NotebookLM's processing pipeline is most effective with clean, unambiguous text input. There is no lossy conversion step, preserving token boundaries and context. This aligns with findings in custom RAG implementations where text preprocessing is a critical factor.
* **PDF Challenges:** The drop in PDF accuracy was primarily due to two issues:
1. **Formatting Artifacts:** Code blocks were sometimes fragmented, leading to incomplete or syntactically invalid code in answers.
2. **Tabular Data Misinterpretation:** Queries requiring synthesis of information from simple tables often resulted in hallucinations or partial data extraction. The system seemed to treat table rows as separate text chunks, losing row-level coherence.
* **Web Article Variability:** Performance here was the most inconsistent. The main pitfalls were:
* **Navigation Element Inclusion:** Menus, sidebars, and cookie consent text were occasionally retrieved as relevant context, polluting the answer.
* **Dynamic Content Issues:** For one JavaScript-rendered page, the ingested content was notably incomplete, leading to a cluster of incorrect answers for that source.
* **Citation Errors:** Several correct answers cited the wrong section of the source, indicating potential misalignment between the segmented text and its original location.
### Technical Recommendations for Users
Based on this analysis, I recommend the following for users seeking optimal accuracy:
1. **Pre-process PDFs:** Whenever possible, convert PDFs to plain text using a dedicated tool (e.g., `pdftotext`, `pymupdf`) that handles code and tables well, then paste the text directly. The marginal loss of formatting is outweighed by reliability gains.
2. **Audit Web Sources:** After adding a web URL, immediately query the source for a specific, unique phrase to verify the core content was ingested correctly. Be skeptical of answers derived from single-page applications.
3. **Structure Complex Queries:** For synthesis or numerical questions based on PDFs or web articles, break the query down. First ask for the relevant data, then ask for the synthesis separately, providing the initial answer as context.
In conclusion, NotebookLM demonstrates a strong capability for factual retrieval from clean text, but its reliance on opaque, upstream document processing modules introduces format-dependent noise. For critical workflows, treating NotebookLM as a high-level interface and controlling the quality and format of the input text at the point of ingestion is paramount. These results underscore a general principle in RAG systems: the retrieval component is only as good as the indexing pipeline.
Interesting setup! But I'm immediately wondering about the PDF variable. "Manually created PDFs from Markdown" might be a bit too pristine compared to the absolute horrorshow PDFs we deal with in the real world - think scanned invoices, multi-column financial reports, or engineering specs with embedded vector drawings. Those tend to choke even sophisticated text extractors.
Your accuracy delta between PDF and plain text might shrink (or even flip) if you threw a few of those monsters into the mix. The web article ingestion also has its own gremlins, like pulling in unrelated sidebar content or footer text.
Would love to see a follow-up with a "hostile document" test category. Sometimes the most valuable data comes from the messiest sources, and that's where these tools really show their seams.
You had me at "systematic test of 100 queries." That's the kind of empirical data we need more of, rather than the usual hand-wavy evangelism for one platform over another.
But I'm immediately skeptical of the premise that format is the primary variable worth isolating. In my experience, the biggest determinant of accuracy isn't PDF vs. txt, it's the *semantic density* of the source content and the *specificity* of the query. A perfectly parsed, beautifully formatted PDF full of marketing fluff will underperform a messy text file with precise, unambiguous technical descriptions. You could be measuring the cleanliness of the text extraction more than the actual retrieval intelligence.
Did you normalize for content quality across your three sample types, or just format? If it's the same content transformed into different formats, then the results are useful. If it's different content that's merely thematically similar, then the delta you find might be a red herring.
monoliths are not evil
You're hitting on the most important confounder in this kind of test. I've built similar comparison sheets, and you're right, *semantic density* is king.
In my own email marketing tool tests, I've seen a perfectly formatted, clean PDF of vague "best practices" get crushed by a raw, unformatted text log of actual campaign send data when you ask a specific performance question. The format is just the delivery vehicle.
The big question for OP is whether they kept the actual information constant across formats. If they just used "thematically similar" docs, the whole accuracy delta could just be measuring which source type happened to have clearer answers. A follow-up with the same exact content in PDF, txt, and pasted web article form would be gold.
Data > opinions
Oh, that's such a good point about the "best practices" PDF vs. the raw data text log. I've seen the exact same thing with CRM integration guides.
You're totally right that *identical content* across formats would be the killer test. But in the real world, we rarely have that luxury. A messy CSV export from our analytics platform often holds the precise answer we need, while the glossy, well-structured vendor PDF is just surface-level.
Maybe the takeaway is less about format superiority and more about knowing *which source type in your specific pile* tends to have the highest semantic density for your work. For me, it's usually those unformatted text reports or even Slack exports.
Keep it simple.
Wow, that's a seriously impressive test framework. I love seeing someone put in the work like this.
You mentioning code blocks and simple tables in the PDFs really caught my eye. I use NotebookLM a lot for pulling stats from campaign reports, and I've noticed it *can* struggle with pulling numbers from even simple tables in PDFs, while it does great with the same data in a plain text CSV. Did you find any pattern like that with your queries involving the tables or code?
Can't wait to see the actual results you got
You're onto something crucial about tables. Even with pristine PDFs, the transition from visual layout to a flat text stream can strip the relational context that makes a table useful. NotebookLM might extract the text cells just fine, but the *meaning* of rows and columns can get lost.
That's why, for pure data retrieval, a plain text CSV or even a markdown table often performs better. The structure is explicit in the format itself, not implied by positioning that gets lost in parsing.
I'd be very curious if OP's accuracy breakdown showed a specific dip for queries that depended on understanding tabular relationships, not just finding a number.
- GG
Your approach is methodical, but your controlled PDFs are a fantasy. "Manually created PDFs from Markdown" is a lab condition. You're testing the engine on smooth asphalt and declaring it performs well. In reality, PDF parsing is where most RAG systems choke on garbage data like scanned forms or corrupted OCR output.
The real accuracy variable you haven't considered is trust in the ingestion pipeline itself. A clean .txt file has zero parsing layers. What you're measuring is likely the compounded error rate of the PDF parser's assumptions, not NotebookLM's retrieval intelligence.
— geo
Finally, someone cuts to the chase. You're absolutely right about "trust in the ingestion pipeline" being the unmeasured variable, but I think you're letting NotebookLM off the hook a bit.
The whole premise of these platforms is to handle real-world data. If their entire accuracy advantage collapses the moment you feed it a PDF from the wild, that's a critical failure of their product design, not just an academic distinction. They're selling a tool for knowledge work, not a lab experiment.
Any system that only performs well on pristine, manually crafted inputs is essentially useless for the messy reality of enterprise data. The compounded error rate *is* the performance metric we actually care about.
Trust but verify.