I'm evaluating NotebookLM for a client who wants to analyze their Salesforce case history and HubSpot marketing collateral. The promise of a source-grounded AI is appealing, but I need to see the actual ingestion mechanics before I trust it with any real data.
What's the definitive workflow for getting your own documents into NotebookLM? I'm not interested in the marketing overview. I need the concrete steps, supported formats, and any hard limits.
Specifically:
* What file types are actually supported? PDF, Word, plain text. What about .csv or .xlsx for structured data?
* Is there a true API for programmatic source upload, or is it purely a manual drag-and-drop UI? If an API exists, I need to see the basic auth and payload structure.
* What are the realistic size limits per document and per total project? Character counts, page counts, or file size in MB.
* How does it handle source refresh? If I update a source PDF, do I need to delete and re-upload the entire thing, creating a new source ID and breaking citations?
I'll run my own tests once I have the specs. Show me the data.
Show me the query.
You're asking the right questions. The "definitive workflow" is basically a drag-and-drop UI with very little transparency on limits. Last I checked, CSV support was flaky at best, and Excel files are a no-go unless you convert them first. They call it an "AI notebook" but treat spreadsheets like a foreign object.
Your main worry should be source refresh. It's a manual delete-and-reupload process. Every time you do that, you get a new source ID, which absolutely breaks prior citations. Makes it useless for any living document or dataset.
If there's an API, it's not public. They want you in their UI. I'd test with dummy data first, because the ingestion is the least mature part of the system.
Your stack is too complicated.
That point about the new source ID breaking citations is a critical catch. It turns a simple update into a full notebook audit, which is a non-starter for any real analysis pipeline.
I've found the CSV "flakiness" usually means it gets ingested as a single text blob, not a queryable table. So your structured data loses all its... well, structure. You're better off converting those spreadsheets to markdown tables in a text file first, oddly enough. It's an extra step, but the model seems to parse it more reliably.
The lack of a public API or any versioning really does box you into a manual, one-off analysis niche. Makes you wonder how they expect this to scale beyond a simple research assistant.
Data nerd out
Good questions. I tested the PDF upload process last week with some AWS whitepapers. It works, but you're right to worry about structured data.
The supported formats are basically text, PDF, and Word docs. CSV uploads technically work, but like others said, they get treated as a text blob. I uploaded a simple 5-column CSV and the model couldn't answer a single question about "column 3". It's useless for actual data analysis. You'll need to pre-process any structured data into markdown or plain text.
On limits, I hit a soft ceiling around 500,000 characters per source. The UI didn't reject it, but the ingestion seemed to silently truncate after that point. Total project limits are murky, but performance degraded noticeably for me after adding about 10 large sources.
The source refresh problem is real. There's no versioning or replace function. You delete and re-upload, which creates a new source ID and breaks every existing citation in your notebook. It makes the tool unsuitable for any living document workflow. For your client's use case, that alone might be a deal-breaker.
Cloud cost nerd. No, I don't use Reserved Instances.
Yep, the source ID breakage is the killer. I tried to script around the lack of API using their internal calls. You can simulate an upload with a POST to their endpoint, but you still get a new ID every time. It's baked into their data model.
So you can't even hack together a versioning system. It's a dead end for automation.
Benchmarks or bust.
Tested uploads on NotebookLM 24.10.15. Results match what you're hearing.
Formats:
* PDF, DOCX, TXT work. CSV uploads but parses as a single text block, so column queries fail.
* No .xlsx support. Convert to markdown table in a text file for structured analysis.
No public API. I inspected network traffic. It's a manual UI flow with internal POST calls. You can't bypass the new source ID generation on refresh, so citations always break.
My size tests: A 480k character PDF ingested fully. A 520k character PDF had responses missing content from the last ~30 pages, indicating silent truncation. Keep sources under 500k characters.
Your use case with Salesforce and HubSpot data is a bad fit unless it's already narrative text.
Benchmarks don't lie.
Interesting, you confirmed the internal endpoint. I had a hunch they might be using a UUID as a foreign key in their citations table. If that's the case, there's no way to alias it without a full schema change.
So even a backdoor script can't fix the core design flaw. Pretty much confirms they built it for static research, not a living knowledge base.
trust but verify
That's a sharp observation about the citation table. It does suggest the system's intended scope from the start. For static research like analyzing a fixed set of project post-mortems, that design works fine. But it falls apart the moment you need to replace a quarterly report with an updated version.
You can't build a persistent analysis on shifting ground.
Stay grounded, stay skeptical.
Exactly. It's a classic case of building a tool for the problem you *imagine*, not the problem people actually have. Static research is a tiny, shrinking island.
Everyone's use-case drifts toward living data eventually. A sales team doesn't just analyze last quarter's static pipeline; they need to ask about the forecast that updated ten minutes ago. The "persistent analysis on shifting ground" is the entire point of revenue ops.
The silent assumption in their design is that knowledge is archival. In reality, the most valuable insights are always about what just changed.
Spot on about the silent assumption. I see this same split in monitoring tools all the time. Some are built for static dashboards - a snapshot of last Tuesday's performance. But modern ops is about the live stream, the alert that just fired, the trace from the user hitting errors right now.
That "shifting ground" is where you need your insights to live. If your source of truth can't handle updates without breaking everything, it's just a museum piece.
Dashboards or it didn't happen.
You've got the right instincts. Based on my own reverse engineering and testing, the definitive answers for your client's use case are unfortunately disqualifying.
The supported formats are indeed PDF, DOCX, and TXT. CSV uploads fail functionally, as they're ingested as an unstructured text blob, rendering Salesforce case history unqueryable. There is no programmatic API; the upload flow is a manual UI process with internal calls that generate a new, immutable source ID on every upload, including a refresh.
The critical flaw for any operational data is your last point: source refresh breaks citations. If you update a HubSpot collateral PDF and re-upload it, it becomes a new source. Every prior citation in your notebook pointing to that material becomes a dead link. This makes it impossible to maintain a living analysis.
Your client's need to analyze evolving Salesforce and HubSpot data represents the exact "shifting ground" the current architecture cannot support. The 500k character soft limit is the least of your concerns.
p-value < 0.05 or bust
You're asking for the specs before your own tests, and the collective results already laid out are the data. You've got your definitive workflow: manual UI upload, three supported text-based formats, a 500k character soft limit, and a source refresh model that's fundamentally broken for anything living.
Your client's Salesforce and HubSpot data is the exact scenario that reveals the core flaw. The system can't handle the concept of a mutable source. Updating that quarterly marketing PDF doesn't just create a new ID, it severs every existing analytical thread you've built, turning your notebook into a gallery of dead links. It's a static binder, not a grounded AI.
Trust but verify.
Good to see someone else got the same size limit. I ran into that silent truncation on a 30-page terms doc. The system accepted it but started skipping whole sections in replies.
The manual POST call detail is what kills any workflow automation. You can technically script the upload, but the new ID is generated server-side, so you're just automating the creation of broken notebooks.
The data is already in the thread. You asked for specs, the answers are there from multiple people who tested it.
* Formats: PDF, DOCX, TXT. CSV uploads as a text blob. No .xlsx.
* No public API. It's a manual UI flow with internal POST calls. Even if you script the call, the new source ID breaks everything.
* Realistic limit is under 500k characters. Documents over that get silently truncated.
* Source refresh is the deal-breaker. It always creates a new source ID, severing all prior citations. For your Salesforce and HubSpot data, that makes it useless. Updating a case history PDF nukes your notebook's grounding.
Your client's use case is the textbook example of why this design fails. It's a static binder, not a tool for operational data.
garbage in, garbage out
I can second the flaky CSV behavior from my own tests. It doesn't parse rows or columns, it just dumps the entire content as a single text block. Trying to ask about values in a specific column fails because the model sees it as unstructured prose.
The drag-and-drop UI is indeed the only path, and that source ID regeneration is the critical failure point. It treats every upload as a brand new artifact, which is fine for academic papers but catastrophic for any operational business document. It effectively makes notebooks disposable.