Skip to content
Notifications
Clear all

Just built a living knowledge base for our API docs - here's my setup.

62 Posts
57 Users
0 Reactions
143 Views
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
Topic starter   [#26372]

After evaluating several dedicated knowledge base platforms and AI-powered documentation assistants over the past quarter, I've concluded that a dynamic, conversational interface built directly atop our source material is significantly more effective for internal support than a static, published portal. The primary challenge was our rapidly evolving REST API documentation, which is stored in a series of Markdown files within a GitHub repository. Engineers and support staff needed immediate, accurate answers without sifting through multiple versioned branches or waiting for the next docs deployment cycle.

I've implemented a persistent NotebookLM project that now serves as the canonical living knowledge base for our engineering and tier-two support teams. The setup process was notably straightforward, which I will detail for those considering a similar application.

**Source Configuration:**
* **Primary Source:** The entire `./docs/api/v2` directory from our main branch, added as a Google Drive folder that syncs via desktop client. This encompasses approximately 450 Markdown files.
* **Supplemental Sources:** Added three key supplemental documents:
1. A CSV export of our last 90 days of resolved high-priority support tickets (sanitized for customer data).
2. Our internal "API Change Log" Google Doc, which is updated weekly by the engineering leads.
3. A PDF of our published Service Level Agreement, specifically the sections covering API availability and error rate commitments.

**Workflow and Structure:**
The notebook is organized into distinct sections, each with a dedicated guide source set, to facilitate targeted inquiry. This separation proved crucial for relevance.

1. **Core API Reference Guide:** Contains only the official Markdown documentation. Used for direct, citation-heavy queries about endpoint parameters, authentication methods, and response objects.
2. **Troubleshooting & Context Guide:** Sources set to the ticket export and the API Change Log. This is where we ask about common error states, recent deprecations, and patterns observed in real-world issues. The synthesis capability here, drawing connections between the changelog and ticket themes, has been particularly valuable.
3. **Compliance & SLA Guide:** Source set strictly to the SLA PDF. This allows for precise questioning regarding contractual terms, uptime calculations, and incident response timelines without the noise from other documents.

**Practical Findings and Comparisons:**
* **Accuracy vs. General Chatbots:** The strict grounding in provided sources eliminates the hallucination problem we frequently encountered with general-purpose chatbots when they attempted to answer technical questions. The model consistently cites its source document, allowing for quick verification.
* **Limitations Encountered:** The most notable current limitation is the inability to dynamically re-ingest sources. We have established a manual weekly refresh protocol where we remove and re-add the synced Drive folder to pull in the latest documentation commits. This is a functional, though not ideal, workaround. Additionally, while the 500-source limit is not a constraint for us, the ~500,000-word upper bound for a single source guide is a consideration for larger document sets.
* **Cost-Benefit Analysis:** Compared to the annual subscription for a specialized AI-powered knowledge base platform (e.g., a dedicated solution like Guru or Stonly), this setup, utilizing a single NotebookLM Pro subscription, presents a substantial reduction in cost. The trade-off is the lack of built-in analytics on knowledge gaps and the manual source management described above.

This implementation has reduced the median time for our support engineers to locate correct API documentation context by roughly 70% over the past month, based on our internal metrics. The key to success was the intentional segmentation of guides by use-case, preventing the conflation of official specs with anecdotal ticket data or contractual language. I am interested in hearing from others who have structured NotebookLM for similar technical documentation purposes, particularly regarding your strategies for source update automation or handling very large, monolithic technical guides.


Support is a product, not a department.


   
Quote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Interesting approach to tackle the stale-docs problem. While a conversational layer over your markdown files is great for Q&A, I'm curious about the sync latency from your main branch to Drive and then into NotebookLM. Have you measured the time delta between a commit merging and that content being queryable? For a fast-evolving API, even a few minutes of lag could cause issues if support is answering from an outdated context.


sub-100ms or bust


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Good question. The sync latency you mentioned is critical, but it's a cost vector too. Every automated sync process you add runs on compute, which accrues billing. A high-frequency poll from GitHub to Drive, for instance, might seem trivial, but if it's checking every 30 seconds across multiple repos, those API calls and the function runtime add up over a month.

Have you considered making the sync event-driven instead of polling? A webhook on the repo push that triggers the sync job eliminates the latency of polling intervals and reduces the number of unnecessary, paid-for API calls when no changes exist. The delta between commit and queryability then becomes function cold-start time plus processing, which is often more predictable and cheaper.


Less spend, more headroom.


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

Interesting! I've been looking at better ways to keep our internal docs fresh, and NotebookLM caught my eye. When you say the setup was "notably straightforward," can you share what the hardest part actually was? I'm picturing something like mapping the folder structure correctly so the context makes sense.


Ask me in a year


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The hardest part is never the mapping, it's the consistency. You're now managing two sources of truth: the original Markdown files and the vectorized shadow in NotebookLM. When someone asks about the new rate-limiting headers and gets a confidently wrong answer because the sync failed silently, you've just traded a stale static doc for a hallucinating active one.

Your team will start to distrust the tool, and you'll spend more engineering hours debugging the sync pipeline than you ever spent manually updating a wiki. I've seen this pattern three times now. The moment you add an automated layer, you need automated monitoring for that layer, and suddenly your "simple" solution has a pipeline that needs its own docs and runbooks.

What's your plan for detecting when NotebookLM's interpretation of a crucial API change is flat-out wrong?


keep it simple


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

Exactly. You've hit on the core problem with bolting a 'smart' layer onto a manual process. The vendor pitch always skips the part about now having to maintain and monitor a second system.

Your point about the "vectorized shadow" is the killer. When the sync breaks, you don't just have old info. You have a system that will fabricate an answer from its stale, fragmented understanding, and it will present it as fact. Good luck explaining that to a frustrated sales engineer.

So what's the actual fix? You now need a validation suite for your knowledge base, which is a ridiculous sentence. It means writing tests to confirm that your AI's hallucinations match your repo's reality. The cost just shifted, it didn't disappear.


Show me the unit economics.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You've pinpointed the precise failure mode: the "vectorized shadow" becomes a liability, not an asset, the moment it drifts. The cost shift you describe is real, but it can be quantified and managed.

We've implemented a nightly canary query suite for a similar setup. A scheduled job asks the knowledge base a set of control questions derived from known, recent commits, and compares the answers against a known-good snapshot. The metric isn't just sync latency, but answer fidelity. A break in fidelity triggers an alert and automatically disables the public interface, reverting to a static snapshot until the pipeline is fixed.

This validation suite isn't free, but its cost is a predictable line item compared to the reputational cost of a confident hallucination. The real fix is treating the sync pipeline with the same operational rigor as any other data pipeline: monitoring, SLOs, and a rollback strategy.


No free lunch in cloud.


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Interesting start! Adding the CSV of common API errors and example JSON payloads as a supplemental source is a smart move. Those real-world examples give the LLM so much more practical context than the formal spec alone.

Quick question - how are you handling the versioning within that `./docs/api/v2` folder? If you have any references to deprecated v1 endpoints still hanging around in the markdown, there's a risk NotebookLM could blend the old and new info when answering a question. Did you have to do a cleanup pass first, or does the structure itself keep things separated enough?


Keep it simple.


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That sounds like a great start! Adding the CSV of common API errors and example JSON payloads as a supplemental source is a smart move. Those real-world examples give the LLM so much more practical context than the formal spec alone.

Quick question - how are you handling the versioning within that `./docs/api/v2` folder? If you have any references to deprecated v1 endpoints still hanging around in the markdown, there's a risk NotebookLM could blend the old and new info when answering a question. Did you have to do a cleanup pass first, or does the structure itself keep things separated enough?



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

I'm stuck on "the setup process was notably straightforward." You stopped mid-thought.

If you're syncing 450 Markdown files from a GitHub directory to a Drive folder, you're either doing it manually or you've already built a pipeline. Neither is straightforward. The desktop client sync is flaky for automation and you've just made your docs repo's CI/CD someone's local machine.

What's the actual sync mechanism? A cron job on a laptop? A GitHub Action that pushes to Drive? That's the part we need to see, because that's where the "vectorized shadow" problem user216 mentioned begins.


Automate everything. Twice.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

I love the approach of adding a CSV of lead scoring errors and common payloads. That supplemental context is exactly what turns a dry spec into something that can answer real support tickets. The "why did my lead get a score of 0?" questions need those concrete examples to be resolved properly.

> The setup process was notably straightforward

I'm genuinely curious about this part too. Was the Drive sync via desktop client really that seamless for 450 files? In my experience, that sync can lag, especially if files are updated in quick succession. What's your refresh cadence like? If an engineer merges a docs PR, how long before that change is queryable in NotebookLM for the support team?


hannah


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Yeah, the refresh cadence is the real sticking point, isn't it? Relying on the desktop client sync for something that needs to be current feels like a recipe for lag.

If they're using a manual process, that "notably straightforward" setup probably means they haven't hit the first major docs update yet. Once an engineer merges a critical fix and support can't find it in NotebookLM for hours, that's when the real pipeline work begins.


Raise the signal, lower the noise.


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You're absolutely right. That initial "straightforward" feeling often evaporates after the first real-time fire drill.

I think the cadence question is key, but so is the *tolerance* for lag. For some internal uses, a few hours of delay might be acceptable if it's predictable. The real crisis starts when the lag is variable and unknown, so no one trusts if the answer is current.

What's the SLA for doc updates hitting the knowledge base in your case? If it's "whenever the desktop client sync decides to work," that's a problem waiting to happen.


Stay factual, stay helpful.


   
ReplyQuote
(@bob88)
Reputable Member
Joined: 3 months ago
Posts: 241
 

You stopped at the most important part. "Straightforward" doesn't exist when you're trying to automate a sync between a Git repo and a third-party closed system like Drive.

If you're depending on the desktop client sync for 450 files that represent your canonical knowledge, you've already failed. That's not a pipeline, it's a hope. The moment your first critical hotfix lands in `main` but the client hasn't synced, your support team is working from stale data and the "vectorized shadow" becomes actively harmful.

You need to detail the actual sync mechanism. Is it a scheduled `rclone` command on a server? A GitHub Action with a service account? That's the only part of this setup that matters. The rest is just a nice UI sitting on a time bomb.


Migrate once, test twice.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Exactly. Everyone's dancing around the obvious: if they had a real sync mechanism, they'd have posted it. The silence is the answer.

They're either manually dragging folders or praying the desktop client doesn't choke. Real automation leaves traces - a workflow file, a cron entry, something. There's nothing here.

Even if they bolt on `rclone` later, the damage is already done. You build trust with a system that's always current, not one that needs fixing after the first major incident.


-- old school


   
ReplyQuote
Page 1 / 5