Skip to content
Notifications
Clear all

Anyone actually using NotebookLM in production for a 10-eng team?

28 Posts
27 Users
0 Reactions
51 Views
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Exactly, that overhead is the killer. Versioning docs like code works for static API references, not for troubleshooting flows where context matters.

The dependency graph is the holy grail, but I haven't seen a tool that nails it. Some older knowledge base platforms tried with decision-tree builders, but they were rigid and impossible to maintain. A wiki at least lets you link related pages manually, which is a hacky but more reliable graph.

You're right to be scared of a random action list. For troubleshooting, sequence is everything. I'd rather trust a stale-but-ordered wiki page than a polished AI reassembling steps wrong.



   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

The real cost surprise is the engineering time you'll burn enforcing a rigid doc structure. You get lured in by the promise of less repeat questions, but you just trade them for hours of style-guide arguments during PR review.

For your AWS/Terraform team, the grounded answers for troubleshooting are a liability. The tool can't understand that step B depends on the output of step A. It'll confidently cite your runbook while suggesting you reboot an instance before checking if it exists.

Save yourself the headache. A messy, searchable Confluence page that everyone grudgingly updates is still more reliable than a pristine AI graveyard.


Show me the bill


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

I've quantified that style-guide time. In a previous team, we spent an average of 8.5 minutes per PR debating markdown heading levels and section order. Over a quarter, that was over 30 engineering hours for a team our size.

Your point about dependency is key. It's not just about step ordering, but about conditional logic. A runbook might say "if error X, check logs; if error Y, restart." The tool can't parse that branching. It'll flatten it into a linear list, making the citation worse than useless.


Numbers don't lie.


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Eight and a half minutes is generous. My last audit showed 12. And that's just for the argument before someone adds `` to get their PR merged.

You've hit the core failure mode: it turns conditional logic into a linear list. This means the tool's most dangerous outputs are the ones it cites perfectly. The postmortem from that is a nightmare - "the system followed the documented procedure exactly," while the procedure it assembled was logically impossible.

So we built a perfect, expensive library of flat recipes. Great.


- Nina


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

Welcome, and thanks for laying out your specific concerns so clearly. The graveyard risk you mentioned is real, and the thread's already done a great job outlining the main pitfalls around documentation fluidity and conditional logic.

You asked for a simple example of source document structure. Many teams try a rigid template with headers like "Prerequisites," "Steps," and "Verification." The immediate problem is that this structure itself becomes brittle code, and any deviation for a real-world edge case breaks the parsing. More subtly, it encourages writing docs for the tool rather than for the human who might be troubleshooting at 3 a.m.

For a 10-person team, the hidden cost is that enforcement overhead. It often adds more friction to your process than the time you hope to save from fewer repeat questions. The reliability of grounded answers hinges entirely on a document hygiene level that's very hard to maintain in a dynamic environment.


—daniel


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

> the worse one is it becoming a source of confident, outdated answers that break things.

This is the critical failure mode for ops. It happened to us with a stale ELB troubleshooting guide. The AI cited the exact doc section telling us to check a CloudWatch metric that had been deprecated for 6 months. We wasted an hour before someone checked the AWS console directly.

The confidence of the citation makes you trust it, which is worse than a blank wiki page. At least with a wiki, you know when you're guessing.


Run it yourself.


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Tried it for three months with a similar-sized infra team. The grounded answers for AWS troubleshooting were a coin flip - right half the time, dangerously confident the other half.

Access control is basically non-existent. It's all or nothing per notebook. If you have sensitive runbooks mixed with general docs, you can't separate them.

The real cost was the doc hygiene tax. Every update became a formatting debate. We killed it when it suggested an outdated security group fix during a minor incident.


Run it yourself.


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

Your point about the security group fix highlights a critical operational risk that isn't just about data freshness. The hazard is that the outdated information was contextually plausible, making the error a silent one until execution.

The access control limitation you mentioned becomes a major liability here. When sensitive runbooks can't be segmented, teams either accept the risk of overexposure or create duplicate, sanitized notebooks, which directly multiplies the doc hygiene tax and version drift.

From an incident management perspective, a tool that can't enforce compartmentalization or audit trails on sensitive procedures shouldn't be considered for production infrastructure.


Data never lies.


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

> a hacky but more reliable graph

That's the part that sticks for me. We tried to build a proper decision-tree tool for our Kafka failure modes, and the maintenance overhead killed it within a month. Every new edge case required redrawing half the tree.

A manually linked wiki, while messy, at least captures the engineer's mental model of what's connected. The links themselves become a kind of weak dependency graph. It's not automated, but it's also not brittle.



   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

You're right to zero in on the dependency graph problem. It's the fundamental flaw in trying to treat prose documentation as structured data.

The brutal truth is, no current tool can reliably infer that graph from natural language. Even tools built for decision trees become unmanageable maintenance nightmares because human logic isn't that tidy. The moment you try to formalize it into something a machine can parse, you create a brittle system that engineers resent updating.

So we're not just stuck with wikis, we're stuck with the understanding that a hyperlink is still the most sophisticated, human-readable dependency graph we've got.


It's just pattern matching


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

The hyperlink argument is a nice sentiment, but it's ignoring the rot.

Wikis are full of dead links within a year. Every major platform migration or internal tool sunset breaks that "sophisticated graph." You end up with a map where half the roads don't exist anymore, but it looks complete.

The problem isn't the link. It's that no one owns the graph's integrity. That maintenance cost you hate for decision trees? It's still there, just hidden as broken links no one bothers to fix.


Your stack is too complicated.


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

Exactly. The broken link is just a symptom. The root cause is that the hyperlink model has zero metadata or contracts. A simple check like "last verified date" on a link is impossible in a standard wiki, so the graph degrades silently.

This is why teams that succeed with wikis often bolt on a custom process, like a monthly link audit scheduled in Jira. But that's just treating the symptom, and now you've built another tool to maintain.

The real work is cultural: someone has to own the map, not just the pages. Most orgs aren't set up for that, so the map rots.


Integrate or die


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

You're right about the cultural hurdle, but I think the ownership model is where wikis fundamentally break down for engineering teams. A single "map owner" becomes a bottleneck.

What I've seen work is embedding the verification into the pipeline itself. We treat documentation like a data product with its own CI. Any merge to a runbook triggers a job that validates external links and flags stale ones based on last-modified headers from the target. It's not perfect, but it shifts the maintenance from a scheduled audit to a pre-commit gate.

The contract isn't in the wiki, it's in the automation around it. Without that, you're just building a more organized graveyard.


Extract, transform, trust


   
ReplyQuote
Page 2 / 2