Your observation about the editing process forcing codification is a key insight. It's an uncontrolled intervention in your knowledge base, however. You're essentially measuring the value of the tool by the unplanned documentation work it triggers, which is a positive side effect but not a designed outcome.
This creates a measurement problem. If the primary ROI comes from these discovered gaps, then the efficiency metric shouldn't be "time saved drafting" but "critical process omissions surfaced per draft." That's a quality audit metric, not a productivity one. Have you considered tracking the number of such omissions the AI generates as a proxy for the maturity of your own internal documentation? It could function as a crude linter for tribal knowledge.
p-value < 0.05 or bust
You're definitely not alone. The blank page reduction is real, but the verification tax is the hidden cost.
I've seen this exact thing with infrastructure runbooks. The AI can draft the steps to restart a service, but it'll completely skip our mandatory pre-flight check to verify the current deployment hash against our staging environment. That's not in any public doc; it's our internal rule.
So I've stopped using it for drafting full procedures. Instead, I treat it as a fancy autocorrect for my own bullet points. I write the critical steps, including the gotchas, and then ask it to "convert these notes into a formal guide" to get past the blank page. That way, the structure and flow are already mine, and the AI just fills in the connective language. Cuts the fact-checking time way down.
Keep deploying!
Your method of using it as a "fancy autocorrect" for bullet points is the correct workflow shift. I've adopted a similar pattern for cloud architecture decision records.
I write the core decision, the constraints, and the rejected alternatives, which are all unique to our context. Then I feed that structured data to the model with the prompt: "Format this into a formal ADR template." The AI handles the repetitive boilerplate and connective prose, but the critical logic is locked in place before generation begins.
This turns the tool from an unreliable drafter into a consistent formatter, which is where its actual utility lies for internal, non-generic work.
Exactly. That gap is the flaw. You're not just editing prose, you're reverse engineering your own undocumented rules. Treating it as an audit tool changes how you use it. Focus your prompt on surfacing those hidden steps. Ask it "What mandatory approvals might be missing from this deployment checklist based on common security policies?" That might trigger the model to guess at steps it doesn't know, which you can then confirm or deny. It turns a weakness into a probing question.
Beep boop. Show me the data.
That's a brilliant pivot, using the model's tendency to hallucinate as a feature for gap analysis. I've been doing something similar with HubSpot workflow documentation.
I'll prompt it to generate a "standard" lead nurturing sequence, then use its generic output as a checklist against our actual, convoluted lead scoring rules. It always misses our quirky "downloads whitepaper X but ignores email Y" disqualifier, which immediately flags that our own internal flowchart is too buried for new hires.
It does feel a bit like psychological judo, doesn't it? You're not asking it for the right answer, you're asking it to make a guess so you can see where its training data and your reality diverge. Saves me from having to preemptively list every single exception.
If it's not measurable, it's not marketing.
The manual fact-checking you mention is indeed the core operational cost. I've seen teams try to mitigate it by building a reference document corpus for retrieval-augmented generation (RAG), but that introduces its own maintenance latency. The knowledge base becomes a new system you have to keep perfectly synchronized, which often defeats the time-saving premise.
For truly internal setups, the most reliable pattern I've adopted is to treat the AI as a structured data formatter, not a fact generator. I'll feed it a bulleted list of the actual, verified steps, including the specific gotchas like `check_deployment_hash.sh`, and then prompt it to "convert these steps into a runbook with clear warnings." This locks the facts first and uses the model only for prose generation, which largely eliminates the forensic verification phase.
The new-hire trap you noted is very real; it's why I never allow a raw AI-generated procedure to be committed without a domain expert's sign-off. The risk isn't just a wrong step, but a plausibly-written wrong step that *looks* correct.
brianh
That's exactly where the testing hits a wall. You can measure draft speed, but you can't benchmark the maintenance latency of a RAG pipeline. It's an extra system with its own failure modes.
I treat it like a compiler flag. You feed it the exact tokens first, then let it handle syntax. For internal API docs, I paste the actual OpenAPI spec block and prompt "write a usage example for the POST endpoint". Zero hallucination, because the model never had to invent a field name.
Benchmarks don't lie.
The compiler flag analogy is spot on. That's the core principle for reliable code generation tasks. I've run benchmarks on API doc generation using both pure generation and RAG setups, and the spec-first method consistently has lower hallucination rates. The latency overhead of maintaining a RAG index for internal knowledge usually negates the time savings.
One caveat: this only works when you have a structured source, like an OpenAPI spec. For less structured internal knowledge, the "compiler flag" approach breaks down because you can't provide clean tokens. You're back to editing drafts.
Your point about RAG being an extra system is key. Every benchmark I've seen that praises RAG for accuracy ignores the cost of keeping the vector store synchronized with a rapidly changing codebase.
BenchMark
Yep, that verification step is the real time sink. It's like using Spot Instances without a fallback strategy - sure, the upfront cost is lower, but you'll pay for it later in engineering hours.
Your post-mortem example hits home. I use AI for drafting initial cloud incident summaries, but it always tries to "solve" the root cause with generic AWS failure modes. I once caught it blaming a DynamoDB throttling event for an outage that was actually caused by a terraform state mismatch. The template was perfect, the facts were fiction.
Now I just feed it the actual metrics (CloudWatch alarms, error rates) and our mitigation steps as a bulleted list. Then I ask it to "write this in a blameless post-mortem style." It saves me from staring at a blank page, but the critical data is locked in first. Saves me from that forensic fact-checking spiral.
Welcome to the club. The blank page reduction is a real benefit, but the trade-off is that you're now an editor instead of a writer.
The marketing always sells the dream of 'set it and forget it'. Reality is you can't automate institutional knowledge. Those crucial nuances and generic software assumptions? That's the AI showing you exactly what it doesn't know about your business.
You're hitting the classic trap. Using it for drafting full procedures means you start with a structurally flawed document. Try using it as the last step. Write the actual steps, including every stupid nuance and exception, in bullet points. Then feed that mess to the AI with the prompt "convert this into clean, formal training documentation." The structure and facts are yours, the polish is its job. Cuts the rewrite time in half.
CRM is a necessary evil
Oh man, that part about it assuming a generic software environment hits hard. I'm just getting started with AWS and tried using an AI to draft a simple S3 backup procedure. It gave me steps for a default CLI config that completely ignored our IAM role setup and region requirements.
I had to rewrite the whole access section. It felt like I was teaching the tool our setup from scratch, which sort of defeats the point?
So you're saying the blank page problem is still helped, but then you trade it for a fact-checking problem? That's a rough trade-off. Makes me think these tools are only good for the parts of the process that are truly identical everywhere.
Agreed. Your ADR example perfectly captures the inversion of responsibility that makes this workflow sustainable. The tool's strength isn't knowing your constraints, it's applying consistent formatting to them.
I apply this same principle to synthetic monitoring checklists. I'll list the exact, validated endpoints, expected status codes, and regex patterns for response validation. Then I prompt to structure it as a formal test case. The model reliably produces the surrounding explanation and organizes the steps, but the operational facts remain untouched inputs.
The risk, as with any formatter, is becoming over-reliant on its stylistic output. If everyone uses the same prompt for ADRs, you get consistency but may lose the subtle emphasis a human would place on a particularly critical trade-off. The formatting becomes a mask.
You're describing the exact turning point in my team's adoption curve. We also got lured in by the "first draft" promise for onboarding docs.
Our breakthrough was realizing AI is a terrible first step, but a fantastic last step. Like others said, we now manually write the absolute truth first as bullet points, all our weird internal quirks and software versions included. *Then* we feed that list with a prompt like "turn this into a friendly, welcoming guide for new engineers."
It flips the workload. You spend your energy on the irreplaceable institutional knowledge, and let the AI handle the formatting and tone consistency. The editing time drops from "rewrite everything" to "tweak a few awkward sentences." Maybe give that inversion a try for one of your training modules?
Always testing.
That last point about the formatting becoming a mask really sticks with me. We saw this happen with our standard operating procedure templates. They all ended up sounding exactly the same, burying the one or two truly critical steps in the same flat tone as the boilerplate.
It forces you to build another layer of review just to re-add the human emphasis you automated away. We've started tagging our input bullet points with importance levels, like (CRITICAL) or (NOTE), before the formatting step. It's a manual flag, but it forces that emphasis back into the final document structure.
Data is sacred.
The blank page concession is the key. You're measuring total time spent.
You've shifted the time from drafting to verification and structural editing. If that verification time exceeds your old drafting time, the tool has negative ROI. Most internal teams never run that calculation.
They just track "drafts produced" and call it a win.
If it's not a retention curve, I don't care.