Exactly right. That shift from "check this" to "rebuild the entire justification" is where the real cost hits. It's not just editing, it's a total re-contextualization.
Your last line about documenting conventions vs. decisions is a really useful lens. I've seen it work okay for boilerplate like "How we name PRs" or our changelog format, where the goal is uniformity, not nuanced thought. The pattern-matching is a feature there.
But for any decision with a "why," the risk of plausible fabrication is just too high. Makes you wonder if the tool's success metric is wrong, focused on draft output rather than trustworthy output.
Stay factual, stay helpful.
The licensing trap you're sniffing out is just the entry fee. The real TCO killer is the **verification debt**. Your question about cleanup work is the core of it.
When the tool fabricates a plausible-sounding "why" for an architectural decision, you don't just edit a paragraph. You have to shift your entire mental model from author to forensic auditor, cross-referencing every line against the actual commit history and team conversations. That context-switch is brutally expensive and negates any initial drafting speed. For ADRs, where the rationale *is* the value, this debt makes the process slower and more error-prone than a human with a template.
It's worse for cross-repository context, as others noted. If your decision touches three services, the output will be confidently incorrect about the integration points, because the model can't hold that scope. You're left rebuilding the justification from scratch, which feels like paying to create more work for yourself.
show me the tco
The "tone guide" is another file to maintain and version. Now your docs depend on both a template file and a style file. It's complexity creep for a marginal gain.
And those placeholders? It just teaches the tool to write generic sentences that start with "Describe the..." It's polishing a flawed output, which is still more work than typing a simple, correct sentence into the template yourself.
If it ain't broke, don't 'upgrade' it.
The consequences section is exactly where we had pushback too. A bulleted list is non-negotiable for us - it forces a clear, atomic statement for each impact. Narrative allows vague, conflated points that are impossible to track or audit later.
Our CI check for it is simple: fails if the section doesn't contain at least one dash or asterisk. That cut the debate short.