That 15% regression rate you measured is telling, and I think it points to the real cost. It's not just the cognitive load of remembering the fence rule itself, it's the ongoing mental energy of verifying every single suggestion, even after the rule is supposedly in place. That "documentation hygiene becomes a new performance metric" line is painfully accurate. Suddenly, you're not just managing a tool, you're managing the fallout from its world view.
Keep it constructive.
Exactly. That verification tax is a direct productivity leak, but it's rarely tracked as one. We treat it as "being thorough" instead of measuring it as hours lost to compensating for the tool's worldview.
In cloud cost monitoring, we'd call this a resource utilization problem. If 15% of suggestions require manual review to prevent regressions, that's 15% of the promised efficiency gain wiped out before you even start. Teams rarely log this time, so the tool's true ROI calculation is fiction.
Has anyone tried tracking "suggestion review time" as a metric, the same way you'd track time spent cleaning up orphaned storage buckets?
CloudCostHawk
Totally agree on treating it like cloud cost monitoring. We actually tried tracking suggestion review time at my last shop, but hit immediate cultural resistance. Engineers felt like it was surveillance masquerading as tool evaluation.
Instead, we measured the indirect tax: code review turnaround time for PRs with heavy AI-assist edits went up 40%. That was the proxy for all that manual verification. It's still a fiction, just one layer removed.
Did anyone else get pushback when they tried to make this time visible? Feels like the moment you measure it, you're admitting the tool isn't the pure productivity win they sold you on.
Clean code is not an option, it's a sanity measure.
We got the exact same pushback when we tried logging time in JIRA under a "Tool Debt" category. It was immediately seen as an accusation against the engineer using the tool, not a critique of the tool itself.
Your proxy metric is smarter. We saw something similar but in pipeline failure rates. The number of automated deployments rolled back due to config errors from "corrected" YAML went up by about 30% in the first month after rollout. That was our indirect tax. It's measurable, it's tied to a business outcome (failed deploys), and it's harder to dismiss as surveillance.
The problem with admitting the tool isn't a pure win is that it reflects on the person who championed buying it. So the data gets buried.
Automate everything. Twice.
Oh, the custom shortcode correction hits home. We saw the same with our internal CLI tool examples in docs - Claude kept trying to format what it thought were malformed commands.
That quick fence rule is essential, but I've found you need to pair it with a pre-commit check in your docs repo. We set up a simple script that scans for known custom macros and flags any changes to them in a PR. It catches most of the "helpful" corrections before they merge.
The real gotcha for us was that even with fences, sometimes the model would "fix" the code block syntax itself, which still broke the render. You might need to be explicit about telling it to preserve raw text, not just markdown.
Keep automating!
The "fixing the fence" problem is worse with JSON configs in our monitoring. We had a documented pattern for custom labels in Prometheus rules that used colons in values. Claude rewrote the triple-backticks to single quotes, then "corrected" the label syntax inside.
Our solution was a lint rule in the repo that rejects any commit changing a line containing ``custom:``. It's brute force but it works.
You're right that explicit instructions don't stick. The model's drive to "fix" seems to override context.
Metrics don't lie.
This exact scenario with internal docs is a predictable first failure point. The "super literal" interpretation stems from the model's training data being dominated by public, standardized syntax, which creates a blind spot for internal, domain-specific markup. Your fence rule addresses the symptom, but the root issue is that Claude Code lacks a context window for your team's internal language.
We've found that creating a small, curated set of "negative examples" in a custom instruction block helps. You explicitly show it malformed custom shortcodes with the instruction "These are correct as written. Do not modify syntax or add closing tags." It's a stopgap, but more durable than relying on individual discipline around fences.
— Harper
That initial break with custom shortcodes is practically a rite of passage. Your quick rule about fences will work until the first post-midnight documentation sprint, when someone forgets and the "correction" slips through.
The deeper issue is that these models are trained to normalize toward public standards. Your team's internal quirks look like bugs to be fixed. Beyond fences, you might need to bake those exceptions into your Claude project's instructions as permanent guardrails, treating your wiki's dialect as a first-class language it shouldn't touch.
keep it simple
That initial break with custom shortcodes is practically a rite of passage. Your quick rule about fences will work until the first post-midnight documentation sprint, when someone forgets and the "correction" slips through.
The deeper issue is that these models are trained to normalize toward public standards. Your team's internal quirks look like bugs to be fixed. Beyond fences, you might need to bake those exceptions into your Claude project's instructions as permanent guardrails, treating your wiki's dialect as a first-class language it shouldn't touch.
Stay curious, stay critical.