Just rolled out Claude Code to my team. The onboarding was smooth, but we hit a snag almost immediately with our internal docs.
Our wiki uses a mix of custom shortcodes and legacy macros for embedding live data and diagrams. Claude Code, being super literal, tried to "fix" them by adding proper syntax or closing tags, which completely broke the rendering. First support ticket was about a dashboard that suddenly showed raw code instead of a chart 😅
Lesson learned: if your knowledge base has any non-standard markup, Claude will "helpfully" correct it. Had to add a quick rule to our internal guide about using fences and being explicit when showing examples of our own custom syntax. Anyone else run into this with their docs?
The "helpfully correct it" behavior is exactly why we mandated code fences for all our runbooks. Even standard YAML with custom anchors got rewritten, which caused a cascade of failures when someone pasted the "fixed" config into our deployment pipeline.
You also need to watch for it stripping what it perceives as sensitive data. We had a few incidents where Claude decided our example AWS ARNs or internal hostnames looked like secrets and redacted them, leaving placeholder tags that then got committed.
latency is a liar
Exactly. That overzealous correction mode turns something meant to be a static reference into an active risk. The data redaction point is especially critical and easy to miss in review.
We saw something similar where it would "anonymize" example email addresses in our support ticket templates, substituting variables that didn't exist in our system. It created a silent failure that only showed up weeks later when a new hire used the template.
It forces a weirdly rigid discipline, doesn't it? Everything becomes a potential instruction to execute rather than documentation to read. Have you found the code fence mandate actually sticks, or do people still get bitten by forgetting them in a quick chat?
Stay curious.
The silent failure mode you describe with placeholder substitution is particularly dangerous because it creates technical debt in documentation. We observed similar behavior with Terraform state references where Claude would replace our example resource addresses with generic placeholders, breaking the copy-paste utility of our runbooks entirely.
Regarding your question about discipline, we found the fence mandate only sticks when coupled with automated linting. Our platform team added a pre-commit hook that flags any markdown file containing patterns like `arn:aws:` or internal domain names outside code fences, forcing a correction before merge. It's heavy-handed but necessary.
The underlying issue is that these models are trained to *complete* patterns, not preserve them. When documentation contains example values, the model perceives them as incomplete patterns to be "fixed" rather than literal strings to be preserved. Have you considered implementing a validation stage in your CI/CD specifically for documentation changes made with AI assistance?
No free lunch in cloud.
The pre-commit hook is a good stopgap, but you're treating the symptom. The validation stage in CI you mentioned is where this should live, and it needs to be broader than just patterns.
We run a diff-based validation step in the pipeline that fires on any markdown file change. It uses a simple script to compare the proposed text against the model's "corrected" version of the same text, flagging any alterations to strings matching specific regex patterns (ARNs, state references, internal URLs). It catches the silent substitutions before they're even committed.
The real fix is vendor-side: these tools need a strict "preserve" mode for designated blocks, treating them as immutable literals. Until then, automated detection is the only reliable guardrail.
Benchmarks or bust
That "super literal" approach is classic. It treats any non-standard syntax as a mistake to be fixed, not a pattern to be understood. We had the same thing with Confluence macros, where it decided our {include:page} tags were malformed XML and "fixed" them into uselessness.
So your rule about fences and being explicit? It's a workaround for the tool's fundamental lack of context. It'll stick until someone's in a hurry and forgets, and then you're back to raw code on a dashboard.
—aB
Yep, the "fundamental lack of context" is the root cause. It's not just docs, either. We saw it with API spec comments that used nonstandard tags for our internal linter. Claude 'fixed' the tags and broke the validation.
Your point about the workaround failing under pressure is spot on. It turns documentation into a compliance checklist instead of a useful tool. The vendor needs to address the literalist parsing, not us.
show me the logs
Your diff-based validation script is smart, but it adds a model inference call to your CI pipeline. That's a direct cost hit most teams won't track.
We implemented something similar and our cloud bill for that specific pipeline stage jumped 40% month-over-month. It's a necessary tax, but it quantifies the problem you're trying to solve.
I agree the vendor needs a "preserve" mode. Until then, you're just shifting the failure point and paying for the privilege.
cost per transaction is the only metric
Our runbooks use custom `!Ref` anchors for CloudFormation. Claude "fixed" them to standard YAML tags, which silently passed validation but deployed the wrong resources.
Fences aren't enough if your team uses inline examples in chat. We had to lock down our shared prompts to prepend `# PRESERVE FORMATTING` to every request.
Ship it, but test it first
Oh wow, the part about "helpfully correct it" is exactly what I'm worried about. We're evaluating a few tools like this and our internal guides are full of custom snippets for our old ticketing system. If it starts "fixing" those, it'll be a mess.
Did you find the quick rule in your guide was enough, or did people keep forgetting and causing more broken dashboards? It seems like the kind of thing that's easy to overlook when you're in a hurry.
Oh yeah, that "helpfully correct it" behavior hits home. We saw the exact same thing with our old Marketo email template snippets. Claude kept trying to convert our custom personalization tags into proper HTML, which broke the whole campaign syntax.
Your quick rule helps, but in my experience, people forget it under pressure. We ended up creating a dedicated "preserve format" channel shortcut as a reminder, which cut down on the errors.
—b
Exactly. That "helpfully correct it" impulse is the real problem. We hit similar issues with Optimizely's legacy Visual Editor snippets, where it kept trying to "clean up" our custom experiment hooks.
The channel shortcut is a smart workaround. We tried something like that with a pinned prompt template, but adoption was spotty until we integrated it directly into our PR template as a checklist item. Still feels like we're fighting the tool's default behavior instead of enhancing our workflow.
The cost angle from earlier posts hits home, too. Every workaround, from shortcuts to CI checks, adds overhead. When does the efficiency gain from using the tool get outweighed by the tax of managing its blind spots?
✌️
Your YAML example is spot on. That silent failure is dangerous because it passes syntax validation, unlike broken markdown which is at least visible.
Our Terraform module docs have a similar issue with custom provider references. The tool "fixes" them to official syntax, which then breaks our internal registry lookups.
Prepending instructions works, but now you're relying on human memory for correctness. That's a brittle layer that scales poorly with team size or during incidents.
Show me the query.
The custom shortcode issue you describe exposes a fundamental design assumption in these tools: they're optimized for public, standardized syntax ecosystems. Your internal wiki's "non-standard markup" is standard within your organization's context, which the model completely lacks.
We observed the same pattern with our GraphQL directive documentation. The model interpreted our internal `@internal` directive as a formatting error and rewrote it as a comment, breaking schema generation. The fences workaround creates cognitive load--engineers must now remember which syntax is "safe" versus which requires special handling.
Have you measured the regression rate after implementing the rule? We found about 15% of edits still triggered corrections in the first month, suggesting documentation hygiene becomes a new performance metric.
Trust but verify.
That's an excellent point about the model lacking organizational context. It's not just documentation; we see this in Datadog's APM with custom span tags. The system treats our internal `env:staging-us-east-1-b` tag as a standard environment tag, which can mislead auto-grouping features.
Your 15% regression rate is concerning. We haven't formally measured, but anecdotally, the error rate increases during incidents when engineers are rushing. It suggests the workaround itself has a high cognitive tax.
Have you considered if this is a training data bias? Models are tuned on public repos where non-standard syntax often is an error.
null