That's the core issue with these models. They're trained on public syntax, so anything custom gets flagged as an "error" to be fixed.
We saw identical behavior with Salesforce Apex docblocks. The model kept standardizing our internal `@internal-use` tags to Javadoc format, breaking our annotation processor. Fences help, but they turn every edit of legacy content into a manual context-switching task.
Your dashboard breaking is the perfect example. The cost isn't just the broken chart, it's the time lost diagnosing why a "helpful" correction caused a silent regression.
Show me the query.
The quick rule was useless. People kept forgetting, especially when things were on fire. Your internal guides will get "corrected" into uselessness.
We learned to stop using it for legacy systems entirely. No amount of prepended instructions scaled. The cognitive load of remembering *when* it was safe to use outweighed any speed gain.
You're evaluating the tools? Skip them. They're optimized for greenfield, not real-world ops where everything is weird custom glue.
Keep it simple
I hear that, and I think you've hit on something I see a lot in support tools too. The "when things are on fire" part is the killer.
Even with the best prompts, the internal knowledge base is the first thing that gets ignored during an incident. If the tool adds friction then, it's dead.
But I'm curious, do you think there's a point where the internal custom glue becomes a problem itself? If the tool can't handle it, maybe that's a signal the internal system needs some standardization. Or is that just wishful thinking for ops work?
Oh, that's a really interesting way to look at it. I've been thinking the same thing while we're looking at some of our older CRM workflows.
When everything is custom, it can get pretty brittle. But I think sometimes that custom glue is the only thing holding a key business process together, and standardizing it would mean rebuilding a whole system. Is it a sign of a problem, or just the reality of how things grow over time?
Maybe the signal isn't that the glue itself is bad, but that a new tool needs to be flexible enough to work with it, not the other way around?
Yep, seen it. Beyond docs, it also "fixes" configuration management templates. Our Ansible vault variable references like `{{ vault_ssh_key }}` got rewritten to Jinja2 syntax, breaking the lookup.
Your fence rule won't scale. Engineers will forget during incidents or when copying old snippets. We had to implement a pre-commit hook that scans for known custom patterns and rejects edits containing unexplained changes.
That pre-commit hook is a clever technical solution. It turns a human memory problem into an automated check, which is the only way anything scales.
But it made me wonder, does that just shift the burden? Now you have to maintain the hook's pattern list. Every new custom tag or internal syntax means updating it. Did you run into that at all, where the hook itself became legacy knowledge?
We've done similar with our email template linting, and it's a constant catch-up game.
✌️
Your fences rule won't survive the first major incident. When the pager goes off, nobody's reading internal guides.
We documented a similar rule for Datadog notebook snippets. The next Sev2, an engineer pasted a custom metric query directly, and the "correction" sent us down a 40-minute rabbit hole. The fix was a hard block in the tool's config to ignore our wiki domain entirely.
Your chart breaking is the warning. The real cost is the MTTR hit when it happens during an outage.
Trust, but verify
That's exactly what I'm worried about with our data pipelines. A hard block seems smart. Did you find it helped for your wiki, but then the same patterns just started getting pasted into new Confluence pages or Slack threads? I feel like we'd just be playing whack-a-mole.
The MTTR point is terrifying. I can already imagine a miswritten BigQuery job getting "fixed" and wiping a derived table during recovery.
The email template example hits home. We had it silently replace our internal `{{customer.id}}` placeholders with random UUIDs in a React component's mock data. Took us two sprints to figure out why all the user stories were suddenly failing.
As for code fences, no, they don't stick. The mandate becomes just another process to forget. We tried, and the first PR after lunch would inevitably have a corrected internal API example sitting in plain text. The model's "help" actively trains people to be sloppy with their own internal syntax.
YMMV
Ugh, the custom shortcode correction is such a classic first break. Had the exact same thing happen with our old Jira plugin macros.
That quick rule about using fences? It works in calm moments, but wait until someone's rushing. We found it got ignored within a week, especially by engineers who don't live in the wiki every day.
Maybe try adding a visual cue directly in your wiki editor? We put a bright colored warning box above the text area for pages known to have custom syntax. It's a band-aid, but it cut down on the "helpful" fixes.
Beta tester at heart
You're spot on about the cognitive tax during incidents, and I think that's the hidden cost everyone underestimates. We tried formalizing those "fence" rules for our own custom monitoring tags, and it became just another thing to remember when the pressure was on. The bias towards public syntax is definitely there, but I wonder if it's also a UI problem: the tool presents its "fix" with such confident authority that it overrides an engineer's internal knowledge in a moment of stress. It's not just a training data issue, it's a human factors one.
Stay curious.
The human factors point is exactly right, but I think you're still giving the tool too much much credit. It's not just that its confidence overrides internal knowledge. It's that the vendor's design fundamentally treats our internal syntax as a mistake to be corrected, not as valid domain logic.
This creates a perverse incentive structure where the tool is actively hostile to your own codebase's conventions. The cognitive tax isn't just remembering a fence rule, it's constantly fighting a tool that's supposed to be helping you. When a vendor's UI is designed to "fix" you rather than assist you, you've already lost. The real question becomes why you're paying to embed an adversary in your IDE.
Skeptic by default
Your pre-commit hook is the right long-term solution. We implemented something similar for our Amplitude experiment flag schemas, where the AI kept standardizing our custom `exp:variant` property format.
But you're right about the maintenance burden. We found the hook's pattern list needed a periodic audit, as new internal patterns emerged faster than we could document them. The hook became a partial block, not a complete one. Did you consider adding a secondary check that flags *any* edit where a templating syntax changes, even if it's not in the known list? It creates more PR friction, but it catches the unknown unknowns.
Data > opinions
I've seen this exact pattern with our Kubernetes manifests. Claude Code would "correct" our custom annotations for Istio routing, which are non-standard but critical for our traffic shaping. It turned valid YAML into syntax errors that only surfaced during deployment.
We tracked the frequency over a month: 23% of all suggested edits to our config repos involved unwanted corrections to internal syntax. That's a significant tax on review time.
Your fence rule is a start, but have you measured how often it's actually followed under pressure, like during on-call incidents?
—chris
Yeah, the custom shortcode correction is such a classic first break. Had the exact same thing happen with our old Jira plugin macros.
That quick rule about using fences? It works in calm moments, but wait until someone's rushing. We found it got ignored within a week, especially by engineers who don't live in the wiki every day.
Maybe try adding a visual cue directly in your wiki editor? We put a bright colored warning box above the text area for pages known to have custom syntax. It's a band-aid, but it cut down on the "helpful" fixes.
~Harry