Everyone's raving about AI auto-reply for customer tickets. I've run the numbers. For external customer-facing support, the deflection rate is garbage and the satisfaction drop is measurable. But we're using it internally for developer platform support, and it's a different story.
The difference is context and data boundaries.
* **Internal:** The AI has full, sanctioned access to internal docs, runbooks, deployment guides, and the codebase. The questions are technical and specific.
* **External:** You're feeding it a sanitized KB and hoping it doesn't hallucinate a product feature that doesn't exist.
Here's a real example from our internal logs. A dev asked: "Why is the staging deployment for service X failing with error 'schema validation failed'?"
The AI-assist pulled the last deployment log, cross-referenced the recent schema change in our `migrations/` directory, and spat this out:
```sql
-- AI suggested the dev run this to see the diff between prod and staging schemas:
SELECT column_name, data_type
FROM information_schema.columns
WHERE table_name = 'your_table'
AND table_schema = 'prod'
EXCEPT
SELECT column_name, data_type
FROM information_schema.columns
WHERE table_name = 'your_table'
AND table_schema = 'staging';
```
That's actionable. For an external customer, the same question would be a black box. The AI would generate a generic "check your deployment logs" response, escalating frustration.
The metrics prove it. Our internal first-contact resolution with AI-assist is up 40%. External ticket deflection? A pathetic 15%, and 30% of those deflected tickets get re-opened within 24 hours because the AI missed the mark.
The tool is the same. The data diet is everything. Feeding it your company's guts works. Feeding it a PR-friendly FAQ does not.
-- bb
-- bb
Yeah, this is a great point. The internal use case is a no-brainer for me too. The difference is the AI isn't just guessing from a vague KB, it's got hooks into the actual infrastructure. That staging deployment example you gave is exactly the kind of time-saver that makes the cost of running the LLM pay for itself vs. an SRE burning 45 minutes digging through log groups.
One thing I'd add though: the internal win is only as good as your observability pipeline. If your logs are scattered across five different tools or your runbooks are stale, the AI will just hallucinate a plausible but wrong fix faster than a human ever could. We saw that happen when we pointed ours at a half-baked Datadog dashboard - the AI confidently suggested a scaling change that would have blown up our lambda concurrency limits. So the real sweet spot is when you've already got clean, structured observability data that the AI can query without hitting auth walls or parsing markdown from 2019.
cost first, then scale
Absolutely nailed it with the point about observability pipelines. That's the hidden cost right there. We learned the hard way after our MongoDB to Postgres migration - we had the AI set up with access to our runbooks, but those runbooks still referenced the old Mongo connection strings and collection names. The AI, trying to be helpful, generated a repair script using a `mongodump` command that obviously failed on the new cluster. It was confidently wrong, just like your lambda example.
So our big lesson was: you can't just point the tool at your docs and call it a day. You need a separate, dedicated, and rigorously maintained "source of truth" context for the AI. We ended up building a small pipeline that takes commit messages from our infrastructure-as-code repo and syncs them as context. It's extra work, but it stopped the hallucination cycle.
I'm curious, after your Datadog incident, did you end up creating a separate, curated dataset for the AI, or did you double down on cleaning up the primary observability sources?
Backup first.
Oof, that migration story hits close to home. We went the other route.
We didn't build a separate curated dataset. The maintenance overhead scared us. Instead, we made cleaning the *primary* sources a mandatory gate for the AI project's success.
Basically, we told the team that if the AI could read it, a human should trust it. It forced a brutal but good docs and dashboard cleanup sprint, which honestly had value beyond just feeding the AI. Our Datadog incident was the catalyst.
It's still a bit messy, but it feels more sustainable than maintaining two separate sources of truth. Curious if that "one source" rule will hold up in six months, though.
—b