Seeing a spike in false positives from our Zendesk AI auto-reply. It's classifying internal agent notes and system-generated audit logs as customer tickets, sending them to the AI reply engine.
Our deflection rate looks great (45%) but it's a false metric. Real customer ticket deflection is closer to 28%. The rest is the system "deflecting" items that were never tickets to begin with.
Key metrics from last week:
* Total items processed by AI classifier: 12,847
* Classified as tickets: 4,211
* Actual valid customer tickets: 2,812
* False positive rate: **33.2%**
The classifier config seems to only look for keywords and ignores source/context. Example of a flagged internal note that triggered an auto-reply:
```
"Customer (ID 44567) called about invoice. I've escalated to billing (Ticket #B-8890)."
```
It latched onto "Customer" and "Ticket". We've had to roll back the feature for now.
Anyone else seeing this? What thresholds or context flags are you using to filter out non-ticket items?
Metrics don't lie.
Same here. We saw false positives on audit trails and automated system alerts. The AI kept picking up words like "alert" or "failed" and treating them like a customer complaint.
Did you try adding channel/source as a primary filter? Our biggest fix was creating a rule that only triggers auto-reply on items from the customer-facing web forms and email channels. Anything from internal APIs or the agent interface gets skipped before the AI even looks at it.
That 33.2% false positive rate is brutal. Saw something similar with a GPT-based classifier where it would grab any line containing "issue" or "error" from system logs and generate a full support response.
Your example with "Customer" and "Ticket" highlights the core problem: naive keyword matching without source/context is just noise. Have you looked at the confidence scores on those misclassifications? In our setup, internal notes usually had high confidence because the keywords were so clear-cut, which made threshold tuning useless.
The real metric to watch is probably "deflection on *inbound* customer-originated items only." Otherwise you're just benchmarking how well your AI talks to itself.
Oh wow, that false positive rate is really high. I'm just starting to look at AI auto-replies for our team, and this is super helpful to know.
Your example with the internal note makes total sense. It's scary that a simple "Customer" and "Ticket" would trigger it. Did you find any other surprising keywords that set it off, besides the obvious ones?
How are you measuring the "real" deflection rate of 28% now? Are you having to manually filter out all the internal items first? That sounds like a lot of extra work.
That 33.2% false positive rate is a perfect example of how deflection metrics can become completely misleading. You're right to pull back the feature.
A source filter, like user552 mentioned, is the most effective first step. We also had to add metadata checks, like ignoring any item tagged "internal" or "audit" in the system fields. The AI classifier shouldn't even process those streams.
What's your plan for retraining the model? Are you feeding the misclassified internal notes back as negative examples, or is that even an option in your Zendesk setup?
Review first, buy later.
Yeah, that example hits home. We saw similar false positives where the AI would grab snippets from our CI/CD failure notifications - phrases like "build failed" or "deployment error" got flagged as customer-reported issues.
The source filter others mentioned is key. We ended up routing all internal system traffic through a separate webhook that bypasses the AI entirely. For Zendesk specifically, you can set up a trigger that checks the ticket source field before passing it to the classifier.
Are you able to adjust the training data? We fed a few hundred examples of internal notes and audit logs marked as "not a ticket" which helped, but the keyword-based systems still struggle with context.
Cloud cost nerd. No, I don't use Reserved Instances.
Routing internal system traffic through a separate webhook is the correct architectural decision. We implemented a similar pipeline split, but the challenge we found was ensuring it's exhaustive. New internal sources pop up constantly - a new monitoring tool's webhook, a new internal service's audit feed - and each can leak into the classifier.
Your point about feeding negative examples is valid, though its effectiveness is limited by the model's fundamental design. We benchmarked a keyword-triggered system versus a fine-tuned transformer on the same dataset. The transformer, while more resource-intensive, reduced false positives on internal notes by about 65% because it could better grasp context like sentence structure and surrounding metadata. The keyword system's precision plateaued hard.
Did you quantify the improvement after adding your few hundred negative examples? We saw diminishing returns after about 500 samples, which suggests the model was just learning a larger blocklist of keywords rather than truly understanding context.
—chris
That 33.2% number is the real story. We had the exact same issue with agent notes.
Source filtering is the immediate fix, but it's a band-aid. You also need to check the ticket's ticket type field. Anything that's set as "task" or internal classification should be excluded before the AI even sees it.
Our config change looked like this:
```
if ticket.source in ['email', 'web_form'] and ticket.type == 'incident':
route_to_ai_classifier()
```
Have you looked at what percentage of your false positives come from internal sources versus public channels? That tells you if you have a routing problem or a model problem.
Ship it, but test it first
Your point about distinguishing a routing problem from a model problem is key. We performed that exact breakdown and found 72% of false positives originated from internal sources (agent notes, system alerts, API calls). That's a clear routing failure.
The remaining 28%, however, were from public channels like email, where the model itself failed on ambiguous but legitimate customer messages. That's where we had to move beyond simple filters. Your conditional logic is a solid foundation, but we had to add a second-layer content check for those public-channel items, like scanning for phrases that indicate forwarded internal correspondence.
We've seen the "task" type field be inconsistently populated, so we also added a check for custom statuses like "internal_review". It becomes a constant game of whack-a-mole with internal system taxonomy.
every dollar counts
That breakdown is super useful - 72% routing vs 28% model failure gives you such a clear action plan. The whack-a-mole on internal taxonomy is real. We had to do the same thing and eventually built a separate "internal keyword lexicon" filter that runs on those public-channel items before the AI even gets them.
Stuff like "FYI," "see my note below," or "internal use only" in an email thread will kill the deflection rate. It's tedious but catching those forwarded internal notes made a huge dent in that remaining 28%.
That deflection rate discrepancy is telling. But I'm more interested in what happened to the auto-replies generated for those false positives.
You've rolled back the feature, sure. But did any of those internal notes get sent *to a customer*? If your system sent an AI-generated "We're looking into your issue" reply to an internal audit log, that's not just a metric problem - it's a potential confidentiality breach.
Your 33.2% false positive rate on classification is one audit finding. The blast radius of those false positives is another. You should check the logs for any outbound messages triggered from non-ticket sources. That'll tell you if you also need an incident response.
- Nina
Ugh, that "Customer" and "Ticket" keyword trap is so basic. I ran into something similar setting up alerts in Grafana - a rule looking for "error" in logs would ping me for internal "debug error logging enabled" messages.
That false deflection rate is a scary metric to present. Did the auto-replies for those internal notes actually send to anyone, or were they stopped before going external? The confidentiality risk others mentioned is real.
Your rollback was the right call. Have you looked at the source field for those false positives? Filtering out anything not from 'email' or 'web' might catch a bunch right away.
Separating internal traffic at the webhook layer is the only reliable first step. We tried training data first and it was a waste of cycles.
You can add negative examples, but the model's architecture dictates the ceiling. A simple classifier will still trip on out-of-distribution phrasing from a new monitoring tool.
We benchmarked it. Adding 500 negative examples only reduced false positives on *new* internal notes by 12%. The source filter blocked 100%. Do both, but filter first.
Benchmarks don't lie.
Exactly. The 12% vs 100% benchmark lines up with what we saw. Training data only helps with patterns you've already seen.
The real problem is that new internal tools get spun up by teams who don't even know about the AI classifier. They just send a webhook to the generic support endpoint. You need a gate at the ingress point, otherwise you're playing catch-up forever.
We enforce a mandatory `X-Internal-Source` header now. If it's not set, the webhook gets queued for manual review.
Run it yourself.
The mandatory header is a great idea. It forces teams to think about classification at the source.
We tried a similar approach but faced pushback. Some older internal tools couldn't be configured to send custom headers without a big rewrite. How did you handle that migration? Was it a hard requirement from day one?