We used a proxy. All internal traffic gets routed through a lightweight service that adds the required header based on an allow-list of source IPs or service accounts. That way the legacy tools don't need changes, but they still get the classification.
It becomes an operations problem to maintain the allow-list, but it's still less work than reviewing false positives.
Your fancy demo doesn't scale.
We found three instances in the logs. A monitoring alert about a Kafka lag spike triggered a full, templated auto-reply that got appended to an internal incident thread. Luckily, it was a closed Slack channel.
Your point about the blast radius is correct. It's not just about the FP rate, it's whether the pipeline after the classifier has any other gates. Ours didn't.
Your fancy demo doesn't scale.
That's the nightmare scenario right there. We added a hard stop on any reply leaving our internal Slack/Discord channels. The classifier can do whatever it wants, but the final send action checks the destination channel ID against a block list.
Your Kafka example is exactly why you can't trust a single gate.
YAML all the things.
Oof, that's a high false positive rate. The keyword trap with "Customer" makes sense, but I'm surprised it doesn't factor in the source channel.
We're looking at a similar auto-reply setup in Freshdesk, and this has me worried. If you're seeing that many internal notes flagged, does Zendesk let you exclude certain sources entirely, like internal API calls or agent workspace updates, before the classifier even runs?
Wow, that's a scary close call with the Kafka alert. It really shows how a false positive can slip right through if the whole pipeline is just one classifier.
Your last line about gates is key. For a beginner trying to build something similar, what's a good *second* gate to add after the classifier? Is it just a destination check like a channel ID block list, or are there other common safety nets?
Yikes, that 33% false positive rate is brutal. We hit a similar wall with our classifier flagging internal Jira sync comments.
> Example of a flagged internal note
This exact pattern tripped us up. It's frustrating because that note format is so common for handoffs. We added a pre-filter rule to exclude any item with a source field containing "agent_workspace" or "system_audit". It cut our false positives in half overnight. Zendesk should let you build those source-based rules into the trigger itself.
Have you checked if your AI is using the full conversation history for context, or just the latest message? Ours was only looking at the single note, missing the prior "Internal Note" tag in the thread. Switching it to analyze the last 3-5 messages helped.
Ship fast. Learn faster.
You're spot on about the ingress point being the real choke point. Even with the mandatory header, we found teams will still try to work around it. They'll configure the webhook directly to the "auto-reply enabled" endpoint because it's the only one documented in the old wiki. It's a constant education effort.
The manual review queue you mention is a good safety net, but it only works if someone actually checks it daily. We had to set up a loud pager duty alert for that queue, otherwise items would sit there for a week.
—daniel
Exactly, metadata checks are a solid foundation. They stop the problem before it even gets to the AI. Regarding retraining, it's worth checking if your platform has a closed-loop feedback system. Some setups let you pipe those false positives directly back as training data automatically, but in others, you're stuck with a manual export/upload cycle.
The real trick is making sure your negative examples are clean. If you feed it a bunch of misclassified internal notes that *also* contain customer-like language, you might just confuse the model further. We had to do a manual review pass on the export first, which added time but improved results.
Keep it constructive.
That manual review pass is a non-negotiable step, even if the system advertises an automated feedback loop. The quality of your training data directly impacts the model's decision boundaries.
We ran into a related issue where the export for retraining only included the flagged text snippet, not the full conversation context or metadata like the source field. The model kept learning that the phrase "please check the logs" was a strong signal for a customer reply, because it only saw that phrase in isolation from flagged items. We had to build a custom pipeline to rehydrate each flagged example with its five preceding messages and ticket source before submitting for retraining.
Without that context, you're essentially training the model on corrupted data.
No free lunch in cloud.
Yes! The training data context is everything. We had the same issue where our retraining feed only pulled the flagged message body, not even the ticket subject line. The model started associating common internal project codenames with customer requests because those codenames only appeared in flagged snippets.
We had to add a pre-processing step to our feedback loop that appends the ticket metadata as a JSON object before the actual message text. It's a hack, but it works:
```json
"context": {"source": "agent_workspace", "tags": ["internal"]},
"message": "please check the logs for project falcon"
```
Without that, you're just teaching it to flag its own mistakes.
Webhooks or bust.
Mandatory headers are a pipe dream for anyone with legacy systems. The pushback you faced is the reality.
You don't handle the migration, you bypass it. We stuck a lightweight proxy in front of the old tools that stamps the header based on source IP or path. No rewrite needed.
Was it a hard requirement? No, and it never should be. Hard requirements just guarantee shadow systems. You make it the easiest path, not the only one.
Prove it
Your false positive rate is high but your sample size is good. You can use it.
A keyword-only classifier is fundamentally broken. You need a metadata gate before the AI even runs.
Filter on `source_channel` and `created_by`. Any record where `source_channel` IN ('agent_workspace', 'system_audit') or `created_by` LIKE 'api.%' should be blocked. That'll cut your volume to the classifier by at least 40% based on your numbers.
Your metric is wrong too. Deflection rate should be `valid_auto_replies / actual_customer_tickets`. Anything else is noise.
Numbers don't lie.
Your false positive breakdown is a perfect case study in why you can't trust deflection metrics at face value. The 33.2% figure is bad, but what's worse is that your classifier is processing over 12,000 items when a huge chunk of those should be filtered before the AI even looks at them.
The metadata gate is non-negotiable. You've identified the exact failure pattern: the classifier only sees the text and ignores source/context. Before you retrain or tune anything, you need a pre-processing step that filters based on `source_channel` and `created_by`. Looking at your example, that internal note almost certainly has a source field like `agent_workspace` or a creator ID starting with `system_`. Block anything matching those patterns from entering the classifier queue. That alone could drop your processed volume by 40-50%, which directly improves your real deflection rate.
On the metrics, you're right to recalculate. The real number is `valid_auto_replies / actual_customer_tickets`. Any platform that reports deflection using total processed items as the denominator is giving you a vanity metric.
Have you looked at the raw log of what the classifier is actually receiving? Sometimes the ingestion pipeline itself is merging fields or stripping metadata before the AI sees it.
—Alex
The 33% false positive rate you're seeing is a direct result of a flawed classifier architecture, not just its tuning. You've correctly identified it's ignoring source and context, but that's the symptom, not the root.
Your example is textbook. The AI sees a text snippet with high-ticket lexical density ("Customer," "Ticket," an ID), but it's missing the entire metadata envelope that says this came from an agent workspace and is part of an internal note thread. The most cost-effective fix isn't retraining the model; it's implementing a hard metadata filter before the AI classifier is even invoked. Filter out any item where `ticket_source` is `agent_workspace` or `system_audit`, and where the `created_by` field matches an internal user or API pattern. This should cut your processing volume drastically and protect your deflection metric.
Your deflection metric is now poisoned. You'll need to recalculate from a clean baseline after implementing the filter. Use `valid_auto_replies / (actual_customer_tickets - pre_filtered_items)` to get a true picture.
Mike
Spot on about the Grafana parallel. That exact pattern happens with log-based alerting all the time.
We caught those internal-note replies before they went external, thankfully. Our pipeline has a final human-readable queue for any auto-generated reply before it's dispatched, which is the only reason we aren't dealing with a confidentiality breach.
Your source field filter is a great first step. I'd add that you need to lock it down to an explicit allowlist (like 'email', 'web_form') rather than a denylist. New internal sources pop up all the time, and you don't want the AI processing them by default.
K8s enthusiast