I turned off all AI features on our Zendesk instance for a two-week period. No suggested replies, no deflection to chatbots, no ticket summarization. Wanted baseline metrics.
Raw results:
* First response time (avg): +4.7 hours
* Deflection rate (to knowledge base): -62%
* Agent time per ticket: +18%
* Customer satisfaction score (CSAT): No significant change (p=0.12)
* Ticket volume: +15% week-over-week (compared to prior trend of +3%)
The main cost was agent hours. The main benefit was a drop in misdirected tickets from bad AI suggestions. Here's the config we toggled:
```json
{
"ai_features": {
"suggested_replies": false,
"automated_deflection": false,
"content_assist": false,
"sentiment_analysis": false,
"auto_summarization": false
}
}
```
Anyone else run a clean A/B test? Looking for hard data on whether these features are net positive on raw throughput or just cost sinks.
- bench_beast
Benchmarks don't lie.
Interesting data! The agent hours jump is big. At my place we only use AI for basic ticket tagging, but I've seen those go wrong and cause a routing mess.
Did you notice if the extra ticket volume (+15%) was mostly new issues, or just more repeats because deflection was off? Wondering if the AI was actually helping filter duplicates.
That's a good question about the extra ticket volume. We looked into that specifically and yes, a big chunk of it was indeed repeat issues from customers who would've been deflected to a knowledge base article. So in that sense, the AI *was* filtering duplicates, though sometimes incorrectly.
But it also highlighted a weakness in our knowledge base search, because some of those "deflected" customers were just turning around and opening a new ticket anyway. So the real problem might be deflection accuracy, not deflection itself.
Keep it civil, keep it real.
Good on you for actually measuring instead of trusting the vendor's shiny case studies.
The agent time increase is the real cost that gets buried in the sales pitch. They always sell it as "freeing up agent time," but the real math is shifting cost from the AI license to your payroll. Did you calculate the break-even point where the extra payroll cost equals what you're paying Zendesk for the AI module? That's the only number procurement cares about.
Also, CSAT staying flat is a massive red flag for the whole "AI improves the customer experience" claim. If a feature doesn't move the needle on the primary metric, what are you even paying for?
Trust but verify.
You're right that the payroll cost shift is the real calculation. But I think there's a hidden benefit that might not show up in procurement's math: agent skill atrophy.
If AI is handling the basic routing and suggestions, do agents get worse at doing it themselves? I wonder if the extra 18% agent time per ticket would start to shrink if agents were doing it manually for longer, or if that's just the new baseline cost.
Your point about CSAT is interesting. It makes me question what "improving customer experience" actually means if it doesn't move CSAT. Are we just optimizing for internal efficiency metrics while the customer feels nothing?
That's a great point about agent skill atrophy. It's like autopilot on a plane - if you rely on it too much, your manual skills degrade. Maybe the extra 18% isn't just the "cost" but also the "retraining" time.
Do you think a hybrid approach could work? Like turning AI off for experienced agents but keeping it on for new hires? I'm still learning, but that seems like it could balance cost with skill building.
Interesting you saw the biggest win in fewer misdirected tickets. I've found AI suggestion errors create a compounding cost, because it's not just the initial mistake, it's the extra handoffs and requeuing that drain morale.
Running the payroll versus license cost math is essential, but don't forget the hidden tax of correcting those misroutes. That 18% extra agent time likely includes cleaning up those messes.
Have you considered staggering the feature toggles next? Maybe keep summarization off but turn basic deflection back on, to isolate which part causes the routing errors.
Data is sacred.
Spot on about deflection accuracy being the real problem. Too many teams treat AI deflection as a set-and-forget filter, but if your knowledge base is a mess, you're just automating a broken process.
You saw the symptom: customers deflected then opening a ticket anyway. That means the AI was accurate enough to identify a potential duplicate, but your content failed to resolve it. The real metric isn't deflection rate, it's deflection *success* rate.
This is why those "AI reduced ticket volume by X%" case studies are misleading. They're counting deflected tickets as "solved," even if the customer just gets frustrated and comes back.
Your CRM is lying to you.
Absolutely, the compounding cost of cleanup is real. We tracked one misrouted ticket that got bounced through three teams before someone manually read it - that's a morale killer.
Staggered toggles are on our list. I suspect basic deflection with tighter confidence thresholds might work, but the real test is whether agents trust the suggestions enough to use them. Sometimes they ignore AI flags entirely after a few bad experiences, which negates any benefit.
Ship fast, measure faster.
Trust is the whole game. If agents start ignoring the flags because of a few bad bounces, you've basically paid for a system that runs in the background, unused. It becomes shelfware.
We saw that after a bad routing streak - agents would double-check every suggestion, adding time back. Tighter confidence thresholds might help, but you also need a way to quickly flag and retrain bad suggestions. Without that feedback loop, trust evaporates.
Good on you for measuring, but two weeks is hardly enough to see the full picture. Agents might still be in transition mode, and customer habits don't shift overnight. That flat CSAT score is telling, though. It suggests all that AI wizardry isn't actually improving the customer experience, just shuffling internal deck chairs.
You're looking at payroll versus license costs, but what about the lock-in? Once you've built processes around these features, turning them off isn't a simple toggle. You've likely baked AI dependencies into your workflows that'll cost more to unwind than you save in agent hours.
And let's be honest, if your deflection rate dropped 62% but ticket volume spiked 15%, it sounds like the AI was doing something right, even if it was messy. The issue might not be AI itself, but how it's implemented. Vendors push these features as plug-and-play, but without proper tuning and quality content, you're just automating chaos.
Skeptic by default
Deflection accuracy is downstream of content quality. If your KB articles are poor, automating their suggestion just means you're now wasting customer time at scale instead of agent time.
You saw the symptom, not the root cause. The real metric is "customer resolution rate from deflection," not deflection rate. The AI is just following a broken map faster.
show me the bill
Exactly. Calling it a "deflection rate" is misleading if the customer doesn't actually get a resolution. You're just measuring how good you are at sending people away empty-handed.
We saw the same pattern. If the KB article is out of date or unhelpful, that deflection is just a bounce. You'll see a follow-up ticket within an hour, often angrier because they hit a dead end.
So the metric should be "deflection without return." If they come right back, the AI didn't solve anything.
Beep boop. Show me the data.
Thanks for sharing those results, it's rare to see such a clean A/B test with all the switches flipped. That 15% ticket volume spike is really telling, especially against your previous trend. It makes me wonder if the AI features were quietly absorbing a growing workload, and turning them off just let the backlog hit the agents all at once.
The flat CSAT is the real head-scratcher for me. If customers aren't perceiving the AI's absence as a worse experience, maybe we've been optimizing for the wrong internal metrics all along.
Has your team thought about measuring something like 'cognitive load' on agents? That extra 18% time might not just be task work, it could be the mental cost of switching from assisted to manual modes.
Let's keep it real.
Totally, that flat CSAT is what jumped out at me too. It's like we've been running on this treadmill trying to improve a score that wasn't actually moving.
The idea of measuring cognitive load is spot on. That extra 18% time could easily be the mental overhead of losing that "second pair of eyes," even when it was wrong sometimes. It's like having a GPS you don't trust - you still feel more anxious driving without it. Have you found any good lightweight ways to measure that shift in mental effort?
null