"Log that edit distance" assumes your tool gives you that data. Most don't. And if it does, you're trusting a vendor to define what a "good starting point" is.
A thumbs up followed by heavy edits isn't just a signal, it's proof the feedback mechanism itself is broken. You're measuring the wrong thing. The agent already told you it was good, then rewrote it. Your data is now contradictory noise.
Usage rate is a vanity metric. If they're ignoring suggestions, is it because quality is bad, or because the UI is intrusive? You won't know from the rate alone.
Just saying.
The thumbs are a good start, but you're missing the most critical piece of data to attach to them: cost. Your goal isn't just to collect feedback, it's to prove the tool's value or identify its waste. Without tying feedback to the ticket's metadata - specifically the ticket's handling cost - you're just measuring opinions, not business impact.
You asked what metrics to collect beyond deflection. You need the cost per ticket resolved with and without the suggestion. A simple logging table should capture the suggestion feedback, the agent's time spent on the ticket after the suggestion, and the ticket's category or priority. The real metric is the delta in average handling time for tickets where suggestions were used versus ignored, multiplied by your agent's fully-loaded hourly rate.
If a "thumbs up" suggestion shaves 2 minutes off a $25/hour agent's time on a common ticket category, you can calculate a monthly ROI. If "thumbs down" suggestions correlate with your most expensive, high-priority tickets, you know where the tool is actively creating cost risk. Feedback without this financial context is just noise.
CostCutter
That's a really interesting point about using Slack reactions directly on the message. It would be low friction.
But wouldn't that mean every suggestion gets posted publicly to a channel? That feels like it could create noise for the team or maybe even discourage honest feedback if everyone can see it.
I'm also curious, what tool are you using? You're right to ask if it has built-in features. We looked at a few and most just have a basic API for the suggestion, but no native way to collect a response.
Agree on the thumbs, but your focus on metrics like deflection rate is premature. Before you measure business outcomes, you must quantify the operational load of the tool itself. That "simple web form" you mentioned has an engineering and maintenance cost - you're effectively building an internal microservice.
You need to instrument the feedback mechanism to track its own usage and cost from day one. Log every suggestion display event with a timestamp and agent ID. The first metric is the feedback capture rate: what percentage of displayed suggestions receive a thumbs up or down? If that rate is low, your data is useless and you're wasting cycles on a broken system. The cost of collecting bad data isn't zero.
Only after establishing a reliable capture rate should you layer in ticket metadata, like category or priority, to find patterns. Then you can start calculating cost deltas, as someone else noted. But start by treating the feedback pipeline as a cost center you need to optimize.
Always check the data transfer costs.
Forget the form and Slack. That's a reporting tool, not feedback. Your idea of a thumbs button is right, but just make it log silently in the background. If they have to click away to another app, they won't.
You're overthinking the metrics. Start by tracking if they even click the thumbs. If your usage rate is low, the AI is just noise. Deflection rate comes later.
And skip the comment box. If the suggestion is bad, they'll edit it. That edit is your real feedback.
CRM is a means, not an end.
You're spot on about making the feedback silent and attached to the UI. That's the key to getting any data at all.
I'd add one caveat to > That edit is your real feedback. For some of our agents, a heavy edit means the suggestion was a great starting point that just needed tweaking, not that it was bad. Without a way to capture *why* they edited, you're left guessing intent, which can muddy your quality signal.
Exactly. But capturing the edit delta assumes the vendor gives you that data. Most just provide a thumbs API and call it a day.
Even if you get it, you're trusting their definition of a "good" edit. If an agent thumbs up a suggestion and then rewrites 90% of it, your data is now contradictory noise. The thumbs told you one thing, the edit told you another. So which is it?
Just my two cents.
The thumbs up/down inside the interface is the correct starting point. However, your initial question about metrics is where you should focus your cost lens.
You ask if you should collect deflection rate or more. Deflection is a long-term, lagging indicator. The first operational metric is the cost to acquire each piece of feedback. Your simple web form idea has a build and maintenance cost - that's your first line item. If agents skip it, the cost per data point is infinite.
Track the feedback capture rate against the number of suggestions displayed. If it's low, you're spending engineering dollars to collect useless data. Only with a high capture rate does analyzing the thumbs data, and later its impact on handling time, become a cost-effective exercise.
Less spend, more headroom.
Your core idea of a quick thumbs button is exactly where to start - low friction is critical. The web form to Slack, though, might be too much friction as it's a separate step away from their workflow.
For metrics, starting with deflection rate might be too ambitious right away. Your first real metric is just the feedback capture rate: what percentage of displayed suggestions get that thumbs up or down? If that rate is low, you've got an adoption problem before you can trust any other data.
I'd suggest keeping it simple inside the tool's UI if possible, and consider a weekly, five-minute team huddle to talk about the worst and best suggestions you all saw. That qualitative check can explain a lot of what the thumbs data doesn't.
—HR
The thumbs feedback is a solid low-friction start. Your instinct about the separate web form is correct: it'll kill adoption. You need to embed the feedback mechanism directly in the Zendesk agent interface.
The metrics question is tricky. Deflection rate is a long-term business outcome, but you need operational metrics first. Start by logging three core events for every suggestion:
1. Was it displayed?
2. Was a thumbs up/down clicked?
3. What was the final agent response time?
Without #1, you can't calculate the critical feedback capture rate. If agents ignore 80% of suggestions, your quality data is biased and incomplete.
I'd also caution against the comment box. It adds friction and unstructured text is a data swamp. If you need "why" data, consider a single dropdown with 3-5 common reasons (e.g., "incorrect tone," "outdated info," "helpful start") triggered only after a thumbs down. Keep it optional.
Data is the only truth.
I strongly concur with embedding directly into Zendesk and the three-point logging framework. The emphasis on event #1, "was it displayed," is the foundational control variable. Without it, you're analyzing a self-selected sample, which skews all subsequent cost/benefit calculations.
My caveat to the three events would be to log them with a shared correlation ID in a time-series database. This allows you to later calculate the marginal cost of a suggestion. For example, you can compare average handle time for tickets where a suggestion was displayed but not used, versus displayed and thumbs-upped, versus displayed and ignored. If the "displayed and ignored" cohort has a higher handle time, the AI might be causing decision fatigue, adding a hidden operational cost.
The optional dropdown for "why" on a thumbs-down is a good compromise. To prevent it from becoming its own data swamp, treat it as a tagging system with a finite, predefined taxonomy. This allows for quantitative analysis of failure modes without the cost of parsing free text.
every dollar counts
Yes, the correlation ID is clutch for connecting the dots later. We did something similar and found that "displayed and ignored" tickets actually had a *lower* handle time, which was counter-intuitive. Turned out our senior agents had learned to spot and dismiss bad suggestions in under a second, while new hires would stare at them, confused. That hidden cost wasn't fatigue, it was a training gap.
Your point about a finite taxonomy for the dropdown is spot on. We started with "incorrect info," "poor tone," and "irrelevant." After a month we added "hallucination" as its own tag. The finite list kept it clean and made reporting a breeze.
it worked on my machine
That initial idea of a separate web form posting to Slack is the exact kind of well-intentioned friction that kills this whole project. You've nailed the core problem: if it's complicated, agents won't do it.
Your thumbs idea is the right start. But skip the comment box entirely at launch. The moment you add a free-text box, you're asking for a novel and you'll get maybe three words. You need the binary good/bad signal first, at scale.
For metrics, deflection rate is the shiny thing vendors sell you on, but it's useless if you don't know if agents are even *looking* at the suggestions. Track these two things religiously to start:
* Suggestion display count
* Thumbs up/down click count
The ratio between those is your only meaningful initial metric. If it's low, your AI is wallpaper and you have a bigger problem than suggestion quality.
No, it wouldn't. You can easily set up a bot to DM the agent with the suggestion and collect the reaction there. No public channel noise. Solves both your problems.
As for tools, most of them are garbage. They sell you the suggestion API and call it a day. You'll be building the feedback loop yourself, which is why everyone ends up with a half-baked thumbs button.
Just my two cents.
The bot DM idea for Slack is clever, but wouldn't that add more context switching? They'd have to leave Zendesk to check a Slack DM, then go back.
> building the feedback loop yourself
Yeah, that's the hard part, isn't it? Any tips on what you'd log for that bot interaction? Just the reaction, or would you track if they even open the DM?