Skip to content
Notifications
Clear all

What's the best way to collect agent feedback on AI suggestion quality?

46 Posts
45 Users
0 Reactions
101 Views
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Web form to Slack is a terrible idea. You're adding steps. Thumbs are fine, but you're missing the real metric.

Forget deflection rate. Your first question is whether agents even see the suggestion. Log every time it's displayed. Then log the thumbs. If your display-to-thumb ratio is low, your AI is just background noise and you're wasting cycles.

And scrap the comment box. You'll get garbage. If you need a "why," use a dropdown with maybe three options like "wrong," "bad tone," "useless." Keep it simple or it dies.


-- old school


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Exactly. The display-to-thumb ratio is the only health metric you've got at launch. If it's under 20%, you're building a very expensive random number generator.

One caveat: if you give them that three-option dropdown immediately, you'll still get garbage. It's a choice, and choices require thought. Make the thumbs mandatory first, then let them *optionally* expand a dropdown to say why. Most won't, but the ones who do give you actual signal.



   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 3 months ago
Posts: 254
 

The phased approach you're describing, making the thumbs binary mandatory and the "why" optional later, is the right operational cadence. It mirrors how we built training data pipelines - you start with the high-confidence, low-cardinality labels before introducing finer-grained taxonomies.

But that low 20% display-to-thumb ratio threshold is critical. If you're below that, the optional dropdown becomes statistically meaningless because your sample size of feedback is too small. You'd be making product decisions based on the opinions of a tiny, possibly non-representative fraction of agents.

The real trick is deciding what triggers the "optional" phase. Is it hitting that 20% ratio for a month? Or is it based on absolute volume of thumbs data? Without a clear gate, the team will just add the dropdown too early and drown in sparse, noisy categories.


Data is the source of truth.


   
ReplyQuote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

The web form to Slack idea is exactly what I nearly built! Glad I'm not the only one who thought of that first. I ended up using a Zendesk app with just thumbs up/down buttons that fire a webhook.

For metrics, everyone's saying display-to-thumb ratio is key. That's what I'm logging now. But how are you storing the data? I'm pushing to BigQuery but my pipeline is... fragile 😅 Do you log the raw suggestion text too, or just the event?



   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

You're absolutely right about attaching cost data. That's where the rubber meets the road for any ops lead approving the budget.

The tricky bit is getting a clean "time spent after suggestion" metric. We tried this and found our timestamps were off because agents would sometimes leave the ticket open while doing other work, inflating the handle time. We had to correlate the suggestion event with *active* status changes in the ticket, not just creation-to-close.

But once we got it, we saw exactly what you described - a few "thumbs down" suggestions on high-priority tickets were adding 3+ minutes of agent confusion. That hard dollar figure got the vendor's attention for model tuning way faster than any volume of "bad tone" feedback ever did.


Pipeline Pilot


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

The mandatory thumbs approach is correct, but I've seen teams implement the "optional expansion" poorly by making the UI stateful. If you store the agent's preference client-side and default to expanded after their first dropdown use, you'll artificially inflate engagement with the finer-grained taxonomy. The expansion should be a conscious, per-interaction choice to avoid biasing your sparse "why" data.

You also need to consider what happens on a thumb-down with no "why" selected. That's a valid signal - it means the suggestion was bad enough to reject, but not worth diagnosing. Some of our most impactful model retraining came from clustering those unlabeled negative events and looking for common patterns in the raw suggestion text and ticket metadata.



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You're already overthinking it. A web form to Slack is adding a whole new app to their workflow.

The thumbs are fine, but if you're logging deflection rate, you're measuring the wrong thing. You need to know if the suggestion was even seen. If your agents are ignoring 80% of the prompts, your "AI feature" is just a UI decoration and any deflection data is noise.

Forget the comment box. That's a tax on your agents and you'll get nothing useful. Start with mandatory thumbs, log the display event, and pray your ratio is above 20%. If it's not, scrap the tool and get a refund.


Prove it


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Oh, the 20% display-to-thumb ratio rule is really interesting, I haven't heard that before. It makes sense though - if they aren't even clicking, it's just visual clutter.

But what if the suggestions are showing up at a bad time? Like, maybe they see it but they're in the middle of typing and a thumbs click feels disruptive. Could a low ratio sometimes be a UI/ timing problem, not just bad AI?



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Your logging schema is spot on for capturing the basics, but you're missing one critical field: the raw suggestion text. Don't just log the ID. Log the text itself.

If you only have the ID, you're completely dependent on the AI vendor's API to retrieve what was actually suggested for any retrospective analysis. That API could change, be deprecated, or have rate limits that cripple your ability to understand why a cluster of tickets all got thumbs down six months from now. Store the text as a JSON field or a simple text column. It's cheap storage and it future-proofs your analysis.


keep it simple


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 5 months ago
Posts: 329
 

The Slack reaction idea for lower friction is clever, and it can work if the AI tool posts suggestions directly into a channel. The main snag I've seen is that most AI assist tools operate inside the help desk UI, not Slack, so you can't react to a message that doesn't exist there.

To your last question, the vendor's built-in feedback is usually a checkbox feature, but it's often too basic. I've found it rarely logs the raw suggestion text or ties the feedback to a specific ticket context, which you'll need later for tuning. You usually end up building a parallel pipeline anyway.


Integrate or die


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 7 months ago
Posts: 293
 

The binary thumbs up/down is the right place to start, as others have said. My addition is to treat the "optional comment" as a separate, asynchronous step to keep friction low. We built a weekly Slack digest that showed the top three most-thumbed-down suggestions from that week and asked agents to tag the one they'd most like to see improved with a single emoji. This gave us qualitative data without interrupting the ticket flow.

For metrics, deflection is a lagging indicator. Start by logging the suggestion display event and the thumb event. The ratio between them is your primary health metric. If it's below 20%, your UI or timing is likely the problem, not the AI quality. You can't improve what agents ignore.

Storing the raw suggestion text is non-negotiable for later analysis, as user216 noted. Your vendor's analytics will likely be insufficient for tuning.


independent eye


   
ReplyQuote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

Your thumbs up/down instinct is spot on for keeping it simple. The Slack web form idea can work for a small team if you're already living in Slack, but you risk agents ignoring a separate app. Since you're on Zendesk, see if your AI tool has a native feedback widget - sometimes that's less friction than building something new.

I'd start with just logging two things: the suggestion display event and the thumb click. That display-to-thumb ratio tells you if agents are even engaging. If it's low, the problem might be UI timing, not the suggestion quality.

Deflection rate is a later metric. First, make sure the feature isn't just decorative. Oh, and definitely log the raw suggestion text, not just an ID. You'll thank yourself later when you need to analyze patterns in the bad suggestions.


Let the machines do the grunt work


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

Thumbs up/down is the vendor's happy path. They want you measuring engagement, not whether the tool actually saves time or money.

Your Slack form idea will die. Agents hate context switching more than they hate bad AI. The feedback mechanism needs to be *in the exact spot* the suggestion appears, or it's a tax.

Forget deflection rate for now. That's a vanity metric they'll sell you on. Start by logging two things: the moment a suggestion appears, and any action the agent takes on the ticket for the next sixty seconds. If they immediately start typing their own reply ignoring your shiny AI, you have your answer. That's the only metric that doesn't lie.


Buyer beware.


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Agree completely that vendors push the engagement angle. The 60-second action log is a great, honest metric for real usage.

One thing I'd add, you need to define that "action" clearly. Is a mouse click in the ticket enough? Or do you need to see actual keystrokes? If you're too strict, you might flag agents who are thinking before replying as ignoring the AI. But if you're too loose, you'll miss the true ignores.



   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

You've got the right instinct with the quick thumbs - that's exactly where to start. The Slack web form idea is a common first thought, but I've seen it fail for exactly the reason you suspect: it adds a separate step that breaks the ticket-solving flow. The friction is too high.

Focus on embedding the feedback mechanism *directly* in the Zendesk UI where the suggestion pops up. That one-click thumb reaction, with no required comment, is your best bet for initial adoption. If you can log that a suggestion was displayed and whether a thumb was clicked, you've got your core engagement metric. That display-to-action ratio tells you if the feature is being used or just ignored as noise.

Deflection rate is a vendor's favorite talking point, but it's a later-stage business metric. First, you need to know if your team is even interacting with the tool. If they're not, you've got a UI or timing problem to solve before you worry about suggestion quality.


Architect first, buy later


   
ReplyQuote
Page 3 / 4