Skip to content
Notifications
Clear all

Step-by-step: testing deflection accuracy with a blind review team

2 Posts
2 Users
0 Reactions
3 Views
(@jacksonj)
Estimable Member
Joined: 1 week ago
Posts: 64
Topic starter   [#7023]

Hey everyone! I’m pretty new to this SaaS-ops space and have been tasked with evaluating our new support tool’s AI deflection feature. My boss wants to know if the “suggested articles” are actually helpful or just guessing.

Here’s my plan: I want to set up a blind test. I’ll gather 50 recent customer questions, have the AI suggest a knowledge base article for each, and then ask three of our support agents (who don’t know which answer is AI-picked) to rate if the suggestion directly solves the query.

Is this a solid approach? Mainly wondering:
- Should the agents just give a yes/no on accuracy, or a score from 1-5?
- Do I need to include the times the AI offers *no* suggestion in the data?
- How many questions is enough for a meaningful sample?

I’m excited to learn how you all measure this stuff. Thanks!


Thanks!


   
Quote
(@cost_optimizer_99)
Estimable Member
Joined: 3 months ago
Posts: 148
 

50 queries? That's a rounding error. Run at least 200 to see any real pattern.

Yes/no for the rating. A 1-5 scale just gives you ambiguous data to argue over later.

And absolutely include the 'no suggestion' results. That's part of the deflection rate. If it only suggests something 30% of the time, your accuracy on the other 70% is zero.


show the math


   
ReplyQuote