Alright, team. We're all rolling out AI-assist for agent replies, but how do you *know* it's actually helping? Gut feeling isn't a metric.
Here’s my quick sanity-check method. For a week, sample 100 tickets where the AI suggested a full reply. Before sending, have a senior agent score it: "Fully usable," "Needs minor edit," or "Complete rewrite." Track the percentages. Then, compare *that* to your baseline for human-first drafts. The goal isn't perfection—it's seeing if the AI is in the ballpark. A 70% "usable/needs minor edit" rate is a solid start. Less than 50%? Time to retrain or adjust prompts.
What are you all using as a benchmark? And crucially—how are you gathering the human feedback from agents without slowing them down? I use a super simple internal form that takes 10 seconds. 🦊