Hey everyone,
We've been testing the new AI-suggested replies on our internal support team for a few months now. The agents are generally positive, but we kept noticing a pattern: the suggestions that landed us in hot water—where a customer would reply with "That didn't answer my question"—were almost always the ones served up with a confidence score between 60% and 80%.
Below that, agents naturally ignore them. Above 80%, they're usually solid. That middle band is the danger zone. Agents feel pressured to use a "pretty good" suggestion, but it's just incomplete or slightly off-topic enough to cause a deflection.
So, our dev lead and I just built a simple workflow into our dashboard that flags any incoming suggestion with a confidence score below 80%. It doesn't block the agent from using it, but it adds a visual caution and a quick "Review carefully" note. The idea is to prompt a deliberate pause.
We're only a week in, but early feedback is that it's reducing "bad" sends. Agents say they appreciate the nudge to tweak the wording or check the ticket history one more time.
Has anyone else tried setting a confidence threshold for human review? I'm curious if 80% is the right bar, or if it varies wildly by platform (we're on HelpGrid). More importantly, how are you balancing the speed boost of AI with the risk of those mediocre suggestions?
Welcome! Let's keep it real.
That's a pragmatic approach, and the 80% threshold is a decent starting line. But be prepared for it to drift.
Our experience with a similar scoring system for code generation was that the "danger zone" shifted over time as the model was updated. What was 70-80% last quarter became the new 75-85% after a retraining cycle. The absolute number is a proxy for something you can't see directly, so it's inherently slippery.
You might want to track the false-positive rate on your 80% flag over time. If it stays noisy, you're good. If agents start blindly trusting everything above the line because the flag is always there, the threshold's lost its meaning.
No SLA, no problem.