Skip to content
Notifications
Clear all

Check out what I made: a script to compare auto-reply suggestions vs actual answers

19 Posts
19 Users
0 Reactions
9 Views
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
Topic starter   [#29114]

Hey everyone! 👋 I've been seeing more and more support platforms roll out those AI-powered auto-reply suggestion features. They promise to save agents time, but I was curiousβ€”how *good* are those initial suggestions, really? Are they close to what a human would actually send, or do they just create more editing work?

As a little side project, I built a Python script to compare the platform's auto-reply suggestions against a set of actual, human-written agent responses from past tickets. I focused on two things:
* **Similarity Score:** Using NLP to check how semantically close the suggestion is to the final answer.
* **Edit Distance:** Measuring how many character changes an agent had to make.

I ran it on a batch of about 500 historical tickets from our old system. The early results were... interesting!

* For simple, FAQ-type queries (like "How do I reset my password?"), the suggestions were often 85-90% similar and needed only minor tweaks for tone.
* For complex or emotional customer issues, the similarity score dropped to around 40-60%. The AI suggestion often missed the nuance or provided a too-generic first step that the agent had to completely reframe.
* The biggest time-savers were for straightforward informational replies. The biggest time-wasters were suggestions that looked okay at a glance but subtly misdirected the solution.

I'm still tweaking it, but the script basically helps quantify when the AI assist is genuinely assisting and when it might be adding cognitive overhead. I'd love to hear if anyone else has done similar internal checks or if you're seeing different patterns with tools like Zendesk Answer Bot, Intercom Fin, or others.

What's your team's experience been? Are your agents accepting most suggestions, or are they constantly rewriting them?

Happy benchmarking!


Always testing.


   
Quote
(@emilyw)
Reputable Member
Joined: 3 months ago
Posts: 188
 

Hey, this is a cool project. I'm a support lead at a small e-commerce company, and we've been using Intercom's AI suggestions in production for about six months now.

Here's what I've noticed from hands-on use, which lines up with your findings.

**Use case fit**: For our team of 10 agents, suggestions are great for common, transactional tickets. Think password resets or order status checks. For complex or emotional issues, like a damaged item or a billing dispute, we usually turn the feature off. The generic "first steps" can actually slow us down.
**Real cost**: Intercom's pricing starts around $74 per seat/month on the Pro plan where this feature is included. That adds up, so the time saved needs to be real. For us, it only pays off because we have a high volume of simple tickets.
**Integration and setup**: Getting the AI to sound like us took real work. We fed it a bunch of past replies and had to constantly tweak the guidelines for about a month. It wasn't plug-and-play.
**The honest limitation**: The biggest gap is context. If a customer's previous ticket isn't fully resolved, the AI often suggests a fresh "solution" that ignores the history, which annoys customers. We have to catch and edit those every time.

I'd pick it only if you have a very clear, high-volume stream of FAQ-type questions. For a team dealing mostly with nuanced support, the editing overhead might not be worth it. What's your average ticket complexity, and what's your current tickets-per-agent volume? That would make the call clearer.



   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

You hit the nail on the head about context. We see the same thing with our internal ticketing bots - they'll scan the last message and spit out a "solution" that ignores the three-week saga documented in the ticket history. Ends up creating more work to clean up.

The cost angle is real, too. When management pushed for an AI feature add-on, we did a quick napkin math check: time saved per ticket needed to offset the license cost. For anything but the most repetitive, simple stuff, the ROI just wasn't there. Makes you wonder if the real value is just in shaving seconds off those password reset flows.


NightOps


   
ReplyQuote
(@finnleyj)
Estimable Member
Joined: 2 months ago
Posts: 111
 

Exactly. That napkin math is the only thing that keeps these projects grounded. We went through the same exercise when evaluating one of the big APM vendors' AIOps suggestions. The feature cost was a 20% uplift on the platform fee.

When you break it down, the "time saved" only materializes on a narrow band of alerts, like low-disk warnings. For anything involving a trace or a log correlation, the suggestion was either a generic "check your service health" or, worse, a confidently wrong root cause that sent engineers down a rabbit hole. The cleanup and context-rebuilding effort erased any marginal gain from the simple cases.

The real ROI question they never answer is the cost of the incident that gets prolonged because someone trusted a bad suggestion.


latency is a liar


   
ReplyQuote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 363
 

Fantastic approach. Quantifying that drop-off from 90% to 40-60% similarity is exactly the kind of data procurement teams need to push back on vendor feature-value claims.

One framework I always use in these evaluations is the "Criticality vs. Complexity" matrix. You've effectively shown the AI suggestions work in the low-complexity, low-criticality quadrant. The problem is, that's often where the least business value is saved. The high-stakes, emotionally charged, or technically complex tickets - where agent time and accuracy matter most - are precisely where the suggestion falls apart and requires a full rewrite.

This is the core of the negotiation. If a vendor's pricing bundles this feature, you can use your data to argue for a cost adjustment, since the utility is confined to a narrow, less critical slice of your ticket volume. Have you considered adding a simple "time-to-edit" metric to your script? Even a rough "seconds to accept vs. seconds to rewrite" estimate would give you a direct hour-dollar figure for those high-complexity cases.


null


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

Love that "Criticality vs. Complexity" matrix framing, it's so spot on for these discussions. That drop-off you mentioned, from high similarity on simple stuff to low on complex cases, is the exact pain point for our UX team when we test these tools.

The "time-to-edit" metric is a great idea for procurement, but from a design perspective, I'd also track "friction." We've seen agents develop a kind of "suggestion blindness" where they just automatically dismiss anything beyond a password reset ticket because the mental cost of evaluating a bad suggestion is higher than just typing from scratch. That erodes the feature's value faster than the raw edit time shows.

Have you found a good way to quantify that trust decay?



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 3 months ago
Posts: 214
 

"Suggestion blindness" is a great term for that. We saw the same pattern with our expense report chatbot.

The trust decay is tricky to measure. We tried tracking the "dismissal rate" - how often agents clicked away without using any part of the suggestion. After a few months of bad suggestions for complex reports, the dismissal rate for *all* suggestions shot up, even for simple ones like mileage claims. They just stopped looking.

Maybe the metric is the crossover point where dismissal time becomes faster than evaluation time?



   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

That "dismissal rate" metric is a solid practical measure of trust decay. In our internal trials, we found it correlated strongly with survey feedback about tool fatigue.

You're right about the crossover point. We actually logged milliseconds for "hover-to-dismiss" vs. "read-to-edit" actions. Once agents learned the suggestions were unreliable for a category, their dismissal time for those tickets approached a reflexive click - sometimes under 500ms, which is basically muscle memory. The system was training them to ignore it.

The harder problem is repairing that trust once broken. We had to create a separate "high-confidence" suggestion queue with a stricter threshold, which helped, but it meant the feature became even more niche.


Architect first, buy later


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

Interesting approach. I'd be skeptical of using past tickets as the benchmark for actual answers. Those historical responses are often full of jargon and internal shorthand that a good suggestion tool should actually clean up, not replicate.

Did you weight the edit distance metric? A change from "Hi there" to "Hello" is trivial, but swapping out a core troubleshooting step is a major redraft, even if the character count is similar. Your drop to 40-60% similarity on complex issues just confirms these features are fancy templates, not reasoning tools.

The real test is whether they reduce net time-to-resolution, not keystrokes saved. I've seen these suggestions add time because agents have to unlearn the AI's bad framing before they can give a correct answer.


Your CRM is lying to you.


   
ReplyQuote
(@charlieb)
Eminent Member
Joined: 1 week ago
Posts: 29
 

Good point about historical answers being a poor benchmark. They're often the problem you're trying to solve. But that just shifts the goalpost - if the 'actual answer' should be a cleaned-up ideal, who defines that? The vendor?

You're dead right about edit distance being a blunt instrument. Swapping a greeting is noise; misdiagnosing a core step is a critical failure that no string similarity metric will catch. That's why these features fail on complex tickets - they're optimizing for character overlap, not correctness.

The "time-to-resolution" point is the killer. I've watched an agent spend more time explaining why the AI's suggestion was wrong to a frustrated customer than it would've taken to type a correct reply from a blank slate. The cost isn't just keystrokes, it's credibility.


Trust but verify.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Interesting project. That drop in similarity for complex issues is a clear signal the tool is working as a template generator, not a reasoning assistant.

One nuance I'd add to your findings: sometimes a low edit distance is misleading. If the suggestion says "Please restart the server" and the correct answer is "Please do NOT restart the server," the characters are nearly identical but the meaning is opposite. It can create a dangerous autopilot moment for a tired agent.

Quantifying the gap is the first step. The harder part, as the thread's shown, is figuring out the real cost of that gap in agent trust and customer patience.


Keep it civil, keep it real.


   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Absolutely. That example about restarting the server hits home. It's the difference between a suggestion being *wrong* and being *dangerously almost right*.

That autopilot risk is huge for agents handling high-volume tickets. We saw a similar issue with lead scoring suggestions where the AI would confidently tag a "Request for Quote" as a "Demo Request" because the phrasing was close. The similarity score would be high, but routing it wrong created a massive delay and a frustrated sales prospect.

It feels like these tools optimize for reducing *perceived* effort (look, a pre-written reply!) rather than *actual* cognitive load, which often increases when you have to debug the suggestion itself.


Cheers, Henry


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 7 months ago
Posts: 427
 

Yeah, the historical answers as a benchmark really got me thinking. When I was building alerts for my team's old ticketing system, I saw so many past replies with internal acronyms and "just reboot it" shortcuts. If an AI learned from that, it'd just bake in the bad habits.

That "unlearning the bad framing" point is so real. It's like getting a wrong map - it takes longer to correct your course than if you started with no map at all.

Has anyone tried using a *curated* set of ideal responses as the benchmark instead, even if it's a smaller dataset?



   
ReplyQuote
(@bent36)
Estimable Member
Joined: 3 months ago
Posts: 114
 

Interesting approach. I've tried comparing auto-suggestions against my own past notes, but using historical tickets adds more context.

The drop in similarity on complex issues makes sense. In my own work, templates fall apart when the problem isn't standard. Did you find any pattern in the type of edits needed for those 40-60% cases? Like were they mostly adding specific steps, or changing the whole tone?



   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

The "time-to-resolution vs. keystrokes saved" mismatch is exactly what we saw in our benchmark. We instrumented a basic ticket flow and found the median save was 7 seconds of typing, but the 90th percentile *added* 112 seconds from corrections and credibility repair. That negative tail kills the average benefit.

For the benchmark problem, we used a set of answers manually tagged "golden" by top-tier agents, but as you say, that just moves the goalpost to vendor/manager judgment. It did highlight something else, though - the suggestions often matched the "golden" answer's structure but swapped out the single most critical piece of information, like a wrong error code or KB article link. That's where the cognitive load spikes.


Numbers don't lie


   
ReplyQuote
Page 1 / 2