That drop to 40-60% similarity on complex tickets is the key finding. It's where the promise of time-saving evaporates.
In our helpdesk, those are exactly the tickets where a wrong or generic first reply escalates the customer's frustration. The agent doesn't just edit text; they have to rebuild the entire conversation's tone from a bad starting point. It adds cognitive load, not reduces it.
I'm really curious about your edit distance findings for those cases. Were most edits small tweaks, or full rewrites? That'd show if it's a template with gaps, or a complete misdiagnosis.
Automate the boring stuff.
You've nailed the core dilemma with "who defines the ideal." In practice, it's often the vendor's product team, which decouples the benchmark from the actual workflow and customer context. This is why these features feel so misaligned - they're optimizing for an abstracted, sanitized version of support.
Your credibility cost observation is critical and often omitted from ROI calculations. When an agent has to say "ignore my first sentence," it undermines the entire interaction. The metric should be "time to *correct* resolution," not just initial suggestion speed.
On edit distance, we found it's worse than a blunt instrument - it's a misleading one. A high similarity score on a complex ticket often indicates the AI suggested the *obvious, generic* steps, which the agent then had to completely replace with the actual, nuanced diagnosis. The character overlap was in the boilerplate, not the substance.
That decoupling from workflow is something I've been struggling to define in our vendor evaluations. We get demos showing replies to perfectly scoped, textbook tickets, but our real tickets are messy. Half of them start with "it's doing the thing again" and reference a conversation from three weeks ago.
> The character overlap was in the boilerplate, not the substance.
This is exactly the kind of metric trap I worry about. A vendor could boast a 90% similarity score improvement, but if all the improvement is in the "Hi [Customer Name]," and "Best regards," sections, it's meaningless. It might even be negative value if it makes agents complacent about skimming the substantive part.
Has anyone seen a vendor successfully measure the "substance overlap" separately from the social framing? Or is that just too subjective to put a number on?
Interesting you used historical tickets as the benchmark. That's like judging a new chef by how well they replicate the lukewarm catering from your last company picnic.
Those old tickets are probably full of the same "just reboot it" shortcuts and internal jargon you're trying to move away from. You might just be measuring how well the AI learned your bad habits. What you really need is a "good answer" benchmark, but then who defines that? Probably the vendor's PM who's never handled an angry ticket.
—aB