So we’ve all seen the vendor promises: “Let our magical AI write all your replies!” and it spits out generic nonsense that makes your senior agents weep. My team decided to test a different approach. Instead of feeding the beast our entire knowledge base (full of outdated articles and hopeful marketing copy), we trained the new AI-assist feature **only on resolved tickets from the last six months**.
The hypothesis was simple: let it learn from what actually worked to close a ticket, not from what some product manager *wished* would close a ticket.
After a month of running it side-by-side with our old model, the deflection rate for Tier 1 inquiries jumped by 15%. Not earth-shattering, but not a rounding error either. More importantly, agent feedback shifted from “I have to rewrite the entire thing” to “I can use this as a decent first draft.”
The key takeaways from our little experiment:
* **Quality over quantity:** A smaller, high-quality dataset of *successful* interactions beats a massive dump of every document you own.
* **Context is (still) king:** The AI trained on closed tickets got better at recognizing the *real* problem behind the initial question, because it saw the full resolution thread.
* **It’s still a tool, not a replacement:** Agents now use the draft replies as a starting point, but they’re still applying judgment and adding the human touch.
The vendors want you to believe their out-of-the-box brain is genius-level. Reality check: it’s only as good as the garbage—or gold—you feed it. Maybe stop feeding it the company propaganda and start feeding it your team’s actual work.
Interesting approach, but I'm skeptical about the long-term effect. Training only on closed tickets means it's learning from a sample that's already been filtered through your current process. It'll just get better at replicating your existing biases and blind spots.
What about the tickets that *should* have been resolved differently, or the novel problems that don't look like past ones? You might just be baking in your team's existing shortcuts and oversights. A 15% bump is nice, but is it because the AI is smarter, or because it's just better at mimicking the path of least resistance?
There's a free alternative: have the model also analyze a random sample of tickets that got escalated or took multiple rounds. Let it learn from the failures, not just the successes. Otherwise, you're building a very expensive parrot.
FOSS advocate
That's such a brilliant, practical insight. You've nailed the core problem with most AI implementations: they're trained on aspirational content instead of historical reality.
Your point about **recognizing the real problem behind the initial question** is everything. Our support KB articles are written for the "perfect" customer journey, but real tickets are a messy map of where users actually get lost. Training on successful resolutions teaches the AI the common detours and shortcuts, not just the highway.
The 15% lift is fantastic, but I'm even more excited about the agent sentiment shift you mentioned. When it goes from being a burden to a useful draft, that's where you get real adoption and time savings. Have you noticed if the suggestions are helping junior agents learn the "style" of your best closers? That could be a hidden long-term benefit.
Clean data, happy life.
I see your point about learning from failures, but your free alternative has a hidden cost: training time and data quality. You think analyzing escalated tickets is free? Someone has to tag, clean, and label that data. That's expensive analyst hours.
And if your current process has blind spots, your escalated tickets are likely just the most *visible* failures, not the most instructive ones. The real value in the OP's method is it's pulling from a clear, measurable outcome: ticket closed, customer satisfied. That's a harder signal to game than "this got escalated."
Still, you're right to be skeptical of any single-number result. A 15% lift in deflection could just mean the AI is really good at guiding users to the existing knowledge base articles that already work. That's useful, but it's not intelligence. It's retrieval. Would love to see if that 15% holds when they measure first-contact resolution on more complex, novel issues next quarter.
show me the bill