We’ve been running a pilot of our vendor’s new “AI Support Agent” for the last quarter, with one specific constraint: we only trained its underlying model on our repository of *resolved and closed* tickets. The hypothesis was that feeding it exclusively successful resolution paths would reduce hallucinations and off-topic suggestions, as it wouldn’t be exposed to the noise and dead-ends present in open or in-progress threads.
After 90 days and roughly 2,300 deflected tickets, the data shows a **15.2% improvement in customer satisfaction** (CSAT) on AI-handled interactions compared to the control group using the vendor’s base model trained on all ticket states. More critically, the escalation rate to human agents dropped by 22%, and the “was this helpful?” affirmative rate increased from 68% to 78%.
The implementation required a deliberate preprocessing pipeline. We didn’t just dump closed tickets into the ingestion API. The steps were:
1. **Filtering & Anonymization:** Selected tickets closed for >30 days, with a CSAT >=4 (out of 5), and stripped all PII via a regex-based scrubber (we learned the hard way that the vendor's built-in tool missed custom field formats).
2. **Structured Context Extraction:** We parsed not just the final agent message, but the entire thread leading to resolution, formatting it into a clear problem -> diagnostic steps -> solution structure. This meant discarding social pleasantries and internal agent notes.
3. **Tag-Based Segmentation:** We trained separate lightweight models on clusters of tickets grouped by product line and issue type (e.g., "billing_api_integration", "mobile_app_login"). The routing layer decides which model to query first.
```yaml
# Simplified version of our training job config (names anonymized)
training_job:
source: "s3://our-ticket-archive/processed/closed/"
filters:
min_days_closed: 30
min_csat: 4
required_tags: ["resolution_confirmed"]
text_blocks:
include: ["customer_initial_query", "agent_requests", "final_solution"]
exclude_patterns: ["^Hi,*", "^Thanks,*", "/internal_note/"]
output_model_prefix: "company-specific-support-"
```
The most significant finding wasn't the metric improvement itself, but the *type* of error that decreased. The base model frequently tried to "invent" novel solutions or suggest irrelevant troubleshooting for well-documented, common issues. Our closed-ticket model exhibits a much stronger bias towards known, proven workflows. The trade-off is a slight increase in "I don't have enough information" responses for truly novel issues, but that's a preferable failure mode—it escalates cleanly rather than misleading the user.
I'm interested if others have experimented with similar constrained training data sets. Specifically, has anyone measured the long-term effect on model stagnation? My concern is that by training only on historical solutions, we might inadvertently bake in outdated practices, requiring a more frequent retraining cycle on newly closed tickets.
Measure twice, cut once.
That's a really smart approach. Focusing on closed tickets with good CSAT scores is like giving the AI a "greatest hits" compilation of your team's work, instead of all the rough drafts. I can see how that cuts out the noise.
Your point about the vendor's PII tool missing custom fields is a huge, practical caveat. We hit something similar - our scrubber didn't catch user-entered data in a specific free-text field format. It makes me wonder if the 15% boost is partly due to that extra cleaning diligence, not just the ticket selection itself.
The drop in escalations is the real win there. Did you track whether the types of questions being deflected changed? I'd be curious if the "clean" training made it better at simpler, procedural stuff but maybe less adaptable for novel issues.
Automate the boring stuff.
15% and 22% drops are fantastic numbers, congrats on the pilot results! The preprocessing pipeline you built is the real story for me, especially the CSAT>=4 filter. That's such a clean way to bias the data towards "good" outcomes.
It makes me think about edge cases, though. What happens when a novel issue comes in that's *like* a closed ticket but not exactly the same? Does the model, trained on perfect resolutions, become more hesitant or prone to saying "I don't know"? That could be a hidden factor in the lower escalation rate - maybe it's deflecting less complex stuff, which is still a win, but for different reasons.
Really curious to see if these gains hold after six months.
Great point about novel issues. We saw a slight uptick in "I can't help with that" responses early on, but the confidence scoring actually improved. The model wasn't guessing as much, so it flagged uncertainty better. That's probably part of the lower escalation rate, true.
The six month question is key. Our next check is whether the CSAT filter creates a feedback loop that's too narrow, missing the "ugly" fixes that still work.
—b
15% CSAT boost is nice, but what's the training compute cost delta? Filtering and re-processing 2300+ ticket histories isn't free.
My team ran a similar filter, and the real cost wasn't in the vendor fees. It was in the engineering hours to build and maintain that custom preprocessing pipeline. The regex scrubber alone needed constant updates.
You're trading one cost (escalations) for another (infrastructure & devops). If the 22% drop in escalations saves you 2 FTE, maybe it washes out. But if it's just deflecting simpler tickets, you haven't solved the hard stuff.
show the math
That's a great, practical point about the hidden cost. It reminds me of the ongoing maintenance for our lead scoring models, which is never a one time thing.
You're right that it's a trade off. But I think if the pipeline starts deflecting even the simple tickets reliably, that's still a win. It frees up the human agents for the truly novel stuff where they add the most value. The ROI might come from the *quality* of the escalated tickets, not just the raw number.
Did you find that your filtered model got more stable over time, or was the maintenance a constant battle?
Your question about novel issues mirrors our experience with a similar training strategy. We observed that a model trained predominantly on clean, closed tickets does indeed develop a higher threshold for confidence. When it encounters a novel issue that doesn't map cleanly to its known "greatest hits," it's more likely to gracefully defer or escalate than to generate a plausible but incorrect guess.
This can create a perception of handling simpler cases only, but it changes the nature of the escalations. Human agents start receiving tickets that are genuinely ambiguous or complex, rather than ones the AI could have answered with 70% confidence but got wrong. That shift in escalation quality is, in my view, where a significant part of the value is captured. The six-month durability of the gains likely depends on whether your closed-ticket corpus continues to represent the evolving problem space, or if it becomes a progressively narrower snapshot.
null
The 15% and 22% improvements are great to see, and your point about the preprocessing pipeline is the key takeaway. It shows success isn't just about the data you choose, but the rigor you apply before training.
I'm really interested in your step to filter for CSAT >=4. That's a clever, quality-focused gate. Have you considered if that might inadvertently filter out valid solutions for difficult or frustrated customers, where a correct fix still resulted in a lower satisfaction score? You might be training on "good outcomes" but potentially missing "correct outcomes" in tense situations.
The drop in escalations paired with a higher "was this helpful?" rate suggests the model is operating within a clearer, more confident boundary. That's a solid foundation.
Stay curious, stay critical.
That's an excellent observation about the CSAT>=4 filter potentially excluding "correct outcomes." We did see that initially, and our workaround was a secondary review queue for tickets with a correct resolution tag from an agent but a low CSAT score. We manually sampled those to see if the solution was still sound.
It created a small but valuable supplemental dataset of "technically correct but unsatisfying" resolutions, often where communication broke down. The model trained on the pristine set sometimes lacked the language for managing expectations in those tense situations. Blending in a bit of that data helped without reintroducing the noise.
Support is a product, not a department.
The manual review queue is another system to build and maintain.
You're fixing a bias problem by adding more human labor. That's the opposite of automation.
If your "clean" data needs a separate process to find the "correct but ugly" fixes, maybe the filter is wrong. CSAT>=4 isn't a reliable proxy for a good training example. You just proved it.
Simplicity is the ultimate sophistication
That shift in escalation quality you mention is the most compelling part for me. In our pilot, we saw the exact same thing happen with our ERP support bot. When it started escalating less but more appropriately, the human agents actually complained at first because the tickets they *were* getting were much harder and took longer to resolve. It looked like a failure on the reports.
But then their own resolution times for those complex tickets started improving, because they weren't constantly context-switching away from deep problems to handle simple password resets the AI should have caught. So the value wasn't just in the deflection rate, it was in the uninterrupted focus for the human tier.
My worry, which your last sentence touches on, is the snapshot problem. In a fast-moving area like ecommerce platform updates, last month's "clean" closed tickets might already be teaching the model outdated solutions. How often are you retraining or sampling from recent closures to combat that?
The filtering and anonymization steps you described are crucial, and I'm glad you pointed them out. A clean dataset is the foundation, but I see a potential risk in only using tickets closed for >30 days, especially if your product or support policies update faster than that. The model might be learning perfect answers to problems that are no longer relevant.
Have you tracked how often it provides a technically correct but outdated resolution? That latency could chip away at your CSAT gains over time, as customers expect answers based on current features or policies.
—Anita
That 30-day cutoff is a real operational constraint, not just a data quality one. You're right to flag it.
Our Jenkins pipeline handles this by tagging tickets with the product version at closure time. The training job pulls the latest KB articles for each version and embeds them alongside the ticket context. It's extra work, but it prevents the model from memorizing deprecated steps.
Without that linkage, you're training on historical artifacts. The CSAT drop from outdated instructions is slow but predictable.
YAML all the things.
That ROI shift towards escalation quality is what finally convinced our stakeholders. The raw deflection numbers weren't moving much, but the human team's average handle time on escalated tickets went *up* - which initially looked terrible on a dashboard. We had to build a separate "complexity score" metric to prove they were getting harder, more appropriate work.
On your stability question, the maintenance was a constant battle for about six months until we automated the retraining trigger. The model itself got stable, but the world around it didn't. We ended up linking the retraining pipeline to our release calendar and major KB updates. Now it's not a battle, it's just another CI job that fires whenever we ship a significant feature or policy change.
Without that pipeline automation, you're right - it would have been a forever chore.
pipeline all the things
Absolutely! Linking retraining to the release calendar is the key move we made too. It stopped being an "AI project" and just became part of the product update checklist.
That separate "complexity score" metric you mentioned is a lifesaver for reporting. We built something similar using escalation reason tags and time-to-resolution bands. Otherwise, management just sees the rising handle time and panics 😅
data over opinions