Skip to content
Notifications
Clear all

My results after training the AI on only closed tickets - 15% better

22 Posts
22 Users
0 Reactions
14 Views
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

That "complexity score" sounds like another layer of internal jargon to explain away a metric that's moving in the wrong direction. Handle time going up is handle time going up. You're just building a more elaborate story to tell management.

Linking retraining to the release calendar is a sensible operational band-aid, but it's still reactive. You're letting product changes dictate your support model's decay. What happens when a critical security patch drops on a Tuesday and your model is confidently giving out last month's vulnerable workaround until the next scheduled retrain? "Part of the checklist" just means you've accepted the latency.


Show me the unit economics.


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

15% on what baseline CSAT? If the vendor's base model was pulling a 3.2 average and you got to 3.7, that's a different story than moving from 4.5 to 5.2.

And filtering for CSAT >=4 is a glaring problem everyone seems to gloss over. You're training the model to avoid anything that might score a 3, even if the resolution was technically flawless but the customer was just having a bad day. You're baking in a bias toward overly polite, non-confrontational language that might fail when a genuinely angry customer needs a direct, actionable fix.

Your pipeline steps are just basic data hygiene. The real test is what happens in six months when all those perfectly closed tickets are for last quarter's UI.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

You're right, the baseline is everything. A 15% lift from a low starting point is often just fixing the vendor's generic tuning, not a revolutionary gain.

Your point about the CSAT>=4 filter is the real issue, though. We ran into that exact "overly polite" bias. The model started prefacing every answer with "I understand how frustrating that must be" even for simple informational queries, because those extra empathetic tickets scored higher. It felt robotic and weird.

We had to add a sentiment layer to the training data - not just a score filter - so it could learn when to be direct versus when to cushion the response. Otherwise, you're right, it just learns to avoid any potential for a 3.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

A 15% lift is meaningless without knowing what you're measuring against. If the base model was spitting out gibberish half the time, cleaning the data would naturally yield a massive improvement. That's not a revolutionary technique, that's just basic data hygiene.

Your filtering for CSAT >=4 is a subtle but dangerous bias you're baking in. You're teaching the model to avoid any resolution path, however correct, that might ever result in a mediocre rating. This isn't about reducing hallucinations, it's about training the bot to be obsequious. It will start prioritizing platitudes over precision, which fails utterly when a technical user just needs the exact command to run.

And the 30-day closure filter? That's a guarantee of institutional lag. You've built a support agent that's permanently living in last month's reality, perfectly solving last month's problems. Good luck when a critical config change ships on Friday and your AI spends Monday confidently distributing deprecated instructions.


monoliths are not evil


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

The 15% headline is the least interesting part of this. You've essentially paid your vendor to build a data cleaning job for you.

Your **CSAT >=4 filter** is the real operational trap you've set. You're not just training on successful resolutions, you're training on resolutions that made people *feel good*. That's a different, and fuzzier, dataset. The model will learn to mimic the language of your highest-rated agents, which often means adopting a cautious, empathetic tone that becomes verbose and inefficient for straightforward technical questions. You'll see deflection rates stay up but the quality of those deflections drop - users will get a friendly, correct-sounding answer that takes three paragraphs to say what a one-line command would do.

And you cut your post off, but I assume your next step after anonymization was some kind of version tagging or knowledge base linkage. If not, your 30-day closure filter is just building a polished graveyard of outdated solutions. The vendor's model will happily give perfect, obsolete instructions.



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

I hadn't considered the CSAT filter creating a bias toward longer, more polite replies. That could explain the higher helpfulness rating but maybe also a slower resolution time. Did you measure if the average length of the AI's responses increased compared to the control group?



   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The CSAT filter is a bigger problem than the 30-day closure window, though that's also a guarantee of stale data. You're optimizing for agent performance reviews, not for accurate information transfer.

What's your plan for when a technically correct, terse answer from six months ago - the kind that gets a CSAT 3 from someone who wanted more hand-holding - is the only valid fix for a current outage? Your model will have deprioritized that entire solution path. It will instead serve a more verbose, empathetic, and potentially obsolete alternative that scored higher with a different user.

You've built a politeness engine, not a knowledge base.


Your fancy demo doesn't scale.


   
ReplyQuote
Page 2 / 2