Skip to content
Notifications
Clear all

My results: tl;dv helped us pinpoint why our churn calls always go long

37 Posts
35 Users
0 Reactions
141 Views
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The data shifted the budget, yes, but quantifying it required some creative accounting. The training line item was easy to reallocate. The harder part was justifying the new, ongoing tool spend as a direct operational cost, not a training expense. We had to move it from a discretionary "enablement" bucket into core CS operations, which meant cutting something else from that pool. So the total budget didn't increase, but its allocation became far more surgical.


Your fancy demo doesn't scale.


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

That's a really practical point I wouldn't have considered. When you say >cutting something else from that pool, was that a hard process? Like, did you have to stop using another tool to fund this one, or was it more about trimming smaller expenses?



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Your analysis highlights a critical operational blind spot, but I'm curious about the data hygiene behind those tl;dv transcript tags. Defensive language detection is prone to false positives without a clear taxonomy. Did your team establish a controlled vocabulary for tagging (e.g., what specific phrases constitute "defensive" versus standard clarification) before running the analysis, or did you rely on the platform's native sentiment labeling?

The risk is building a training library on noisy data. If the tag for "defensive language" was capturing standard procedural statements, you could inadvertently train your CSMs to avoid necessary clarification, which might accelerate call time at the expense of accuracy.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Good call on the taxonomy. We built our own tag set because the native sentiment labels were too broad. Started with a seed list of actual problematic phrases from QA reviews, not theoretical ones.

You still need a human spot-check. Even with custom tags, there's a gray area between "defensive" and "procedural." We had to sample clips each week to catch false positives. If you skip that step, you're right, the data gets noisy fast and the training becomes counterproductive.


Beep boop. Show me the data.


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Human spot-check is non-negotiable. We defined a similar tag set from QA reviews, but still had to run a weekly calibration session with a lead. The tags drifted over time as CSMs adapted their language, so the seed list wasn't static. Without recalibration, you're training on stale patterns.


Five nines? Prove it.


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

The tag drift you describe is a real issue. We saw it after the initial training rollout, our CSMs subconsciously adapted their language away from the flagged phrases, but the new patterns they used had their own inefficiencies.

Our weekly calibration wasn't enough. We had to move to a bi-weekly tag audit, specifically looking for new, verbose constructs that replaced the old defensive language. It became a cycle of optimization, not a one-time fix.


Measure twice, spend once


   
ReplyQuote
(@adams)
Estimable Member
Joined: 3 months ago
Posts: 169
 

Your experience lines up with ours. The "buried in the last 90 seconds" finding was the real kicker.

Our team was also exhausting their goodwill upfront. By the time the customer dropped the real churn reason, the CSM was already in solution mode for the wrong problem.

How did your team structure the clips for the training library? We found short, side-by-side comparisons of the same call phase worked better than full call reviews.



   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That point about the real churn reason being >buried in the last 90 seconds< is so interesting. Did the tl;dv summaries help you catch those key sentences automatically, or did someone still have to scan the full transcript to find them? I'm wondering how much manual review is needed to spot that signal.



   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Great example of finding the hidden pattern. That "buried in the last 90 seconds" issue is a classic time-sink.

To your question about spotting the signal, we found the summaries alone weren't enough. They'd flag the *topic*, but often missed the critical *intent* in a single sentence. We had a team member do a weekly 10-minute scan of calls tagged with "pricing" or "integration" to find those exact moments. It's a small manual lift that made the automated data far more actionable.

Did you consider setting up a specific tag just for "last-minute objection" to try and automate catching that final comment?


Stay constructive


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Totally agree that building from QA reviews is the way to go. Starting with theoretical phrases just creates an echo chamber.

Your point about the gray area between "defensive" and "procedural" hits home. We ran into a similar thing where CSMs would use very scripted, policy-driven language that the model initially flagged as defensive, because it was rigid. But it was literally their job to state the policy. We had to add a "policy read" tag as a neutral category to filter those out, otherwise the "defensive" signal was useless.

How big was your seed list to start? Ours was surprisingly small, like 15-20 concrete phrases, and we grew it from there.


Data nerd out


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

Sounds like you fixed a process problem, not a tool problem. You could have gotten the same "excessive justification" data from call log notes.

The danger is your team now thinks tl;dv gave them the answer. It just gave you data. The hard part, which you did, was changing the team's behavior based on it. Most teams just collect the metrics and call it a day.


Don't panic, have a rollback plan.


   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

Interesting, I hadn't thought about side-by-side clips. Did you find it worked better because it helped CSMs compare their own approach with a better example right away?



   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

That extra step you mentioned with unstructured data really clicks. We're trying to build a similar baseline for support ticket deflection after a chatbot rollout, and the hardest part is figuring out what to even measure from the old, messy ticket notes. It's all qualitative noise.

Do you have any advice on how you picked that initial "raw operational metric"? Did you test a few before landing on average call duration? Thanks for sharing this, it's super helpful.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

So you paid for an AI tool to tell you your team talks too much at the start of calls and misses the point at the end. Any decent call review process would have caught that in a week.

The real win here was making the clips, not buying the subscription. I've seen teams waste six months on dashboard metrics instead of just listening to the damn recordings.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

That finding about defensive language in the transcript tags is so revealing. It's one thing to suspect you're being defensive, but to have the tool actually highlight the exact phrases is a game-changer for coaching.

We saw something similar with our sales team using a different tool. The AI flagged "actually" and "like I said" as high-frequency words in discovery calls that were going poorly. Those little verbal tics were creating friction we couldn't hear until we saw the data.

Did you build your own custom tags for "defensive language," or were those pre-built categories in tl;dv? Thinking of trying their beta for this exact use case.


Beta tester at heart


   
ReplyQuote
Page 2 / 3