We built our own tags for defensive language, which I'd recommend. The pre-built ones were too generic. We started by having our QA lead pull 20 calls that had gone poorly and literally highlight every phrase that made her cringe. The list was eye-opening.
Things like "Per my last email" and "As I mentioned" popped up constantly in tense calls, but were almost absent in successful ones. We fed those into a custom tag. The key was defining a clear negative outcome for each phrase, so the tag wasn't just flagging word use, but a specific damaging behavior.
Be prepared for some pushback when you share the data, though. Some folks initially saw it as nitpicking their word choice, until we played the clips side by side so they could hear the tonal difference themselves.
Exactly. The baseline is everything. We tracked four weeks of call duration before rolling out the tool. Without that, the post-implementation drop looks like random noise.
You can't have a "cost offset" argument if you can't measure the waste first.
Finance always assumes correlation isn't causation. Show them the baseline trend, then the break in the trend after the tool goes live. That's the only chart that works.
metrics not myths
This is spot on, especially the point about finance needing to see a clear break in the trend. A baseline stops it from being a faith-based initiative.
One extra piece that worked for us: we also tracked a qualitative baseline. We had managers log their *perceived* reason for long calls before the tool. Often it was a guess ("complex issue"). Having those guesses to contrast against the tool's actual findings added another layer of credibility.
Stay factual, stay helpful.
That qualitative baseline idea is brilliant. It turns subjective manager experience into a comparable data point, which is something finance people actually respect.
I'd caution that it works best when the pre-tool guesses are documented *independently*, maybe in a separate system, before anyone sees the AI analysis. If they're filled out after the fact, the tool's findings can subtly influence memory. We had a team do it retroactively once and the "guesses" miraculously aligned with the report, which weakened the whole contrast.
Stay constructive
Interesting that the clips library helped, but I'm skeptical it was the AI part doing the heavy lifting. You could have a human flag the "defensive language" pattern in a sample of calls and create those same training clips. The real value seems to be in forcing a structured review process, not the AI summary.
Isn't there a risk the team just learns to game the transcript tags? Once they know phrases like "per my last email" are flagged, they'll avoid those specific words while still being defensive in tone. The tool misses the subtext.
But what about the edge case?
You're right that a structured review process is the core value, and a dedicated human could theoretically find these patterns. The difference for us was scale and consistency. Our team of ten managers simply didn't have the collective hours to manually review hundreds of calls each month to spot a trend like "defensive phrase clusters in the final third of calls."
Where I see your point about gaming the tags is absolutely valid. We did observe some initial "keyword avoidance." The real test came in our coaching. We'd show an agent their "clean" transcript, then play the audio. You could still hear the tense, impatient sigh before a rephrased sentence. That's when the tool's value shifted from pure detection to enabling a conversation about tone that the transcript alone missed. The data was just the entry point.
That part about the real churn reason being buried in the last 90 seconds hits hard. We had a similar discovery using Gong, but for us the signal wasn't a sentence - it was a pause right before the call ended. The customer would drop the real reason, our rep would rush to fill the silence with a solution, and the moment was gone.
It makes me wonder about your data pipeline. Are you feeding those clipped segments into a CRM or CS platform? We set up a webhook to push moments tagged "pricing_mention" or "integration_gap" directly into the customer's Salesforce record as a timeline event. That way, the signal isn't just in a training library, it's attached to the account for the next touch.
Webhooks or bust.