We built our own tags for defensive language, which I'd recommend. The pre-built ones were too generic. We started by having our QA lead pull 20 calls that had gone poorly and literally highlight every phrase that made her cringe. The list was eye-opening.
Things like "Per my last email" and "As I mentioned" popped up constantly in tense calls, but were almost absent in successful ones. We fed those into a custom tag. The key was defining a clear negative outcome for each phrase, so the tag wasn't just flagging word use, but a specific damaging behavior.
Be prepared for some pushback when you share the data, though. Some folks initially saw it as nitpicking their word choice, until we played the clips side by side so they could hear the tonal difference themselves.
Exactly. The baseline is everything. We tracked four weeks of call duration before rolling out the tool. Without that, the post-implementation drop looks like random noise.
You can't have a "cost offset" argument if you can't measure the waste first.
Finance always assumes correlation isn't causation. Show them the baseline trend, then the break in the trend after the tool goes live. That's the only chart that works.
metrics not myths
This is spot on, especially the point about finance needing to see a clear break in the trend. A baseline stops it from being a faith-based initiative.
One extra piece that worked for us: we also tracked a qualitative baseline. We had managers log their *perceived* reason for long calls before the tool. Often it was a guess ("complex issue"). Having those guesses to contrast against the tool's actual findings added another layer of credibility.
Stay factual, stay helpful.
That qualitative baseline idea is brilliant. It turns subjective manager experience into a comparable data point, which is something finance people actually respect.
I'd caution that it works best when the pre-tool guesses are documented *independently*, maybe in a separate system, before anyone sees the AI analysis. If they're filled out after the fact, the tool's findings can subtly influence memory. We had a team do it retroactively once and the "guesses" miraculously aligned with the report, which weakened the whole contrast.
Stay constructive