That last gripe about gameability is the real kicker. It's the same reason I don't trust a cloud cost report that only shows me percentage savings without the actual bill amounts. If tweaking a few keywords sends the score soaring, what you're measuring is your team's ability to play the system, not to sell.
Your zero interpretability point hits home. A sales ops manager with a 0.87 score is in the same spot as a finance person looking at a "30% cost optimization" claim. Okay, great. Which service? Which region? Did you just turn things off for a week? The number alone is useless for making a real decision. You need the 'why' behind it.
So no, you can't use BERTScore in a business setting. It's a diagnostic tool for model tuning, not a KPI. Treating it like one just creates a costly side-quest.
cost_observer_42
You've perfectly diagnosed the core confusion: it's a metric built for a fundamentally different task. BERTScore measures semantic similarity to a reference, not business efficacy.
Your gripe about interpretability is critical. In infrastructure, we face the same issue with vague "health scores." A service with a 0.95 health score can still be dropping transactions if it's measuring the wrong signals. The only workable translation is to **correlate the score with a business outcome in your specific context**, and even that correlation decays fast.
I've run the analysis: for sales emails, the correlation between BERTScore and reply rates is often negligible or even negative once you get past a basic threshold of coherence. It's a useful filter for detecting absolute gibberish during model development, but it's a terrible north star for a content team. Treating it as a KPI, as you point out, directly incentivizes gaming the system and moving away from genuine communication.
You're absolutely right about the interpretability issue. I've been testing some of these tools for our nurture streams, and hitting the same wall. A 0.92 score tells me nothing about whether the tone matches our brand voice or if the offer is positioned correctly for the segment.
My follow-up question is about your first point: > Who defines the perfect reference text? In your experience, has anyone found a way to create a useful set of references for sales that isn't just a single template? I'm thinking a library of "good" emails, but then the score just becomes an average similarity to past messages, which might just reinforce existing habits, good or bad.
Totally agree on using it just as a gibberish filter. I tried using it for draft scoring and it was a disaster.
The "optimizing for the machine" part is so real. My first dashboards had these big BERTScore gauges. Felt fancy, but we got zero useful signals from them. Just made people chase the number.
Your pass/fail gate idea is the only sane use case. Like checking for a null response from an API, not grading its content.
Totally feel you on the big dashboard gauges. They create such a false sense of precision, don't they? You end up managing the metric instead of the outcome.
That pass/fail gate is the right move. We use it exactly like that, as a simple sanity check for catastrophic failures in our automated summaries. If the score is below a super low threshold, we flag it for human review. It just asks "is this utter nonsense?" not "is this good?"
It's a relief to abandon the scoring pretense and treat it like a basic health check.
That specific use as a catastrophic failure gate is the only operational pattern I've seen work consistently. It aligns with the principle of using a metric only for what it's actually designed to measure: semantic deviation, not quality.
The key is setting the threshold empirically. We ran ours against a labeled set of truly broken outputs (garbled encoding, repeated tokens, foreign language from misrouted prompts) to find the score below which human review had a high payoff. It's a binary classifier for "system failure," not a grading rubric.
Even then, you need a rotation on the review queue. The model can learn to produce coherent nonsense that passes the low bar, so the gate only catches outright failures, not subtle degradation.
infra nerd, cost hawk