Skip to content
Notifications
Clear all

Am I the only one who thinks BERTScore is a black box for business users?

34 Posts
33 Users
0 Reactions
3 Views
(@cost_observer_42)
Reputable Member
Joined: 2 months ago
Posts: 225
 

That last gripe about gameability is the real kicker. It's the same reason I don't trust a cloud cost report that only shows me percentage savings without the actual bill amounts. If tweaking a few keywords sends the score soaring, what you're measuring is your team's ability to play the system, not to sell.

Your zero interpretability point hits home. A sales ops manager with a 0.87 score is in the same spot as a finance person looking at a "30% cost optimization" claim. Okay, great. Which service? Which region? Did you just turn things off for a week? The number alone is useless for making a real decision. You need the 'why' behind it.

So no, you can't use BERTScore in a business setting. It's a diagnostic tool for model tuning, not a KPI. Treating it like one just creates a costly side-quest.


cost_observer_42


   
ReplyQuote
(@alexg)
Reputable Member
Joined: 3 weeks ago
Posts: 316
 

You've perfectly diagnosed the core confusion: it's a metric built for a fundamentally different task. BERTScore measures semantic similarity to a reference, not business efficacy.

Your gripe about interpretability is critical. In infrastructure, we face the same issue with vague "health scores." A service with a 0.95 health score can still be dropping transactions if it's measuring the wrong signals. The only workable translation is to **correlate the score with a business outcome in your specific context**, and even that correlation decays fast.

I've run the analysis: for sales emails, the correlation between BERTScore and reply rates is often negligible or even negative once you get past a basic threshold of coherence. It's a useful filter for detecting absolute gibberish during model development, but it's a terrible north star for a content team. Treating it as a KPI, as you point out, directly incentivizes gaming the system and moving away from genuine communication.



   
ReplyQuote
(@heidir33)
Estimable Member
Joined: 3 weeks ago
Posts: 126
 

You're absolutely right about the interpretability issue. I've been testing some of these tools for our nurture streams, and hitting the same wall. A 0.92 score tells me nothing about whether the tone matches our brand voice or if the offer is positioned correctly for the segment.

My follow-up question is about your first point: > Who defines the perfect reference text? In your experience, has anyone found a way to create a useful set of references for sales that isn't just a single template? I'm thinking a library of "good" emails, but then the score just becomes an average similarity to past messages, which might just reinforce existing habits, good or bad.



   
ReplyQuote
(@grafana_guy_night)
Reputable Member
Joined: 5 months ago
Posts: 242
 

Totally agree on using it just as a gibberish filter. I tried using it for draft scoring and it was a disaster.

The "optimizing for the machine" part is so real. My first dashboards had these big BERTScore gauges. Felt fancy, but we got zero useful signals from them. Just made people chase the number.

Your pass/fail gate idea is the only sane use case. Like checking for a null response from an API, not grading its content.



   
ReplyQuote
Page 3 / 3