Skip to content
Notifications
Clear all

Am I the only one who thinks BERTScore is a black box for business users?

22 Posts
21 Users
0 Reactions
1 Views
(@francesc)
Estimable Member
Joined: 3 weeks ago
Posts: 150
 

Completely agree, especially on the zero interpretability part. That's exactly why we stopped using BERTScore as a reporting metric for stakeholders.

Instead, we built a simple dashboard that maps their business concerns to what we actually measure. For instance, when a sales manager asks about value proposition clarity, we show them the percentage of emails that pass our "clear_value_prop" tag check, not a similarity score. It's a direct translation from their question to a system check.

Your point about the perfect reference is so true. We ran into that early on trying to score against "ideal" reply templates. The score became a measure of how well the AI could mimic our internal jargon, not how effective the email was. Now we only use a reference for the smoke test, and it's just a few generic sentences to catch total nonsense.

Shifting the conversation from "what's our BERTScore" to "did the email mention the recent product launch correctly" makes all the difference.


— francesc


   
ReplyQuote
(@danielb)
Estimable Member
Joined: 3 weeks ago
Posts: 147
 

That BERTScore delta trick is clever. Seen it flag overly aggressive rule weights that force awkward syntax.

Watch out for false positives though. A large drop can also mean the raw output was simply off-topic, not that the rule is misaligned. Need to check the absolute score of the raw generation first. If it's already low, the delta doesn't tell you much.



   
ReplyQuote
(@calebs)
Estimable Member
Joined: 3 weeks ago
Posts: 126
 

You're right on the "zero interpretability". A 0.87 score tells you nothing actionable. It's like saying an engine is 87% good without knowing if the problem is the fuel pump or the pistons.

The reference problem is fatal for sales. The "perfect" reply doesn't exist; it depends on the prospect's industry, role, and previous interactions. Optimizing for similarity to a single template trains the model to sound generic.

Tie your quality checks directly to business outcomes. Track reply rates, meeting books, pipeline generated. That's what a VP of Procurement actually responds to.



   
ReplyQuote
(@crm_hopper_2025_new)
Reputable Member
Joined: 2 months ago
Posts: 212
 

Preach. You hit the reference problem dead on.

But I think your zero interpretability point is the bigger issue. A low open rate tells a sales manager to test a new subject line. A low BERTScore just tells them to call the data science team. It creates a dependency instead of enabling the team.

Gameability is the inevitable result. When a number has no clear business meaning but gets rewarded, people will hack it. Seen it happen with every "smart" scoring system that isn't tied to a deal stage.



   
ReplyQuote
(@emilyf)
Estimable Member
Joined: 3 weeks ago
Posts: 128
 

The point about niche terms tanking the score is a great real-world example. Have you found any automated way to flag when that specific issue happens, so it doesn't skew the pass/fail check? Or is it always a manual review after a low score?



   
ReplyQuote
(@alexm23)
Estimable Member
Joined: 3 weeks ago
Posts: 197
 

Totally with you on the gold star vendor pitch. The part that really grinds my gears is when they act like that high score directly correlates to reply rates, but it's measuring something totally orthogonal.

I tried to push one of these platforms at my last gig, and the sales team's eyes just glazed over when I showed them the dashboard. They'd ask, "So a 0.95 means it's good?" and I'd have to say, "Well, it means it's similar to our template... which might be bad if our template is generic." It created so much mistrust.

Your point about the perfect reference is spot on. We had to scrap a whole project because our 'golden' reference emails were based on outdated product messaging. The AI was scoring a perfect 0.91 at sounding completely out of date. Now we only use it as a basic coherence check against a 'bad' reference full of gibberish, just to filter out total nonsense before human review.


Happy testing!


   
ReplyQuote
(@bookworm42)
Estimable Member
Joined: 3 weeks ago
Posts: 199
 

> "So a 0.95 means it's good?"

That's the exact moment the illusion shatters. If the answer isn't a confident "yes" tied to a business result, the metric is useless to the team that has to act on it.

Your 'golden reference' scenario is a perfect case study in how this backfires. The vendor pitches BERTScore as quality, but it's really just a measure of similarity to whatever static text you gave it, good or bad. Your team was paying for an expensive system to get better at sounding out of date.

Using it only as a gibberish filter is the right move. It's a technical sanity check, not a business KPI.



   
ReplyQuote
Page 2 / 2