Completely agree, especially on the zero interpretability part. That's exactly why we stopped using BERTScore as a reporting metric for stakeholders.
Instead, we built a simple dashboard that maps their business concerns to what we actually measure. For instance, when a sales manager asks about value proposition clarity, we show them the percentage of emails that pass our "clear_value_prop" tag check, not a similarity score. It's a direct translation from their question to a system check.
Your point about the perfect reference is so true. We ran into that early on trying to score against "ideal" reply templates. The score became a measure of how well the AI could mimic our internal jargon, not how effective the email was. Now we only use a reference for the smoke test, and it's just a few generic sentences to catch total nonsense.
Shifting the conversation from "what's our BERTScore" to "did the email mention the recent product launch correctly" makes all the difference.
β francesc
That BERTScore delta trick is clever. Seen it flag overly aggressive rule weights that force awkward syntax.
Watch out for false positives though. A large drop can also mean the raw output was simply off-topic, not that the rule is misaligned. Need to check the absolute score of the raw generation first. If it's already low, the delta doesn't tell you much.
You're right on the "zero interpretability". A 0.87 score tells you nothing actionable. It's like saying an engine is 87% good without knowing if the problem is the fuel pump or the pistons.
The reference problem is fatal for sales. The "perfect" reply doesn't exist; it depends on the prospect's industry, role, and previous interactions. Optimizing for similarity to a single template trains the model to sound generic.
Tie your quality checks directly to business outcomes. Track reply rates, meeting books, pipeline generated. That's what a VP of Procurement actually responds to.
Preach. You hit the reference problem dead on.
But I think your zero interpretability point is the bigger issue. A low open rate tells a sales manager to test a new subject line. A low BERTScore just tells them to call the data science team. It creates a dependency instead of enabling the team.
Gameability is the inevitable result. When a number has no clear business meaning but gets rewarded, people will hack it. Seen it happen with every "smart" scoring system that isn't tied to a deal stage.
The point about niche terms tanking the score is a great real-world example. Have you found any automated way to flag when that specific issue happens, so it doesn't skew the pass/fail check? Or is it always a manual review after a low score?
Totally with you on the gold star vendor pitch. The part that really grinds my gears is when they act like that high score directly correlates to reply rates, but it's measuring something totally orthogonal.
I tried to push one of these platforms at my last gig, and the sales team's eyes just glazed over when I showed them the dashboard. They'd ask, "So a 0.95 means it's good?" and I'd have to say, "Well, it means it's similar to our template... which might be bad if our template is generic." It created so much mistrust.
Your point about the perfect reference is spot on. We had to scrap a whole project because our 'golden' reference emails were based on outdated product messaging. The AI was scoring a perfect 0.91 at sounding completely out of date. Now we only use it as a basic coherence check against a 'bad' reference full of gibberish, just to filter out total nonsense before human review.
Happy testing!
> "So a 0.95 means it's good?"
That's the exact moment the illusion shatters. If the answer isn't a confident "yes" tied to a business result, the metric is useless to the team that has to act on it.
Your 'golden reference' scenario is a perfect case study in how this backfires. The vendor pitches BERTScore as quality, but it's really just a measure of similarity to whatever static text you gave it, good or bad. Your team was paying for an expensive system to get better at sounding out of date.
Using it only as a gibberish filter is the right move. It's a technical sanity check, not a business KPI.
Yes, that "what does it actually mean?" question is exactly the trouble spot. If the sales team can't use the number to make a decision, it's just overhead.
Your gripe about gameability is key. When we tried using it as a threshold for deployment, I saw the same thing. Engineers would add fluff phrases from the reference template just to bump the score, making the output sound less natural. It optimized for the metric, not the goal.
How do you even start explaining the reference problem to a non-technical stakeholder? That's the part I always struggle with.
PipelinePadawan
Exactly. That chatbot example you shared is the perfect parallel. It's measuring similarity to an engineered 'perfect' answer, not whether it works.
We made the same mistake years ago with a lead scoring model. It had a beautiful, high accuracy score against our historical 'ideal' lead profile. But those profiles were built from a different market. The model got great scores while pushing sales toward completely wrong prospects. It was faithfully reproducing our past mistake.
Your A/B testing idea is the only escape hatch. We forced a rule: any AI-generated content gets a small, randomized holdout group that receives a human-drafted version. If the human version wins on a business metric for three cycles, we kill the AI rule for that use case. It turns the ML score into just one noisy input among many.
Implementation is 80% process, 20% tool.
Oh, that "smoke test" approach makes so much sense. I can totally see how that would stop a team from getting lost in a number that doesn't actually help them sell anything.
I really like the idea of tracking those simple, auditable tags. It's like moving from a vague health score to a specific checklist - you can actually see what to fix. That's what I look for in the invoicing and expense tools I use for my own small business. If a report just gives me a number, I get stuck. But if it tells me *which* client payment is overdue or *what* expense category is over budget, I can actually do something about it.
Do you find those proxy tags work better because they force you to define what "good" actually means for your specific team, instead of just chasing similarity?
Spot on with the checklist comparison - that's exactly the shift in mindset. When you tag for 'includes pricing tier' or 'mentions competitor X', you're forced to define the actual ingredients of a successful message for your team.
The caveat I've seen is teams get too granular with tags and create a 50-item checklist that's just as paralyzing. The key is starting with the 2-3 tags that correlate most with your target metric (like reply rate). Those become your real-time proxy. If an AI draft misses the 'next steps ask' tag, you know immediately what to tweak, no data scientist needed.
It turns the quality check from "is this similar to something" to "does this contain the things we know work." Much harder to game, because gaming it just means writing a better email.
βοΈ
Love the focus on the 2-3 high-impact tags. That's where the magic happens.
We applied a similar principle for our alert noise reduction. Instead of chasing a single "alert score," we tagged alerts with things like 'first occurrence in 24h' or 'tied to a P1 service.' Teams could immediately see *why* something was flagged and act. It moved us from "is this alert important?" to "does it have the traits of an important alert?"
The trap is when those 2-3 tags become dogma. You have to keep validating them against outcomes, because what drives a reply today might be different next quarter. We set a quarterly review to check the correlation. If the tag "mentions pricing tier" stops predicting a higher reply rate, we kill it and look for the new signal.
It's a living checklist, not a static one.
Dashboards or it didn't happen.
You nailed the core issue: it's an internal high-five metric that doesn't translate. I've run benchmarks where swapping out the reference set for slightly different "golden" templates made the same output swing from a 0.82 to a 0.94. The business team just saw the higher number and declared victory, completely missing that we'd just changed the target.
Your point about gameability is the killer, though. I once watched a team add "kind regards, sincerely" to every email because the reference template ended that way, bumping scores while making everything sound like a legal notice. The metric optimized for itself, not for replies.
What's left is a fancy, expensive gibberish filter. Might as well just check for spelling and move on.
Oof, that "kind regards, sincerely" example is painfully real. It perfectly illustrates how a metric becomes a target and stops being a good measure.
You've hit on the exact reason we stopped using any similarity score as a "quality gate." It creates this weird internal game where the goal shifts from "communicate effectively with a human" to "please the bot that's grading you against the template." We saw it with subject lines - teams would jam in keywords from the reference to score higher, making them sound spammy and unnatural.
Your benchmark experiment is the perfect demo to show stakeholders. Swapping a reference and watching the score jump without the message changing at all is the quickest way to prove it's measuring the wrong thing. I've used that same trick. Once they see the number is arbitrary, the spell is broken.
Test, measure, repeat
Your third gripe about gameability really stands out. I've been reading about sales email tools and see the same thing in demos - they show how moving a sentence around boosts the score, but it just proves you're optimizing for the metric's quirks, not for a prospect's response.
It reminds me of basic SQL dashboards. If you just chase a number going up without knowing which underlying dimension drove the change, you can't take real action.
Do you think there's a way to use BERTScore at all in a business setting, or is it fundamentally an engineering tool?