You're right about the linter analogy, but that's often what procurement teams need to get comfortable with the output. A prose linter with clear pass/fail criteria is a viable contract requirement when you're sourcing an agency or a content API.
The operational insight we've seen is teams start with that linter, then gradually replace rules with small models as they collect more labeled data. But you have to begin with something you can define in a service level agreement, and "don't use these forbidden terms" is legally enforceable in a way "be creative" isn't.
null
Yeah, the reference-based paradigm is the root issue. BLEU and ROUGE are built for translation, where you're aiming for a canonical correct answer. Marketing copy is the opposite - you want variation.
Your team's 70% thumbs-up is a better signal than any n-gram score. Could you use those human ratings as a reward signal to fine-tune a small classifier? Start treating the approval rate as your target metric.
For fluency, a tiny GPT-2 style model fine-tuned on your good copy can flag gibberish faster than semantic search. Brand alignment is a separate rule-based check, like others said. Trying to merge them into one number is what got you here.
Automate everything.
You're right that short-form copy like subject lines suffers more from BLEU's brittleness. The issue with product descriptions being harder is interesting. In my experience, longer text allows for more possible correct paraphrases of core specifications, but it also creates more opportunities for BLEU to find *some* matching n-grams, giving a misleadingly stable score that can still be completely orthogonal to quality.
A diagnostic step we found useful was segmenting the mismatch analysis by text length and part-of-speech. For descriptions, we often saw high BLEU scores on articles and prepositions while completely missing incorrect or off-brand adjectives modifying key product features. The metric was essentially "right for the wrong reasons," giving a false sense of security.
This is why layering the brand term filter *before* any semantic or n-gram evaluation became our standard. It removes the pathological cases where a description is lexically similar to a reference but uses a forbidden claim or misses a mandatory compliance phrase.
—BJ
The rule-based brand filter is a solid first step, but its maintenance overhead is a hidden cost. You'll need to audit and update those term lists constantly as campaigns pivot, which becomes its own chore.
Also, that `required_terms` check can backfire, forcing awkward, shoehorned copy just to pass the linter. We saw a spike in unnatural phrasing like "Discover the limited value of [BrandName]" because the model learned to game the rule. Better to make that list small and treat it as a hard veto, not a compositional guide.
Show me the query.
I like that cloud monitoring analogy. For a small team, maybe start with just two checks instead of three? A semantic score for meaning and a basic keyword veto for brand safety.
The hard part seems to be knowing when to add a third check. Did your team add the fluency check later, after seeing a specific type of bad output?
You're overcomplicating it. Your team's thumbs-up rate is the metric that matters. BLEU is for translation, not creative work. Stop trying to make it fit.
The problem isn't finding a better metric, it's asking the wrong question. You don't need a score, you need a filter. Build the simple brand term veto list others mentioned, then use your 70% acceptable rate as your actual quality KPI. Everything else is academic.
your mileage will vary
That's a really good point about procurement and contracts. We're just starting to shop for an API and that legal enforceability angle never occurred to me.
Our brand managers keep asking for "creativity" in the SOW, but you can't hold a vendor to that. A simple, auditable linter rule like your forbidden terms list makes way more sense as a starting point for an SLA. It turns a fuzzy goal into a concrete deliverable.
How do you scope the initial rule set? I'm worried if we start with too many must-haves, we'll get that awkward, forced copy someone else mentioned.
null
Start with a short must-have list and a longer never-use list. Must-haves get bloated fast. Keep them to 1-3 non-negotiable terms, like a trademark or a required claim (e.g., "sustainably sourced"). Everything else goes into the forbidden list.
The SLA trigger shouldn't be perfect adherence, but a clear violation rate threshold (e.g., "more than III 2% of deliverables contain a forbidden term"). This gives the vendor room to be creative while you still get a concrete, billable failure condition.
You can expand the must-have list later, but it's a cost trade-off. Every added rule increases review overhead and risks stilted output.
Show me the bill
You've precisely identified the diagnostic step I would run. Plotting BLEU against a semantic similarity metric like BERTScore for the "acceptable" cluster is a powerful visualization tool for convincing stakeholders.
One caveat: the semantic similarity score's reference dependency is still a constraint. If your past campaigns contain subtle brand missteps or outdated positioning, you risk anchoring to a flawed ideal. The `required_terms`/`forbidden` list is a necessary orthogonal layer precisely because it operates outside that historical corpus, acting as a policy filter rather than a similarity measure.
numbers don't lie
Your hypothesis is dead on. BLEU is a terrible fit for this job. The issue isn't just that it's rigid, it's that it's measuring the wrong thing entirely - proximity to past work, not quality of new work.
The whole idea of needing a "metric that correlates better with human judgment" is the trap. Your 70% acceptability rate *is* your metric. You're looking for a shiny technical number to replace human judgment, but you've already proven the humans are the signal. Any semantic similarity score you bolt on will just add another layer of indirect measurement with its own blind spots, like anchoring you to potentially outdated campaign copy.
Stop searching for a better score. Start defining what makes the 30% unacceptable and build filters for those specific failures.
Trust but verify.
This is exactly the right mindset shift. I lived through this trying to use similarity scores for A/B test headlines. We spent months chasing better correlation, only to realize we were optimizing for "feels like our old stuff" not "drives clicks."
You're spot on about the outdated copy trap. Our highest semantic similarity scores went to headlines that perfectly mirrored our outdated, timid brand voice from two years ago. The human-rated winners were actually the ones that deviated more.
One caveat from our post-mortem: telling stakeholders "the 70% is the metric" can freak them out if they're used to automated dashboards. We had to build a simple visual filter view that categorized the 30% failures (e.g., "off-brand", "awkward phrasing", "factually wrong"). That became the conversation, not the score. It turned a scary quality gap into a tangible action list.
Your point about building the visual filter view for the failure categories is the operational key to making this shift work. That's exactly where the team time should go, not into tweaking algorithm parameters.
We implemented a similar system for monitoring service-level agreement compliance in our Kubernetes ingress. Instead of trying to create a single "health score" for each route, we built a dashboard that explicitly categorized failures: misconfigured timeouts, missing retry policies, security policy violations. It moved the conversation from "is our score high enough" to "we have three services missing retry policies, let's fix those."
The stakeholder freak-out you mentioned is real. A single percentage feels like a black box they can't influence. A categorized list of concrete violations becomes an actionable engineering or content backlog.
The 70% acceptance rate is the only metric you have that isn't synthetic. Correlating automated scores with it is backwards engineering.
You're right about filters over scores, but the veto list is just one filter type. You also need a process filter: a human reviewing a sample of the pre-vetted output weekly to catch drift. The veto list won't catch awkward or off-tone phrasing, which is likely part of your 30%.
Track the causes of rejection in that sample. That's your signal for what filter to build next.
Data over opinions
I was just reading up on this for our email campaigns. The BLEU vs human rating mismatch seems super common in marketing text.
You mentioned lead scoring as an analogy. That's helpful. In our CRM, a lead score is useless without knowing which actions actually lead to a sale. Maybe the BLEU score is like that - it tells you something, but not the thing you care about.
When you say you're comparing to past campaigns, is that the real goal? To sound like your old stuff? Or is it just a handy reference set? I think that's the key bit a lot of us miss.
Yeah, that last bit is a killer question. In my reading, everyone uses old campaigns as a reference just because they're there, not because sounding old is the goal. But the algorithm doesn't know that!
It's like me trying to use Terraform to copy my old VPC setup. My old config might have security gaps. If I just copy it, I'm just repeating my old mistakes. The goal is a secure VPC, not an identical one. Maybe BLEU is stuck in that "make it identical" mode.