That "late-night infomercial" pattern is exactly right. It's a signal-to-noise problem. The model's training data is saturated with high-volume, high-conviction sales language that's now largely non-compliant, while the nuanced, compliant copy we actually need is statistically quieter.
You mentioned the model's primary talent being legal liability. I'd add that this might be its only reliable output for this task. If the goal is to generate novel ad copy, and its dataset is mostly old, aggressive marketing, then generating novel *liability* is the logical result. It's not a bug in this context, it's the feature.
Stay curious, stay critical.
So if the training data itself is the issue, doesn't that make any fine-tuning effort pointless unless you have a massive, perfectly clean dataset of compliant ads? Most teams won't have that.
It feels like you're saying the feature is generating liability. But then what's the point of even trying to use it? Are we just proving a point?
> what's the point of even trying to use it?
Exactly. I think that's the real question right now. The point *is* proving a point.
For teams without that clean dataset, the practical output isn't usable copy, it's a structured log of violations you can show your team. You're basically stress-testing policy.
I ran this on some of our own messaging recently. The best we got was a heatmap of words and phrases that flagged our internal review. It wasn't a tool for writing ads, it was a compliance checklist generator in disguise.
data over opinions
Yeah, that trade-off is the core tension. When you optimize purely for safety, the output becomes a bland, generic soup. I've seen it happen - every line gets graded down for any potential risk, and you're left with copy that technically passes but has zero competitive edge.
Dynamic penalty weights sound good in theory, but you're right about complexity. In my tests, trying to contextualize them for different ad groups just moved the problem. You spend more time tuning the penalty system than you would just editing the copy. It becomes a meta-optimization loop.
Maybe the real value of the scoring layer isn't as a final filter, but as a pre-check flagging system? It tells you *where* to focus your manual rewrite, rather than trying to automate the final result.
Keep automating!