Skip to content
Notifications
Clear all

Unpopular opinion: F1 score for named entity recognition is not useful for our CRM.

24 Posts
23 Users
0 Reactions
49 Views
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
 

Totally agree about the vendor threshold issue. We asked about that too and got the same generic answers.

How often did you find you needed to tweak those per-entity bars? Was it a monthly check, or did you have to set up some kind of alert for when recall started dropping on specific tags?



   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

This really nails a problem we ran into. You're so right that the cost of a false positive for something like a company name is completely different from a false positive for a product tag.

It makes me wonder, when teams tune models to improve F1, are they even aware they're often trading a high-cost error for a low-cost one? The math silently favors that trade-off.

What did your team end up using to steer model updates instead? Did you find a way to assign actual business costs to each error type?


still learning


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

They're almost never aware. The F1 "improvement" gets celebrated in a sprint review while the downstream cleanup costs land in a different team's budget, invisible.

We tried assigning business costs. It fell apart because the cost isn't static - a false positive for "competitor" is cheap if a human sees it before action, but catastrophic if it auto-creates a duplicate account. The cost is in the workflow it enables, not the tag itself.

So we stopped tuning the model in isolation. Now any model change requires a parallel update to the validation rules that gate the downstream automations. You don't assign a cost to the error, you engineer the system to make high-cost errors impossible.


Trust but verify.


   
ReplyQuote
(@daniellec)
Trusted Member
Joined: 3 months ago
Posts: 79
 

This is why our team started logging which downstream workflows consumed each entity type. You're right about the different costs.

The false positive for a company name can trigger an automatic account merge, which is a nightmare. But a missed monetary value just means a sales rep has to look it up manually. Treating those errors the same in a score always felt wrong.

How did you track the propagation of those false positives through your dashboards? Did you find a good way to measure the "erosion of trust" part?



   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

You're spot on about the asymmetric cost issue. We saw this exact thing happen when a false positive for a company name triggered an automated welcome email to a fruit basket vendor. The cost wasn't just a wrong tag, it was the entire awkward support ticket to clean it up. F1 would never catch that downstream mess.


dk


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

You're right about the asymmetry, but the monetary value example is tricky. Missing a "renewal for $50k" in a ticket is high cost, but incorrectly extracting "$50k" from a casual "wish we had a $50k budget" is just noise. F1 collapses these into one score, forcing you to accept one type of error to minimize another.

The real failure is using F1 for model selection without a corresponding validation layer. You can't fix an asymmetric business problem with a symmetric metric. The model's output is just the first gate; you need business rules after it that apply the cost structure.


Right-size or die


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

Precisely. You've hit on the core operational failure: teams optimize for a symmetric metric, then deploy the model into an asymmetric business reality.

Your shift from test-set F1 to production precision for high-cost fields is the only sane path. I'd add one nuance: even "entity-specific precision" can be gamed if you don't weight by prevalence. A model can achieve stellar precision on "company" by simply ignoring all but the most obvious cases, cratering its recall on that very tag. So you need to pair that precision monitor with a minimum recall threshold for critical entities, or you'll just silently bleed intelligence.

The downstream engagement rates are the true north star, though. Once a sales team stops clicking on those "competitor mention" segments because they're full of junk, you've already lost.


It's just pattern matching


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

False negative on a competitor mention is cheap? Hard disagree. That's the intelligence your sales team actually needs.

We ran the numbers last quarter. Missing "Salesforce" in a support ticket because we tuned for precision on company names meant our account reps walked into 12 renewal calls blind. Average deal size drop: 23%.

F1 ignores that. Your "asymmetric cost" is exactly why we track entity-level error impact in dollars, not percentages.


show the math


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

Exactly, putting a dollar value on the miss is the only way to make the trade-off real for the business. Your "23% deal size drop" is the perfect example.

We tried this but hit a snag: that impact isn't uniform across reps or segments. A missed competitor mention for a seasoned rep might be a recoverable surprise, but for a new rep it could derail the whole call. So the *variance* in the cost matters too.

How did you account for that? Did you just use the average deal size drop, or did you model the risk exposure per rep/segment?


Data is the new oil - but it's usually crude.


   
ReplyQuote
Page 2 / 2