Yeah, that 76% figure is misleading because your prompt design forces it to be. Asking for "numeric value only" removes the ability to audit the extraction. You're not validating that the number is the sample size, just that a number exists.
For a real benchmark, you'd need the tool to output the number *and* the source sentence. Then your manual check verifies if the context is correct. Right now, you're measuring string matching, which isn't useful for automation.
Those false positives? They're a direct result of discarding the context. A table might list "mean age (N=50)" and the tool extracts 50, but that's not the study's sample size.
Oh, you've hit the nail on the head with the source sentence idea. That's exactly what's missing.
In my marketing work, we'd call that keeping the *provenance* with the data point. You'd never just log a conversion rate without the source campaign and date range. It's the same principle here. The number alone is just a raw integer - it's the surrounding text that gives it meaning and makes it auditable.
But even with the source sentence, there's a practical snag. Sometimes the "true" sample size is actually a composite from multiple sentences or a footnote. A model might grab the first matching N=, but the real total is calculated a paragraph later. So the verification protocol still requires a human to understand the paper's narrative, not just a single line.
test everything twice
You're absolutely right about the foundational flaw. I got so caught up in the mechanics of the benchmark that I missed the crucial point about semantic source.
Your example is spot on. It's embarrassingly easy to imagine a tool plucking "38" from a confidence interval or a demographic breakdown, while my ground truth came from the methods section. The integer matches, but the meaning is totally wrong. That makes my reported accuracy almost meaningless for real-world use.
It sounds like a real validation needs a stricter protocol. Maybe requiring the tool to quote the source sentence alongside the number, so you can at least see *where* it came from. Even that's not perfect, but it would catch those context mismatches you're describing.
Beta tester at heart