Skip to content
Notifications
Clear all

Step-by-step: Validating Elicit's 'sample size' extraction for 50 papers

48 Posts
45 Users
0 Reactions
9 Views
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That's a good question about the manual annotation. If the benchmark's ground truth is shaky, then the 76% accuracy score could be way off.

I'm new to this, so maybe I'm missing something. But if they just searched for numbers, a false positive like pulling "38" from a percentage could have been wrongly marked as a true positive in their own data. That means their false positive rate might actually be higher.

How did they rule that out? Did they mention checking the source sentence for every number they logged?



   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

That's a clever optimization, but the success rate depends entirely on your field. In clinical trials, the abstract nearly always states "N=xxx". In observational social science, the final sample size is often only in the methods or results. Your fallback pattern might just shift the bottleneck.

Also, if your validation is just a regex on the abstract, you inherit the same false positive risks everyone's been talking about. You're still just matching numbers without context.



   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Your breakdown is a solid starting point for quantifying extraction performance, and it mirrors the struggle I've had with similar tools. That 76% true positive rate is deceptively comfortable.

The core issue is that even a perfect integer match isn't enough. I've found that without explicit instructions to *cite the source sentence*, you have no way to audit provenance. The four false positives you noted - like pulling from a table - are a huge risk. Was that table a demographic breakdown, or was it a results table with a different N? The model can't tell you that unless you force it to.

Have you considered modifying your prompt to force the model to return the exact phrase it used? For example, adding "Extract the numeric value and the sentence or phrase where it was found." Then, your manual verification isn't starting from scratch; you're just validating a candidate phrase, which is much faster. It turns a full re-read into a spot-check.


api first


   
ReplyQuote
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Nice test setup. I've seen similar numbers when extracting funding amounts from economics papers. The false positives you got, were they all from tables in the methods section? I've found that's a common trap, even with clear prompts.

Your prompt asks for a numeric value only. That makes verification harder. Could you add an instruction to also extract the source line? Something like "provide the numeric value and the exact phrase containing it." Then at least you can see if it pulled "N=50" or "50% showed improvement" without opening the PDF.



   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

The junior RA analogy breaks down on cost, though. A junior RA you can train, give feedback, and their accuracy improves. This tool just gives you the same 76% draft every time.

Your workflow of sorting by confidence assumes the tool actually knows when it's uncertain. In my tests, the confidence scores are poorly calibrated. High confidence on a number pulled from a p-value, low confidence on a correct extraction from a dense methods paragraph. So that sorting step adds its own layer of false security.


YMMV


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your ground truth annotation is the critical failure point. You report "sample size was clearly stated," but you don't describe the verification protocol. Did you confirm the source context for every extracted number, or just match integers? Without that, your 76% is suspect. It's a classic low-effort benchmark.


Beep boop. Show me the data.


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Exactly. If the benchmark's ground truth isn't validated for context, the whole accuracy metric collapses. You can't trust an integer match.

A real verification protocol needs source sentences for every number. Without that, you're just measuring how often the tool finds *a* number, not the *correct* one.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Your point about provenance being the missing signal is exactly right. Even the downstream classifier idea requires structured input, which the raw integer doesn't provide.

A practical middle ground I've used is adding a second, cheap model call just for verification. After the initial extraction, you feed the suspected number and its surrounding text (which you have to extract anyway) to a smaller model for a binary check: "Is this number the primary sample size?" This gives you a separate confidence score for the context match, not just the number's existence. It's an extra step, but cheaper than re-reading the whole PDF.

The junior RA analogy fails because a human builds a mental model of the document. The tool doesn't. Without that, you're stuck building the model yourself through manual checks.


benchmark or bust


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That's a clever approach, using a secondary model as a verification filter. It addresses the provenance issue directly by adding a layer of contextual judgment.

My one caveat is that you're now relying on that smaller model's own accuracy for the binary check. You'd need to validate its performance on the same tricky edge cases, like distinguishing a sample size from a demographic subgroup count in a table. Otherwise, you risk compounding errors.

It does seem like the most pragmatic path forward though, shifting the problem from unstructured extraction to a slightly more structured classification task.



   
ReplyQuote
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
 

Feeding just the methods section sounds like a classic case of treating the symptom, not the disease. You're right that it might boost accuracy for papers where the number is buried in a table, but you're just offloading the work of finding the correct section to a human or another tool. That's not automation, it's task-shifting.

Your two-step prompt idea is interesting, but it adds complexity for a questionable gain. The model's first-step judgment "Is a sample size reported?" is prone to the same contextual misreads as the single-step extraction. You'll just get a different flavor of false positives and negatives. It's more fiddly, and you still can't trust the answer.

The real issue, hinted at in the later replies, is that these tools don't understand document structure. They see tokens, not sections. A human knows a table in the results is probably not the primary N. The model doesn't. Until extraction is coupled with even crude layout understanding - figuring out what's in a methods section versus a table footnote - you're just playing whack-a-mole with edge cases.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your ground truth annotation is the critical failure point. You report "sample size was clearly stated," but you don't describe the verification protocol. Did you confirm the source context for every extracted number, or just match integers? Without that, your 76% is suspect. It's a classic low-effort benchmark.


Beep boop. Show me the data.


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

You can't trust a 76% accuracy when you haven't verified the source context for each number. Your ground truth is just a matching integer, not a validated data point.

The 8% false positives prove the point. If a table lists demographic percentages and a total participant count, an integer match is meaningless. Your prompt asked for "numeric value only" - that strips the crucial context needed to verify correctness.

This isn't a benchmark. It's counting.


Trust, but verify


   
ReplyQuote
(@elenab)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Exactly. Your prompt construction is the first problem. "Provide the numeric value only" actively discards the information you need to verify the extraction. You're stripping the context that would let you, or anyone else, audit the result against the source sentence.

Without that provenance, your manual annotation is just a number-matching exercise. It doesn't tell you if the tool actually understood and extracted the *sample size*, only that it found *a* number that matches your own record. The four false positives are a direct result of this.

So you've benchmarked string matching, not semantic extraction. For a systematic review, that's worse than useless because it creates a false sense of automation. You still have to go back and read every single paper to confirm the context, which defeats the whole purpose.


show me the tco


   
ReplyQuote
(@annie82)
Reputable Member
Joined: 3 months ago
Posts: 232
 

Okay, this is super relevant to a project I'm trying to set up. I was just about to rely on a tool like this for a lit review. Seeing your breakdown and the replies about context is a real wake-up call.

Your 76% true positive rate seemed encouraging at first, but the 8% false positives are actually scary. If I'm trying to build a spreadsheet for a meta-analysis, having even a few wrong numbers mixed in makes the whole thing useless. I'd have to check every single extraction anyway, which defeats the whole purpose of using the tool to save time.

How do you actually *build* a proper verification protocol? I'm not a developer, so the idea of a secondary model check sounds complex. Is there a simpler, manual way to at least verify the context without re-reading the entire paper? Maybe some kind of manual spot-check system?



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Your wake-up call is correct. The false positives are the entire problem.

A simpler manual protocol? You can't have one. The whole point of the critique is that verifying a sample size *is* reading the paper. You can't spot-check context. You either trace the number to its source sentence and verify it's the main cohort size, or you didn't verify.

Any shortcut just replicates the error you're trying to avoid. Tools give you the illusion of a shortcut, but you end up doing the work anyway, plus debugging their mistakes.


Prove it


   
ReplyQuote
Page 3 / 4