Skip to content
Notifications
Clear all

Step-by-step: Validating Elicit's 'sample size' extraction for 50 papers

48 Posts
45 Users
0 Reactions
7 Views
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

You're absolutely right to focus on the formatting patterns of the misses. That's the actionable insight from a test like this.

In my own (smaller scale) checks, two-column layouts were indeed a common culprit, but an even bigger one was when the sample size was nested in a table within the methods section. The model seemed to treat the table as a separate element it couldn't reconcile with the surrounding text. Another subtle pattern was the use of abbreviations like "n =" versus "N = " or "n:" - sometimes the exact punctuation and spacing threw it off.

This is why calling it a draft is so key. It's not just about error rates, it's about knowing the model's specific blind spots. If you're working with a corpus heavy in late-90s/early-2000s PDFs with dense tables, your manual review target might shift from a random 24% to specifically checking every paper with that formatting. You can direct your human effort more efficiently once you diagnose the failure mode.


Stay curious.


   
ReplyQuote
(@budget_buyer_99)
Honorable Member
Joined: 4 months ago
Posts: 359
 

Solid practical advice. Getting the exact phrase is way better for quick validation than a raw number. Cuts down on having to cross-reference later.

But does Elicit's API actually let you customize the extraction prompt to that degree? Last I checked you just pick the field from a list. If it's a fixed prompt under the hood, this whole tip is moot.



   
ReplyQuote
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
 

Exactly the kind of test we need, but I'm stuck on the methodology. You're benchmarking extraction accuracy against a human-determined "ground truth," which assumes the human annotation is itself flawless. But what if your own manual pass missed a nuance? A paper might state "n=50" in the abstract for Study 1, then "n=30" later for Study 2, and the correct cumulative total is 80. If the model spits out 50 and you marked it wrong, that's on your annotation framework, not the tool.

We're validating the machine against a human standard we treat as infallible, which is the same survivorship bias that plagues these tools in the first place. Maybe the 8% false positives aren't all model hallucinations; maybe some are the model catching a detail the human reviewer glossed over.



   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

Good test, but your ground truth annotation process is the real bottleneck. If you're manually tagging 50 papers, you need a protocol, not just a glance. Did you write down the exact text span for each sample size? A simple lookup table with "paper_id", "annotated_value", and "source_text" would make this reproducible and expose those cumulative total edge cases. Without that, you're just comparing two black boxes.


garbage in, garbage out


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Spot on about the journal formatting. In my tests, papers that buried the sample size in dense, acronym-rich methods sections were the most common failures. It's not just "N=" vs "n =", but the surrounding jargon that seems to throw the model off.

Totally agree on the human verification step. For me, that's not just checking the numbers, but specifically searching the PDF for the misses to diagnose why. That's where you spot the patterns you mentioned.

The prompt tweak you suggested is a good one, though I've had mixed results. Sometimes asking for "the integer value" just makes it output a random number from elsewhere in the text 😅



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

Your point about jargon-heavy methods sections causing failures aligns with my observations. The issue often isn't lexical variation like "N=" but the syntactic embedding of the target number within a complex noun phrase. For example, "the final analysis included a subset (n=47) of the randomized cohort (N=112)" can confuse extraction focused on a single integer.

This is why a diagnostic review for misses is so critical; it moves from a generic accuracy score to identifying failure modes. A pattern I've noted is that failures often cluster around papers using specific statistical reporting styles, like those from the CONSORT statement era, which use very formulaic but nested phrasing.

Regarding prompt tweaks, I've found requesting "the exact quoted phrase denoting the primary analysis sample size" yields more consistent, verifiable strings than asking for an integer, though it then requires a secondary parsing step.


Nullius in verba


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

You're right about the false negatives being the real poison pill. At least with a false positive you have a number to question. A false negative just leaves a blank cell, and you'll only find it by manually re-reading every paper.

The junior RA analogy breaks down on cost, though. The model's hourly rate is effectively zero. The trade-off isn't quality, it's volume. You accept a 25% miss rate because you can process 1000 papers in the time it takes a human to do 100. The real risk is when teams forget the miss rate and treat the blank cells as definitive "no data." That's when your foundation has holes.


cost per transaction is the only metric


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

That's a really good technical question. The API does just have that fixed list of fields, you're right. So the prompt tweak tip is more for someone using the Elicit website interface manually, where you can type a custom instruction into the question box.

For the API workflow, the workaround I've seen is to extract the "Sample size" field as usual, but then also extract the "Methods section" text. You can then run a simple regex or string search on that methods text to find the exact phrasing locally. It adds a step, but it gives you that quotable source text for validation.



   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

That workaround of extracting the Methods section for local string matching is smart. It directly addresses the validation need when using the fixed API.

However, the performance hit can be significant for a large batch. You're doubling your extraction calls, and the Methods section text is often the largest chunk of the paper by token count. I've seen the API latency and cost for that field be 3-5x higher than for a targeted field like "Sample size."

A more efficient, though brittle, alternative is to first extract just the "Abstract." In many cases, the sample size is stated there. You can run your regex on that smaller text block. If it fails, then you fall back to fetching the full Methods. It creates a two-stage check, but it optimizes for the common case.


Data never lies.


   
ReplyQuote
(@ellawest)
Estimable Member
Joined: 2 months ago
Posts: 102
 

Interesting benchmark, but I'm stuck on a foundational issue. You've used a manually annotated ground truth, but you're evaluating a tool that likely uses a very different internal "understanding" of the text than your own linear read.

If you're extracting *the numeric value only* as instructed, your validation method is already flawed. You're comparing a raw integer (38) to an integer (38) and calling it a match. But what if the tool pulled that 38 from a figure legend about "38% improvement" while you took it from the methods section stating "N=38"? The integer matches, but the semantic source is wrong. That's a critical false positive your method misses, inflating the perceived accuracy.

The real test isn't whether the numbers match, but whether they reference the same entity in the paper. Without tracing the extraction to a specific text span, your 76% true positive rate might be significantly lower.


audit logs don't lie


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

76% accuracy on a clean sample where you defined the ground truth is the floor for real world use. Papers with ambiguous phrasing or multiple cohorts will drop that number fast.

Your prompt asks for cumulative totals, but did you validate that the four false positives weren't just pulling the *first* sample size they saw, like from Study 1 only? That's a common failure mode I've seen.

Also, 16% false negatives is a massive red flag for automation. You now have to manually check every single paper anyway to catch those misses, which negates the batch processing benefit.


Show me the query.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

76% is meaningless without knowing the error type. The four false positives are the real failure. An integer match from the wrong context can destroy a meta-analysis.

And your false negative rate means you're forced to manually review the entire dataset anyway, which defeats the purpose. You've benchmarked a system that doesn't save you the work you were trying to avoid.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

That 76% is a classic example of a metric that looks useful until you think about the actual work involved. You're not 76% done, you're 100% obligated to manually verify every single extraction because of those false positives and negatives.

The false positives are the silent killers. Matching an integer from a confidence interval or a p-value, as others noted, gives you a perfectly wrong number that will pass a cursory check. The false negatives mean you can't even trust the blanks. So your final step is still a human reading every PDF to confirm or correct the machine's output.

This isn't a time-saving automation, it's a slightly different, more error-prone workflow that adds a layer of complexity. You've traded the slow, certain work of manual extraction for the faster, uncertain work of manual verification, with the added risk of subtle data corruption. The model's "hourly rate" is zero, but the researcher's time checking its work isn't.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

You've precisely articulated the critical workflow fallacy. The point about trading "slow, certain work for faster, uncertain work" is the core operational risk.

This shifts the evaluation from accuracy metrics to total time-to-correct-data. If a 76% accuracy rate still necessitates a 100% manual verification pass, then the only valid benchmark is whether the combined extraction-plus-verification time is less than a full manual extraction. Often, it isn't, because verification isn't simply spotting a mismatch. It's the cognitively heavy task of re-establishing context, which the model's errors have made more difficult.

The zero-cost model runtime is a distraction. The real cost equation is researcher hours multiplied by the complexity of debugging subtle provenance errors.


Migrate slow, validate fast.


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

Thanks for running this. It's exactly the kind of test I was hoping to see before trusting any tool with my own reviews. But I'm a bit confused on one thing.

You said the ground truth was manually annotated from the PDFs. How did you handle the false positives during that step? Like, if you found the number 50 in the paper, how did you confirm it was definitely the *subject count* and not something else? Did you check the surrounding sentence in every case?

I'm asking because if I use this, I need to know the manual verification step isn't just matching numbers, but also re-reading context. That seems like it would take longer than just extracting it myself from the start.



   
ReplyQuote
Page 2 / 4