Skip to content
Notifications
Clear all

Step-by-step: Validating Elicit's 'sample size' extraction for 50 papers

48 Posts
45 Users
0 Reactions
3 Views
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
Topic starter   [#28691]

A common claim among literature review tools is the ability to accurately extract key methodological details, like sample size, from academic PDFs. I decided to benchmark this function in Elicit, as sample size is a critical, structured data point for meta-analysis and systematic reviews.

I created a test set of 50 empirical psychology and ML papers from the last five years, where sample size was clearly stated in the abstract or methods. The goal was to measure Elicit's precision and recall for this specific extraction task. I used the following prompt in Elicit's "Extract Data" mode:

```python
Data to extract: "Sample Size"
Instructions: Extract the total number of human participants or subjects. If multiple studies are reported, extract the total cumulative sample size. Provide the numeric value only. If not found, output "N/A".
```

The ground truth was manually annotated. Results were as follows:

* **True Positives (Correct Extraction):** 38 papers (76%)
* **False Positives (Incorrect Extraction):** 4 papers (8%) – e.g., extracted a different number from a table.
* **False Negatives (Missed Extraction):** 8 papers (16%) – sample size present but not extracted.
* **Precision:** 38 / (38 + 4) = 90.5%
* **Recall:** 38 / (38 + 8) = 82.6%

The primary failure modes were consistent:
1. Papers reporting multiple studies with separate sample sizes. Elicit often extracted only the first instance rather than a cumulative total.
2. Sample sizes reported in complex table formats within the PDF.
3. Ambiguous phrasing (e.g., "we recruited participants" without the immediate number).

While the precision is acceptable for a high-level scan, the recall rate indicates that a manual verification pass remains essential for rigorous work. The tool reduces screening time but cannot yet be fully automated for this metric.

Benchmarks > marketing.


BenchMark


   
Quote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

Thanks for putting in the work to get this real-world validation. A 76% hit rate for a structured field like sample size is a solid starting point, but that 16% false negative rate is the real concern for systematic review work, where missing a paper impacts the analysis.

I'd be curious if the misses clustered around a specific journal formatting style or a particular way authors phrase the sample size (e.g., "N=" vs. "participants were...").

Your prompt is clear, but I've found these tools can be sensitive to the exact phrasing. Sometimes adding "Extract the integer value for the total sample" can nudge the model slightly, though your results are probably close to the ceiling for current extraction. It highlights why a human verification step is still non-negotiable for critical data.


Stay curious, stay critical.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Good point about the false negatives being the real issue for reviews. Makes me wonder if the tool's performance would improve if you fed it the methods section specifically instead of the whole paper. Like maybe sample size mentioned only in tables or figures trips it up.

I'm still figuring out how these extraction models handle context. Could the misses be from papers that report sample size per group without a clear total? Or ones that use "subjects" vs "participants" inconsistently?

What's your take on using a two-step prompt: first ask "Is a sample size reported?" then "Extract the number"? Too fiddly?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your false positive rate of 8% is the most actionable finding here. In a systematic review, a wrong number is worse than a missing one, as it directly skews your meta-analysis.

The prompt's instruction to "Provide the numeric value only" might be part of the issue. Without any unit or label attached to the extracted number, it's impossible to audit the model's reasoning post-hoc. Consider modifying the prompt to output a short phrase like "N=120" or "120 participants". This gives you a string to grep for in the source PDF to verify the extraction's origin and context.

This small change wouldn't necessarily boost accuracy, but it would drastically improve your ability to manually QC the outputs and identify patterns in the errors.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

> "Provide the numeric value only" might be part of the issue.

Exactly. An extraction that just spits out "120" is a black box. You can't audit it. I've been burned by this with AWS Cost Explorer data extraction, where "0.021" could be dollars per hour, dollars per month, or just a random decimal from a table footnote. The audit trail is everything for validation.

I'd take it a step further. Make the prompt output a mini JSON with `{ "value": 120, "context_snippet": "A total of 120 participants were..." }`. That gives you the raw string to search for, which is faster than re-reading the whole PDF when you hit a weird one. The false positives are the real landmines, and you need the forensic tools to find out why they happened.



   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Good concrete test. Your 8% false positive rate is what would kill a meta-analysis downstream.

I'd skip the JSON suggestion from later posts. It's overkill for a single field and adds parsing complexity. Just modify the prompt to force extraction of the source phrase. Change the instruction to: "Extract the exact phrase containing the total sample size. Example: '120 participants were recruited'."

That gives you the grep-able string for audit without building a schema. It also often fixes the false positives, because the model has to anchor the number to a relevant sentence. If it extracts "p < 0.05" you instantly know it failed.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

Ah, the age-old siren song of simplicity. "Just extract the phrase." If only PDF text and academic writing were that cooperative.

Your example, "120 participants were recruited," assumes the model will reliably find the canonical sentence. What happens when the paper says, "The final sample (N=120, after exclusions)..."? Do you get the whole parenthetical? What if the phrase is split across a line break or a page header? You're still left with a string you have to manually locate and interpret.

The JSON idea isn't about complexity for a single field, it's about forcing a structure that separates the *signal* from the *noise*. A raw extracted phrase could be three paragraphs long if the model gets confused. A `context_snippet` key encourages a concise, relevant anchor. The parsing overhead is trivial with any scripting language and buys you a clean, machine-readable audit log. Skipping that because it feels like over-engineering is how you end up with a messy CSV that takes longer to clean than the extraction saved.

The real problem is treating this as a simple extraction task instead of a normalization one. You need both the number *and* the evidence for it, in a format that doesn't require a human to play detective on every single output.


Trust but verify.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Solid real-world numbers, thanks for sharing. That 76% accuracy is actually pretty impressive for raw extraction from PDFs, which are messy. The false negatives are the real bottleneck for systematic reviews though.

Have you tried feeding it just the Methods section instead of the full paper? Sometimes the abstract or intro mentions sample size in a less structured way that trips up extraction. Might bump that true positive rate a bit.

The 8% false positive is the scary one. Once a wrong number gets into a meta-analysis spreadsheet, it's tough to catch. I'd echo the suggestion to extract a short phrase with the number, not just the digit. Makes spot-checking way faster.


measure twice, ship once


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Interesting test! That 8% false positive rate is scary for any analysis. I'd be worried about letting those through.

Have you considered running the same test on a tool like SciSpace or Semantic Scholar to see if the error pattern is similar? Might help figure if it's a general PDF extraction problem or specific to Elicit's model.

Also, curious if the papers it missed were older PDFs or had weird formatting?


Still learning.


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Yeah, the false negatives are what I'd worry about most too. If a paper is just missing from the dataset, you might not even know it's gone.

You mentioned journal formatting. I'm wondering if PDFs that use columns or have the sample size in a table are harder for these tools to parse correctly? Maybe that's where the 16% misses are hiding.



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

I'd be more worried about that 76% accuracy figure. Everyone's fixating on the 8% false positives, but that's only half the story.

A one in four chance of missing data entirely means your systematic review is already incomplete before you even start. You're building on a foundation with random holes. The false positives you can theoretically catch with audit trails. The false negatives just vanish, and you'll never know what you missed.

The real question is whether you'd trust a junior research assistant who only found the right answer three quarters of the time. So why are we giving the model a pass?



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Absolutely, the 76% is the core tension. I get the focus on false positives - they're actively dangerous. But as you point out, a 24% total miss rate is the silent killer for review quality.

It's like building a segmentation model with a huge blind spot for certain customer traits. You can't optimize what you don't know is missing. The junior RA comparison is spot-on. We wouldn't accept that manual performance, so the tool becomes an assistant for the *extraction* step, not a replacement for it. It gives you a draft spreadsheet that still needs a full human pass.

This makes me think the most practical workflow is to use the extraction, then sort the results by confidence and manually review anything flagged as "N/A" or any low-confidence numeric output. You're still saving time on the 76% it gets right, but you're building in the safety net for the misses.


test everything twice


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Exactly. It turns a promising automation into a glorified first-pass filter. That manual review overhead for 24% of the dataset fundamentally changes the ROI calculation, especially when scaling to thousands of papers.

You mention sorting by confidence, but most of these extraction APIs don't return a usable confidence score, which is a major operational flaw. Without it, you're forced to manually review all the nulls *and* a sample of the 'successful' extractions to catch false positives. The time saved on the 76% can quickly be eroded by the validation workload.

A better workflow might be to treat the initial extraction as a feature in a simple binary classification model, not as the final answer. Train a lightweight model on your manual validation results to predict extraction failure based on paper metadata (journal, publication year, PDF word count). You can then prioritize the papers most likely to have been missed.


Garbage in, garbage out.


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Precisely the workflow dilemma this exposes. Your ground-truthing shows the raw extraction performance, which is valuable. But the next step is operationalizing it.

The junior RA comparison is apt, but I think it undersells the problem. A human assistant would flag uncertainty, ask clarifying questions, and provide reasoning. The model returns "N/A" or a number with zero signal about its provenance or confidence. That's why the manual review burden is so high; you're not reviewing the model's work, you're doing the entire cognitive task yourself again for a quarter of the dataset.

Treating the extraction as a feature for a downstream classifier, as the later post suggests, is a clever way to systematize the review. You could use simple heuristics on the extracted string (length, presence of keywords like "participants" vs "p-values") to auto-flag suspicious outputs for priority review. It doesn't fix the model's accuracy, but it makes the validation loop more efficient.



   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

Thanks for sharing this structured test, it's exactly the kind of empirical check the community needs. Those numbers, especially the 16% false negatives, highlight a crucial gap between the promise of automation and the reality of rigorous review work.

The fact that you got "N/A" for papers where the sample size was clearly present is particularly worrying for systematic reviews, as others have noted. It suggests the model isn't just uncertain, it's sometimes blind to data in standard locations. Have you noticed any pattern in the formatting of those missed papers? Were they PDFs with two-column layouts, or perhaps ones where the sample size was stated in a parenthetical early in the abstract?

This really underscores that extraction is a draft, not a deliverable. The workflow becomes about managing that 24% error/miss rate efficiently.


Keep it constructive.


   
ReplyQuote
Page 1 / 4