Hey everyone, I've been using Elicit for a few weeks now to help with literature reviews for my analysis projects, and I wanted to share a workflow I've been refining. I often need to pull out specific population details (like sample size, age range, country) from papers, and the default columns weren't always giving me exactly what I needed. So I started experimenting with custom columns.
I found that being really specific in the prompt for the custom column is key. For example, instead of just asking for "population," I now create separate columns for each detail. Here's what one of my prompts looks like for extracting the **mean age**:
```text
Extract the mean age of the study participants as reported in the abstract or methods. Return only the number, or "Not stated" if not found.
```
And another for **sample size**:
```text
What is the total number of participants in the study? Provide only the numeric figure.
```
This has made my data so much cleaner for exporting to Excel. I can just export the table and then sort or filter studies by sample size right away. It saves a ton of time compared to manually scanning each paper.
Has anyone else tried something similar? I'm curious if there are other specific data points you've found useful to extract with custom columns, especially for meta-analysis type work. Also, I sometimes get a mix of formats (like "45" vs "45 years old") – any tips for making the extraction more consistent?
That's a solid approach for a systematic review, but I'm always a bit wary of trusting extracted numbers without a spot check. These tools have a nasty habit of "hallucinating" plausible-looking data, especially with numeric fields.
I'd recommend adding an audit column with a prompt like "Copy the exact sentence containing the sample size." Lets you verify the output against the source text later. The export is clean, but garbage-in-garbage-out still applies.
How often do you find yourself having to manually correct the figures it pulls?
Trust but verify.
The audit column is a good idea in theory, but it's just moving the problem. If the tool hallucinates a number, why wouldn't it also hallucinate or misattribute a source sentence for the audit trail? You're building a verification step on a potentially faulty base.
The core issue is treating these outputs as data. They're not. They're probabilistic text completions. You need to design the system expecting the extraction to be wrong.
I'd skip the manual column verification and push for a programmatic check instead. Flag any extracted figure where the supposed source sentence from the audit column doesn't actually contain a number, or where the number doesn't match. That's the only way to scale a review.
Least privilege is not a suggestion.
Exactly. This is why the "source sentence" column is the only one you can trust. The model can invent a number, but it can't invent a credible sentence fragment from a PDF it's processing. The mismatch is the signal.
You build your sanity check by comparing the extracted number against the text in that audit column using a simple regex. If "27" is extracted but the quoted sentence says "participants were aged 18-35", your script flags it. The hallucination isn't in the audit trail, it's in the gap between the trail and the extraction.
Treat the raw text as ground truth. The extracted column is just a candidate.
Prove it.
I agree that >the "source sentence" column is the only one you can trust< as a baseline, but this trust hinges on the model accurately identifying the relevant text span. In practice, I've seen instances where the extracted sentence fragment is adjacent to but not containing the key detail, or it's from a different section like the discussion instead of the methods.
Your regex sanity check is useful, but it's brittle when dealing with implied or computed values. If the source sentence states "ages ranged from 18 to 35" and your custom column prompts for mean age, a correct extraction of "26.5" would be flagged as a mismatch. You'd need to incorporate some logic to handle ranges, medians, or other derivations.
This isn't so different from monitoring metrics in distributed systems, where you correlate logs with metrics but always account for parsing errors. The audit trail gives you a trace, but you still need human judgment for edge cases.
Okay that makes sense, treating the source text as the ground truth. But what do you do if the audit column itself comes back with something like "Not stated"? Then you don't have a sentence to check against at all.
Is the workflow then just to manually review all those "Not stated" entries? Seems like that could still be a lot of papers.
That's a great starting point for structuring your extraction. I use a similar method when comparing vendor proposals, asking for things like "Incident response time SLA in hours" in a dedicated column.
One thing I'd add from that experience is to build a "confidence" column right from the start. For your mean age prompt, you could have a second custom column with a prompt like: "Does the text explicitly state 'mean age was X'? Answer Yes or No." It gives you a quick filter later. You'll see which papers reported it directly versus where Elicit might be calculating or inferring, which is where a lot of the verification issues other folks mentioned can creep in.
buyer beware, but buy smart