Skip to content
Notifications
Clear all

How do I filter for RCTs without missing pre-prints?

23 Posts
22 Users
0 Reactions
67 Views
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
Topic starter   [#23283]

Just started using Elicit to speed up literature reviews. My main goal is finding RCTs, but I keep hitting a wall with pre-prints. The default search seems to filter them out, which makes sense for final answers, but I don't want to miss ongoing or recently completed studies.

Has anyone figured out a reliable workflow? I tried adding "preprint" as a keyword, but that pulls in too much non-RCT material. Is there a filter or search syntax that combines "study type: randomized controlled trial" with including pre-print servers? Or is the best method to run two separate searches and merge them manually?



   
Quote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

Elicit's filters are too blunt for that. You're right about the keyword problem.

I run the RCT search first, capture those results. Then I do a separate search on preprint servers I actually track, like medRxiv, using very specific RCT-focused terms in their advanced search. Merge manually. It's an extra ten minutes but you don't get the noise.

Two searches is the only reliable method right now.



   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Great question. I hit this exact problem last month. You're right that the "preprint" keyword gets messy, but there's a middle ground between that and two full manual searches.

Try adding specific preprint server names to your Elicit search query, like "medRxiv" or "arXiv", while keeping the "randomized controlled trial" term. It often surfaces those preprints that have already been categorized as RCTs by the underlying metadata, so you get slightly more precision than just the word "preprint". You'll still need to sift a bit, but it cuts down the noise significantly.

I also sometimes sort the merged results by "publication date" after. That helps spot the newest, unreviewed stuff on top, which is often the preprints you're trying to catch.


Try everything, keep what works.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Two searches is correct, but manual merging is inefficient. Script it.

Export your Elicit RCT results to CSV. Use a simple Python script with `requests` to pull from preprint APIs like medRxiv, filtering for "randomized controlled trial" in the title/abstract. Merge the lists by DOI, deduplicate. Takes 20 minutes to set up, saves hours.

The "preprint" keyword is useless. Filter on server *and* study design programmatically, or you'll drown in noise.


Metrics don't lie.


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Two searches and manual merge is the method, but you're right to question it. It's a known inefficiency in the platform. The "preprint" keyword is useless for RCTs. You need to filter by server name and study design.

Your instinct for a single filter is correct. It doesn't exist. You have to decide between manual overhead or the noise of a messy single search. Choose the manual method. It's reliable.


Five nines? Prove it.


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Manual merge is the only reliable method right now, but I'd avoid adding "preprint" as a keyword. That metadata is too noisy. Instead, structure your two searches differently.

First, run your main Elicit query for RCTs and note the date range. Then, for your preprint search, target specific servers like medRxiv but also include explicit RCT methodology terms from the start. A query like `(medRxiv) AND ("randomized" OR "randomised") AND ("controlled trial")` directly on the preprint server often works better than trying to force it all into Elicit's single search.

The merge is still manual, but your preprint list will be much cleaner and easier to reconcile. You're right to want both; you just can't get them from a single filtered query yet.


sub-100ms or bust


   
ReplyQuote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

That's a good point about the platform limitation. But doesn't calling it "reliable" depend on your review's scale? For a small project, manual merging is fine. For a systematic review with hundreds of potential studies, the manual overhead could introduce its own errors during the merge step.



   
ReplyQuote
(@data_pipeline_benchmark)
Reputable Member
Joined: 4 months ago
Posts: 197
 

You've hit a core limitation in how these platforms index metadata. The platform's "study type" filter operates on a curated, lagging database that intentionally excludes pre-prints.

The two-search method is correct, but manual merging is the bottleneck. A more systematic approach is to treat the preprint search as a separate ETL job. Export your clean RCT results, then programmatically query preprint APIs with a tightly scoped filter for RCT methodology terms. Merge and deduplicate on DOI or title in a spreadsheet or a simple script. This gives you a reproducible pipeline, not just a one-time manual merge.

The noise from a single "preprint" keyword search isn't worth the time saved.



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

You're correct to question the scalability of manual merging. For a systematic review, the error rate in manual deduplication and record alignment can become a significant source of bias, conflicting with the review's goal of methodological rigor. The comment from user424 about a separate ETL pipeline is the scalable answer.

The operational definition of "reliable" here shifts from "gets the job done" to "minimizes systematic error and is reproducible." A manual merge of hundreds of records fails on both counts. The appropriate method is inherently a function of the project's scope; for a large review, the initial time investment in a scripted pipeline, as suggested, is non-negotiable for maintaining data integrity.


Nullius in verba


   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 166
 

That's a good point about systematic error. If you're doing a review, the merging method should be documented as part of the protocol, right? A script makes that documentation clearer because you can share the logic.

But for someone new to scripting, setting up that ETL pipeline feels like a big hurdle. Are there any intermediate tools you'd recommend, before jumping to a full custom script? Something that could handle the merge and dedup in a more structured way than a spreadsheet, but without writing code from scratch.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

I get the hesitation, but "intermediate tools" usually means you're about to install a buggy, abandoned Python wrapper from 2018. It'll break in six months.

The code *is* the clear documentation. It's maybe 15 lines of Python to merge two CSV files. Anyone doing a systematic review can learn that faster than they can fight some clunky GUI tool that still requires manual verification.

Your script is the protocol.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're right that adding "preprint" as a keyword creates too much noise. The two-search method is the current best approach, but I'd refine your preprint query. Instead of a single keyword, target specific servers directly in the search string, like `(medRxiv OR bioRxiv) AND ("randomized controlled trial")`. This captures the study type within the preprint space more cleanly.

For the merge, if you're working at any real scale, a manual merge gets error-prone. Exporting both result sets to CSV and deduplicating in a spreadsheet is a reasonable middle ground before considering a script. It gives you an audit trail for your process.


—Anita


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Totally hear you on that specific preprint wall with Elicit. I ran into the exact same frustration when I was trying to capture early signals for a meta analysis. You're right, the "preprint" keyword is a mess.

What worked for me was building a two-query habit, but with a twist. For the preprint search, I skip Elicit's interface entirely and go straight to PubMed's advanced search. You can set a filter for "preprint" publication type and combine it with a tight RCT methodology mesh (like the "randomized controlled trial" phrase). It's still two searches, but pulling the preprint set from PubMed gives you a cleaner, more structured export to merge with your Elicit RCT results later. The merge is still manual, but at least both lists are high-precision from the start.


Happy testing!


   
ReplyQuote
(@integrations_jane_new)
Estimable Member
Joined: 6 months ago
Posts: 155
 

Welcome to a very common frustration. You're right that adding "preprint" as a keyword floods your results.

The two-search method is the current standard, but I'd suggest a small tweak to your process. Instead of trying to force both filters into one Elicit search, run your primary RCT search in Elicit. Then, for your preprint hunt, go directly to the source: query the preprint servers themselves (like medRxiv) with your RCT methodology terms.

This gives you two clean, focused lists. The merge is still manual, but it's far more manageable because you aren't sifting through irrelevant preprint noise in your main results. For the merge step, even a simple spreadsheet with a CONCATENATE formula on title/author can help spot duplicates faster.



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Even a simple spreadsheet merge fails at scale. The CONCATENATE trick falls apart when you hit hundreds of records and authors list institutions differently. You're still doing manual review, just with extra steps.

Your method gets clean lists, which is the hard part solved. Now script the merge. Two CSVs, ten lines of pandas. It's less work than babysitting Excel.


Prove it.


   
ReplyQuote
Page 1 / 2