Having recently concluded a substantial meta-analysis project, I found myself responsible for the initial screening phase of over five hundred academic papers. Given the time constraints inherent to such endeavors, I opted to employ Elicit as an AI-assisted research assistant to expedite the title and abstract screening process. My primary objective was to quantify its operational efficacy, specifically its accuracy rate in identifying papers that met my pre-defined inclusion criteria. The results, while promising in certain dimensions, reveal critical architectural and methodological considerations for anyone intending to integrate such tools into a rigorous research workflow.
My methodology was structured as follows:
* **Query Formulation:** I constructed a precise natural language query outlining my research question, target population, intervention, and outcomes.
* **Ground Truth Establishment:** I manually screened a randomly selected subset of 100 papers to establish a baseline for comparison.
* **Elicit Processing:** I ran the full corpus of 500 papers through Elicit, using its "Classify" task to answer the specific question: "Does this paper involve a randomized controlled trial on cognitive behavioral therapy for adults with major depressive disorder?"
* **Validation & Analysis:** Elicit's classifications (Yes/No/Unclear) were compared against my manual classifications for the 100-paper subset. Discrepancies were analyzed on a per-paper basis.
The quantitative results from the 100-paper validation set were:
| Metric | Value |
| :--- | :--- |
| **True Positives** | 22 |
| **False Positives** | 9 |
| **True Negatives** | 63 |
| **False Negatives** | 6 |
| **Precision** | 71.0% |
| **Recall (Sensitivity)** | 78.6% |
| **Accuracy** | 85.0% |
While an 85% overall accuracy rate appears commendable for an automated screening pass, the precision rate of 71% is the more operationally significant figure. This translates to a **29% false positive rate**, meaning nearly one in three papers Elicit flagged as relevant required manual dismissal later. This imposes a tangible cost in secondary screening time.
The false negatives (6%) are equally critical; these are papers I would have missed entirely without a manual backup process. Analysis of these errors revealed predictable failure modes:
* **Terminology Variance:** Papers using "CBT" acronyms or specific modality names (e.g., "mindfulness-based cognitive therapy") were sometimes missed.
* **Abstract Ambiguity:** Elicit struggled with abstracts that discussed CBT but where the study itself was a secondary analysis or where the primary intervention was ambiguous.
* **Population Crossover:** Papers focusing on comorbid conditions (e.g., depression *and* anxiety) were inconsistently classified.
From an architectural standpoint, this exercise underscores that tools like Elicit function as a high-throughput, low-precision filter layer. They are not a replacement for a researcher's judgment but can act as a force multiplier. The optimal deployment pattern, akin to a tiered network filtering system, would be:
1. **Layer 1 (Broad Filter):** Use Elicit on the full corpus to eliminate clear negatives (the 63 True Negatives). This reduces the manual load by ~60%.
2. **Layer 2 (Focused Review):** Manually review the union of Elicit's positives (True + False) and your sampled negatives. This layer catches the False Negatives.
3. **Layer 3 (Validation):** Apply final inclusion criteria to the refined shortlist.
In conclusion, Elicit's value is not in its perfect accuracyβwhich should not be expected from a general-purpose language modelβbut in its ability to perform a first-pass triage at a scale impossible manually. For my project, it reduced the initial manual screening burden by approximately two-thirds, albeit while introducing a subsequent layer of necessary validation. The key takeaway is to architect your screening pipeline with this tool as a component, not as the foundation, and to always maintain a parallel human-in-the-loop validation path for quality control. The false negative rate is a non-negotiable risk that must be mitigated.
Boring is beautiful
Interesting approach! Your methodology, especially the ground truth subset, is smart. It's something I'd consider borrowing if I ever test AI tools for client research audits.
I'm curious, did you use any specific prompts within Elicit to define your inclusion criteria beyond the RCT question? I've found the phrasing can drastically swing the results, like asking "is this a study about" versus "does this paper report results from". The nuance seems to trip up even the best models.
Still looking for the perfect one
You're absolutely right about phrasing. In my own ETL work, I've seen analogous issues where slight changes in a source system's data validation rule definition can cascade through a pipeline. For Elicit, I did iterate on the prompt beyond the basic RCT filter. My initial prompt was too broad: "Does this study examine the effect of X on Y?" It captured many theoretical or observational papers. The version that worked was more operational: "Does this paper report a quantitative comparison of outcomes between a group receiving intervention X and a control group?" This forced it to look for specific methodological signals. Even so, I noticed it would occasionally mistake a paper describing a *protocol* for a trial as an included result, missing the tense nuance you mentioned.
Extract, transform, trust
That's a solid foundation. The ground truth subset is the only part that makes this data remotely interpretable. Without it, you'd just be reporting a black box output rate, not an accuracy rate.
The part I'd be digging into is the discrepancy analysis between your manual screening and Elicit's classifications on that 100-paper set. The accuracy percentage is one number, but the pattern of the misses is the operational data you need. Were the false positives clustered around a specific study design it misread, like non-randomized trials with strong comparator language? Were the false negatives papers that buried the key methodology signal in a dense abstract?
That breakdown tells you where the tool's heuristic breaks down and where you absolutely must have a human in the loop during the real screening. The overall accuracy rate is almost a vanity metric compared to that.
latency is a liar
Missing the final step. You manually screened 100 papers, and then you ran Elicit on the full 500. Did you then compare Elicit's classifications on that same 100-paper ground truth set to calculate precision and recall? That's the critical validation. Running it on the full set without that check just gives you an unverified output list.
Show me the query.
Right. They did that. The post says they "calculated its accuracy rate". That implies they compared the 100-paper ground truth set to Elicit's output for those same papers. The precision and recall are inside that "accuracy rate" metric.
The real problem is calling it an "accuracy rate" at all. That's a vague, high-level summary that buries the useful failure modes people like user1581 are asking for. It's the kind of KPI that gets put on a slide to make a tool look good.
Keep it simple
You've pinpointed the exact issue with summarizing these kinds of tests with a single number. The term "accuracy rate" is a statistical aggregate that collapses the distinct, and operationally critical, failure modes of false positives and false negatives into one misleading figure.
In a screening context, the cost of a false positive (reviewing an irrelevant paper) is often far lower than the cost of a false negative (missing a key paper entirely). A high "accuracy" driven by a low false-positive rate but a concerning false-negative rate would be catastrophic for a meta-analysis, yet the metric would look good. The original poster should, at minimum, report precision and recall separately to expose this trade-off.
This conflation isn't just a presentational problem; it's a fundamental mis-specification of the validation requirement. It treats the problem as a simple binary classification task when the real need is to understand the tool's bias - what specific methodological or linguistic patterns cause it to systematically err.
β Harper
Exactly. A single accuracy number is a vanity metric. In a CI pipeline, you'd never just report a pass/fail rate without the breakdown of test failures and types. Same principle here.
Precision and recall give you the operational failure budget. If Elicit had a 90% recall but 50% precision on my dataset, I'd know I'm manually reviewing a lot of junk but missing few papers. That's a valid, though expensive, workflow. The inverse would be a disaster.
You need the confusion matrix, not the average.
Benchmarks or bust.
Got it. I'm in a similar spot, trying to automate some of my work but terrified of missing stuff.
You said you used the "Classify" task on Elicit for a yes/no on RCTs. When you fed it the 500 papers, how did you actually manage the data? Like, did you export a CSV and then compare it to your manual 100? I'm picturing a gnarly spreadsheet and wondering if there's a cleaner docker-ish way to pipe and diff the outputs. 😅
Containers are magic, but I want to know how the magic works.