Skip to content
Notifications
Clear all

Elicit for clinical trial screening - real user experience for a 200-user hospital

15 Posts
14 Users
0 Reactions
2 Views
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
Topic starter   [#29221]

Alright, let's get this out there before another round of "AI will solve all our clinical research problems" hype. We've been using Elicit for clinical trial screening at a 200-user regional hospital for the past eight months. The pitch was compelling: automate the systematic review process, find relevant trials faster, reduce manual PubMed/ClinicalTrials.gov slogging. The reality, as usual, is a mixed bag of genuine utility and frustrating over-promises, wrapped in a workflow that doesn't quite fit the messy, compliance-heavy world of actual hospital operations.

First, the good—because there is some:
* **Rapid literature sifting:** For broad, initial sweeps across a wide range of conditions, it's undeniably faster than a junior resident manually crafting perfect Boolean strings. You get a spreadsheet output in minutes that would take hours manually.
* **Concept extraction:** It's decent at pulling out PICO (Population, Intervention, Comparison, Outcome) elements from abstracts, which saves some highlighting and note-taking.
* **Cost for scale:** Compared to hiring additional full-time research coordinators purely for screening, the subscription fee is a rounding line in the budget. That's the business case that got it approved.

Now, the parts that make me want to throw my keyboard, usually stemming from the naive assumption that clinical research is a clean, academic exercise rather than a bureaucratic minefield:

* **The "last mile" problem is a canyon.** Elicit gives you a CSV. Great. That CSV then needs to be:
* Manually validated against source abstracts (because you cannot, under any circumstances, trust AI hallucinations with trial eligibility).
* Imported into our clinical trial management system (CTMS), which involves a byzantine mapping exercise because our CTMS API is from the Stone Age.
* Annotated with internal institutional review board (IRB) statuses, principal investigator (PI) interest, and resource availability—none of which Elicit knows or could know.
* This "last mile" eats up 70% of the effort. The automation saves the first 30%.

* **It's built for researchers, not hospital systems.** No real user management beyond "shared login." No audit trail compliant with 21 CFR Part 11 if you're doing regulated research. No integration with hospital Single Sign-On (SSO). We had to build a clunky wrapper around it with our own logging to track who ran what query and when, for compliance.

* **The search is only as good as the source data, and it's opaque.** When it misses a key trial—and it does—debugging why is a black box. Was it the prompt? The model's interpretation? A lag in its database update? You're left guessing, which is professionally unnerving when the stakes are patient eligibility.

Here's a snippet of the kind of Frankenstein workflow we've ended up with, which is the opposite of the sleek, automated future we were sold:

```python
# Not actual code, but a sad depiction of our process
1. User prompts Elicit via manual web interface.
2. Export CSV, upload to secure internal SharePoint.
3. Custom script (homegrown) validates DOI/PMID links, flags missing abstracts.
4. Manual review by senior research nurse (gold standard).
5. Another script attempts to transform CSV into CTMS-compatible XML (fails 30% of the time).
6. Manual import + data entry for the failures.
7. Weekly reconciliation meeting to discuss discrepancies. Yes, a *meeting*.
```

So, is it worth it? Cautiously, yes, but with massive caveats. It's a powerful **assistant**, not a solution. It shifts the workload from "finding needles in a haystack" to "validating the needles the AI found and searching for the ones it missed." If your organization expects a button that says "Find All Relevant Trials," you will be disappointed. If you have the in-house, grumpy infrastructure talent (like yours truly) to build the guardrails and integration glue, and your research staff understand it's a tool for generating a *starting point*, not an answer, it can provide a marginal efficiency gain. Just don't believe the marketing slicks, and for the love of all that is holy, budget for the significant hidden costs of integration, validation, and training.


monoliths are not evil


   
Quote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your point about the cost comparison is interesting, but I think it's missing the operational latency overhead. The subscription fee might be a rounding line, but have you measured the human-in-the-loop delay introduced by the platform? In my own benchmarks of similar tools, the time spent validating the AI's "decent" concept extraction often negates the initial speed gain. You're still committing a human to fact-check every PICO element against the source, and the tool's API call latency can add seconds of idle time per abstract during a large screening session. The real cost isn't just the license, it's the senior researcher's wait states.


--perf


   
ReplyQuote
(@alexh)
Estimable Member
Joined: 3 months ago
Posts: 103
 

That's a solid point about the hidden time cost. We found the same - those seconds per abstract really add up when you're screening 500+ results, and it creates this weird micro-interruption rhythm.

Did you notice if the validation fatigue changes based on the researcher's experience level? Our junior staff seem to second-guess every AI extraction, which drags things out even more.



   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

That's a really helpful breakdown. The cost angle is something I'm trying to understand for our own team's business case.

You said the subscription is a rounding line compared to a full-time coordinator. But does that still hold if you have to factor in the senior researcher time for validation, like the others mentioned? It seems like the tool shifts the workload rather than eliminates it, and senior time is usually the most expensive.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

It holds, but barely. The business case assumes you're reallocating that senior time from manual searching to validation, not adding a new task. If they were already spending 20 hours a month searching, and now they spend 10 validating, that's a net gain even at a higher hourly rate.

The trap is when it creates *new* screening work because the tool surfaces more low-quality leads. That's where the senior time cost flips the equation. You need to track pre and post-tool time allocation, not just output.


Beep boop. Show me the data.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

That's the critical pivot. Your point about tracking pre and post-tool allocation is the only way to measure real impact, but it's also where most pilot projects fail.

They track abstracts screened per hour, not the senior researcher's calendar. If the tool surfaces a 30% increase in leads that need human triage, you've just expanded the workload instead of shifting it. The net gain disappears.

The business case only holds if the tool's precision is high enough to keep the total volume of human-touched items the same or lower. Has anyone seen a tool that actually manages that consistently in a clinical setting?


—AF


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

That's exactly the question I'm wrestling with as we try to build our own case. When you say it shifts the workload, do you think that shift could actually be helpful if it's moving senior time from searching to more strategic validation? Or is the validation part just as tedious?

I'm coming from a project management background where we've seen similar things with automation tools - they promise time savings but sometimes just move the time expenditure to a different, more expensive part of the process. The senior time cost seems like it could quietly blow up the whole value proposition if you're not careful.

How are you measuring the "senior researcher time" in your own calculations? Is it just estimated, or are you tracking it somehow?



   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Totally feel that micro-interruption rhythm you mentioned. It's like the tool adds a tiny loading screen to every single decision point.

On the junior staff thing - we saw the opposite, actually. Our juniors tend to trust the AI extractions too much, maybe because they're less confident in their own reading? They just blaze through. It's the senior folks who get bogged down second-guessing. They spot the subtle mismatches that a junior would miss, but it makes the whole process slower for them.

Have you tried mixing up the screening pairs? Maybe a junior and senior together on one session to balance speed vs accuracy?



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

"the subscription fee is a rounding line"

For a 200-user hospital? Let's see the math. Even at their cheapest team plan, scaled for 200 seats, you're looking at tens of thousands annually. That's not a rounding line, that's a dedicated part-time coordinator's salary.

You're also assuming zero integration or training overhead, which is never the case. The actual cost is license + senior time for validation + IT overhead. That adds up fast.


show the math


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

You've nailed a key dynamic there. That trust spectrum between junior and senior staff is real, and it directly impacts how you calculate the tool's efficiency.

We tried the pairing idea you mentioned. It helped accuracy, but it tanked throughput. You're paying two people for one job, and the conversation overhead itself becomes a new time sink. The senior kept stopping to explain *why* they were questioning an extraction, which is great for training but terrible for a high-volume screening sprint.

Our compromise was to split the workflow: juniors do the initial pass with the tool's extractions flagged, and seniors review a random sample plus all the juniors' "maybe" pile. That at least contains the senior time to a defined batch instead of a per-abstract drag.


Ask me about my RFP template


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The "rounding line" claim is where these pitches always get slippery. Compared to a full-time coordinator's total comp, sure. But you're not hiring a coordinator *just* for screening, you're hiring them for a dozen other tasks, and the screening gets folded in. The tool's subscription is a new, dedicated line item that buys you exactly one thing.

So the real comparison isn't tool vs. salary, it's tool vs. the *portion* of that salary allocated to manual search hours. That's a much smaller denominator, and suddenly the subscription looks a lot less like a rounding error.


Show me the data


   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That's a really good point about the denominator. It shifts the whole cost-benefit look.

When you say "portion of that salary allocated to manual search hours," how are people actually calculating that? Is there a standard method, or is it usually just a rough estimate based on time logs?



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

I'm with you on the mixed bag feeling. That rapid sifting is fantastic, but the spreadsheet output can be a false summit. We've found we spend almost as much time cleaning and reformatting that export for our internal systems as we would have spent searching manually.

And that concept extraction for PICO... it's good until it's critically wrong. We had it swap an intervention and a comparison in a cardiology abstract, which totally flipped the study's meaning. It wasn't a common error, but it only takes one to undermine trust in the whole batch. Now we treat every extraction as a "suggestion" that needs cross-checking, which adds back a lot of the time we were supposed to save.

The subscription cost point is interesting, but have you hit any walls with their API or bulk processing limits? That's where our "rounding error" started to add up quickly when we tried to scale.



   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

That rapid sifting is a huge appeal. We're exploring similar tools at my smaller clinic. But the spreadsheet output - how do you handle the results after you get that CSV? Is it easy to feed into your existing trial management system, or does it just become another manual data entry step?



   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Totally agree on the rapid sifting being a game-changer for that first pass. That speed is so compelling.

But when you say "the subscription fee is a rounding line," I'm curious if that's just for the screening license itself? Our admin folks are already asking about the hidden costs - like if we need to pay for extra API calls if we scale up, or costs for storing/backing up all those generated spreadsheets securely. It feels like the base subscription might be the tip of the iceberg.



   
ReplyQuote