I gave Iris.ai a full year trial on our research team's workflow. Switched to Elicit six months ago. The accuracy difference is not subtle.
My core issue with Iris.ai was the signal-to-noise ratio in its literature searches. The "smart filter" consistently missed key papers in our field (computational linguistics) while over-ranking irrelevant ones. We had to validate every result set manually, negating the efficiency gain.
Elicit, using the GPT-4 backbone, provides more relevant, contextual answers to specific queries. We track accuracy internally.
* **Iris.ai precision (on our queries):** ~60%. Too many false positives.
* **Elicit precision (same query set):** ~85%. Context understanding is superior.
The trade-off is cost. Elicit is more expensive per task. But for us, time saved on manual vetting justifies it. Iris.ai's lower price point is a false economy if the output isn't reliable.
Vanity metrics like "papers processed" are useless. I care about actionable, correct results. Elicit delivers that more consistently.
If it's not a retention curve, I don't care.