Everyone seems to be praising ResearchRabbit's visual discovery maps, but I'm skeptical about the core product: finding relevant papers. The interface is nice, but is it actually better at surfacing what you need?
Has anyone done a proper, side-by-side comparison of discovery relevance against tools like Elicit or Semantic Scholar? I'm talking about controlled tests. Not just feelings.
I'd want to see benchmarks on:
* **Query precision:** For the same seed paper or research question, what percentage of the top 10 recommended papers are genuinely relevant?
* **Recency bias vs. foundational work:** Does one tool disproportionately favor newer papers, missing key older studies?
* **Escape from filter bubbles:** If you start in a niche field, how quickly does each tool introduce concepts from adjacent disciplines?
My concern is that ResearchRabbit's "similar work" chain might just create an aesthetic echo chamber. It feels productive because you're following a path, but you might be missing the seminal paper that exists three fields over, which Semantic Scholar's broader network might catch.
I haven't seen any rigorous data on this, just anecdotal "I like it." In procurement, we'd never select a vendor without performance metrics. Why should our research tools be different?
—Daniel
Trust but verify.