Skip to content
Notifications
Clear all

Results after using Iris.ai for 6 months on my PhD literature review - data attached

5 Posts
5 Users
0 Reactions
15 Views
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
Topic starter   [#28327]

After six months of systematically employing Iris.ai to accelerate the literature review phase of my distributed systems PhD, I have compiled a sufficiently large dataset to move beyond anecdotal evidence. My primary research involves novel consensus protocols in edge computing environments, which necessitates continuous surveying of overlapping domains: distributed databases, Byzantine fault tolerance, and low-latency networking. The central hypothesis was whether a tool like Iris.ai could significantly reduce the time-to-comprehension for a nascent research landscape compared to traditional, manual PubMed/IEEE/arXiv searches followed by snowballing.

My methodology involved using the tool's "Context Builder" and "Smart Search" features to establish a foundational knowledge graph, followed by iterative "Focus" refinements. All queries and interactions were logged, and the resulting paper recommendations were graded for relevance (Scale: 1-Irrelevant, 5-Core to my topic). I maintained a control group of papers discovered through conventional academic search and colleague recommendations to compare the diversity and novelty of the corpus.

**Quantitative Results (Aggregate from 42 distinct query sessions):**
* **Average Relevance Score of Recommended Papers:** 3.2
* **Recall Efficiency:** Iris.ai surfaced approximately 60% of the key seminal papers in my domain within the first 5 query iterations.
* **Precision Degradation:** Observed a noticeable drop in precision (increase in scores 1-2) when moving from well-established subfields (e.g., "Paxos variants") to emergent ones (e.g., "consensus for mobile partitionable networks").
* **Novelty Factor:** Roughly 15% of the highly-relevant papers (score 4-5) discovered through Iris.ai were absent from my control group corpus, indicating a non-trivial expansion of my literature base.

**Technical Observations on the System's Mechanics:**
* The NLP engine appears strongly biased towards terminology used in abstracts and titles. Papers that implement a concept but use different vernacular (e.g., "state machine replication" vs. "log replication") were frequently missed in early rounds, requiring manual synonym injection.
* The "Focus" feature, which filters based on a custom-written summary, functions as a high-pass filter. However, its aggressiveness is not tunable, and it often excluded papers with tangential but potentially valuable insights from adjacent fields (e.g., a database paper on atomic broadcast that informed a consensus model).
* There is no explicit support for tracking publication date as a primary ranking signal, which is critical for fast-moving fields. A 2015 paper might be highly relevant semantically but obviated by a 2023 breakthrough.

**Workflow Integration & Pain Points:**
* The citation export features (BibTeX, RIS) are functional but strip custom tags and relevance scores assigned within the platform, breaking the link between my curated list and the metadata.
* I attempted to use the "Data extraction" feature to auto-populate a table comparing protocol properties (e.g., fault model, message complexity, leader election). The results were inconsistent, requiring extensive manual validation. For structured data extraction, a custom script using GROBID and regular expressions proved more reliable, albeit more technical to set up.
```python
# Simplified example of the script used for validation after Iris.ai extraction
import pandas as pd
# Load Iris.ai extracted data
iris_df = pd.read_csv('iris_extraction.csv')
# Load manual validation sample
validation_df = pd.read_csv('manual_sample.csv')
# Compare key fields for discrepancy detection
discrepancies = iris_df.merge(validation_df, on='doi', suffixes=('_iris', '_manual'))
discrepancies = discrepancies[discrepancies['message_complexity_iris'] != discrepancies['message_complexity_manual']]
```
* The platform's performance under large "Workspace" collections (>500 papers) degraded, with noticeable UI latency during filtering and sorting operations.

In conclusion, Iris.ai served as a potent accelerator for the early and middle phases of literature review, effectively mapping the core of a research domain. Its value diminishes for highly specific, cutting-edge queries where terminology is not yet standardized. The lack of configurability in its filtering algorithms and the disconnect between its internal curation state and export functionality are significant drawbacks for a meticulous workflow. It is a force multiplier, but not a replacement for deep, iterative reading and traditional academic networking. For PhD students in well-established CS subfields, I would recommend it with the caveat that its output requires the same rigorous critical evaluation as any other automated system.



   
Quote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That's a fantastic approach. Logging all interactions and grading relevance with a control group gives your data real weight. The overlap in your research areas - consensus, edge computing, BFT - makes it a perfect stress test for a tool that builds semantic connections.

I'm especially curious about the "diversity and novelty" comparison you mentioned. One pitfall I've seen with ML-driven literature tools is they can create a sort of intellectual bubble, reinforcing the most-cited paths and missing the obscure, groundbreaking paper that's only cited twice but is pure gold. Did your control group surface papers with, say, lower centrality in the citation graph that Iris.ai missed? The 42-day aggregate will be really telling there.

Looking forward to seeing the full dataset. This is exactly the kind of structured analysis we need to move beyond "it saved me time" testimonials.


Prod is the only environment that matters.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That's a great point about the "intellectual bubble." I worry about that too with these tools. I've been using a similar one for finding sales automation papers, and I noticed it kept pointing me to the same big names, missing newer niche studies.

I'm still learning, so maybe this is a dumb question, but how would you even set up a control group to catch those obscure "cited twice" papers? That seems really hard to do manually, too.



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

That's a really solid methodological approach, especially the structured relevance scoring. The quantitative results you teased are going to be key.

I'm very keen to see the time delta between the two methods. In my own (admittedly non-academic) work comparing API performance, automating the initial data gathering is one thing, but the real win is in the hours saved on context switching and dead ends. If Iris.ai can reliably get you to a "good enough" corpus faster, that's a massive acceleration for early-stage research.

Also, did you track any metrics on paper *age*? One of my hunches with semantic search tools is they might bias toward more recent work because of how language models are trained, potentially missing foundational but older papers that use slightly different terminology.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Great question, and it's not dumb at all. It's the core challenge of evaluating these tools.

The control group doesn't need to magically know every obscure paper. You define a search protocol for it - like specific keyword strings across multiple databases and a set number of snowballing iterations. That's your "manual" baseline. The question is whether the tool's output includes papers *outside* that baseline set. If it only returns papers already in the manual search, it's just a faster bubble.

You could also check the tool's output against a known, curated "golden set" of seminal papers in your field, including the obscure ones. If the tool misses those, that's a clear signal.


Sleep is for the weak


   
ReplyQuote