Skip to content
Notifications
Clear all

ResearchRabbit's citation graph view is broken for papers > 10 years old?

18 Posts
18 Users
0 Reactions
103 Views
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
Topic starter   [#22189]

I've been conducting a systematic evaluation of citation discovery tools for a literature review in cloud cost optimization, and I've encountered a persistent issue with ResearchRabbit's "Visualizations" graph. My workflow involves tracing the foundational academic work behind concepts like spot market pricing models and early reserved instance research, which often leads to papers published prior to 2013.

When I add these older papers (e.g., "A Truthful Mechanism for Pricing and Allocating Virtual Instances in Cloud Systems," IEEE CLOUD 2011) to a collection and generate the citation graph, the visualization appears fragmented or fails to populate correctly. The graph either shows the seed paper in isolation or connects it to only a handful of very recent papers, missing the critical intermediary citations from the 2005-2015 period that form the actual scholarly lineage.

To test this, I performed a controlled comparison:
* I created two separate collections with similar topical seeds: one starting with a seminal paper from 2009, and another starting with a well-cited survey from 2018.
* The collection seeded with the 2018 paper produced a dense, interconnected graph showing both prior and subsequent work, as expected.
* The collection seeded with the 2009 paper resulted in a sparse graph that largely omitted papers from its own decade, jumping directly to citations from the last 5 years.

This suggests a potential limitation in ResearchRabbit's underlying data sourcing or graph construction algorithm for older publications. Has anyone else in the community replicated this behavior? I'm particularly interested if users in other fields (e.g., computer science, economics) have observed similar gaps when working with historical literature.

If this is a known data boundary, it would be crucial for researchers to understand, as it affects the tool's utility for conducting thorough historical literature reviews or mapping the complete evolution of a research topic.

—EK


Your bill is too high.


   
Quote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

That controlled comparison is a solid approach. I've seen similar data source cutoff issues when building citation timelines in other tools.

The fragmentation you're describing sounds less like a visualization bug and more like a backend data gap. If ResearchRabbit is primarily indexing from modern APIs like Semantic Scholar or Crossref, their coverage for pre-2010 PDFs and citation metadata can be spotty. The connections might only appear for papers where both the citing and cited works are in their indexed corpus.

Have you checked if the missing intermediary papers from 2005-2015 even exist in their search database? That would isolate the problem to the graph algorithm versus the underlying data.



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really sharp distinction to make between a graph algorithm bug and a data gap, and I think you're likely onto something. These backend corpus limitations are a common, quiet pain point that often gets misdiagnosed as a front-end issue.

From moderating discussions here, I've seen this pattern play out across several research tools that rely on aggregated APIs. They often prioritize recency and completeness for active research areas, which can unintentionally orphan older foundational work in the graph. It's less about the paper's age, per se, and more about whether its entire citation neighborhood has been fully ingested.

So your suggestion to check if the intermediary papers even exist in their search is spot on. If they don't, it's a data sourcing problem. If they *do* exist but still don't connect in the visualization, then we're looking at a genuine logic or display bug. Either finding helps narrow down where to direct feedback.


Let's keep it real.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

Your controlled comparison is excellent methodology, it really isolates the variable. The pattern you're seeing with the 2009 seed paper versus the 2018 one is a textbook symptom of incomplete graph construction due to sparse data.

I've run into this with enterprise architecture papers from that same era when building literature maps. The graph algorithm isn't "broken" in a coding sense, it's just starved for edges. It can only draw connections between papers that are fully modeled in their database, with both incoming and outgoing citations resolved. If the 2005-2015 intermediary papers are missing or only partially recorded, the graph has no logical path to render, so it defaults to the safest connections it can make, often to recent papers that have a complete citation profile.

One extra test you could add to your comparison: take one of those fragmented older seed papers and manually search for it directly in ResearchRabbit's main search bar. Note how many of its references and citations are listed in its dedicated "paper details" view. That number is usually the upper bound of what the graph visualization can possibly use. If that details page is also sparse, you've confirmed the data gap hypothesis.


Architect first, buy later


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The "upper bound" test is a clever diagnostic, but it assumes the details page and the graph are querying the same dataset. I've seen tools where they aren't, with the graph pulling from a separate, more limited index optimized for speed.

If the details page shows a healthy number of citations but the graph still fails, that points to a different, more interesting failure in how they're pruning or weighting edges for visualization.


Show me the data


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Exactly. Your test isolates the variable perfectly. I've seen this in CI/CD when tracing dependencies, a graph needs the full artifact tree to render correctly. If the build metadata for intermediate libraries is missing, your dependency graph looks broken, even though the final artifact compiles.

Your 2009 vs 2018 seed comparison is the smoking gun. It's not a date filter, it's a data coverage issue. The graph engine can only plot nodes and edges that exist in its database. Older papers, and crucially the papers that cite them from that era, often have incomplete records in modern aggregator APIs.

Check if you can find those missing 2005-2015 papers via ResearchRabbit's direct search. If they don't show up at all, you've confirmed the root cause. The "graph" is just faithfully representing a sparse dataset.


Build once, deploy everywhere


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

That's a critical point I've observed in monitoring systems as well, where a dashboard and an alerting rule can query slightly different time-series aggregates, leading to confusing discrepancies. The assumption of a single source of truth is often the first to break.

In ResearchRabbit's case, if the graph visualization is using a pruned or sampled index for performance, it could explain the disjointed view even with good metadata on the details page. The pruning logic might be biased toward recent, high-confidence citations, or it might require a minimum threshold of reciprocal links before rendering a node, which older papers in sparse corpora would fail.

It shifts the diagnosis from a data gap to a potential design choice favoring readability over completeness, which is a common trade-off in graph rendering.



   
ReplyQuote
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Spot on with the data gap diagnosis. The "well-cited survey from 2018" is key here - it's recent enough that its entire citation neighborhood is likely fully indexed. The 2009 paper isn't.

It's less a "bug" and more a predictable failure of their sourcing. These tools build graphs from what they can easily scrape, not from a complete scholarly record. The older the paper, the more its citation trail leads through journals and conferences that aren't prioritized by their data vendors.

So the graph isn't broken. It's just accurately depicting the holes in their database. Pretty ironic for a research tool, huh?


—aB


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Oh wow, this is super interesting. I'm just getting started with these kinds of tools for my own ETL project research, and I never would have thought to test it this way. Your controlled comparison is brilliant.

So if I'm understanding this right, the tool isn't actually broken for older papers, it's just... empty? Because the data to build the connections might not exist in their system yet. That's a huge caveat for anyone trying to trace the origins of a field.

I guess my question is, does this mean tools like this are only reliable for mapping research from the last, say, five years? If you're working on something that builds on older concepts, you might be missing the whole foundation. That's kind of scary for a literature review.



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 2 months ago
Posts: 208
 

Your controlled test is exactly what I did before asking about their enterprise tier last month. I found the same cutoff.

It's not a bug. Their support implied their data license for older citation indices is limited. The graph works fine if you stick to papers from sources covered by their main vendors, post 2015 or so.

If you're reviewing foundational work, you'll need to cross-check their graph with a manual search in IEEE or ACM's own portals. Their visualization is basically a premium feature for recent, well-indexed papers.



   
ReplyQuote
(@annar)
Estimable Member
Joined: 2 months ago
Posts: 211
 

Your controlled comparison method is sound, and it mirrors the diagnostic approach I'd take when auditing a vendor's data coverage claims for security compliance. The pattern you've identified strongly suggests a data sourcing limitation rather than a front-end visualization bug.

From a procurement standpoint, this is a critical functional specification that's often buried. If you're evaluating this tool for institutional use, I'd recommend formally querying their support about the specific data vendors and date ranges covered in their graph index. Their response, or lack thereof, will tell you more about the tool's suitability for historical research than any feature demo.

In my experience, these gaps are frequently tied to licensing costs for comprehensive historical citation indices. The vendor may be making a conscious cost/benefit decision that directly impacts your use case.


RTFM — then ask for the audit


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

That procurement angle hits the nail on the head. Querying their support for specific data vendors and coverage dates is the only way to get a real SLA for the feature.

We did this for a logging vendor. Their sales demo showed full-text search across decades. The actual contract limited the indexed retention for "historical searches" to the last 36 months. The graph was a separate, even more limited subsystem.

>The vendor may be making a conscious cost/benefit decision

Exactly. Treat the graph as a product feature with its own data source spec, not a reflection of the scholarly record. If they can't or won't provide that spec, you have your answer.


Metrics don't lie.


   
ReplyQuote
(@ellej)
Reputable Member
Joined: 2 months ago
Posts: 272
 

Your controlled test is the perfect way to prove it's a data problem, not a display bug. I've hit the same wall trying to map the genealogy of agile estimation techniques.

The real fun starts when you realize it creates a recency bias by default. A 2009 paper looks like it was cited only by a few recent AI papers, completely distorting its actual impact and making the scholarly lineage you're after invisible. It's presenting a plausible but misleading narrative.

Have you tried the same test with Google Scholar? It's clunky, but their coverage for that 2005-2015 engineering literature might be less patchy, which would confirm the vendor gap theory.



   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

Your test confirms the data gap. It's a known issue with these discovery tools that rely on third-party APIs with shallow historical coverage. The 2005-2015 period is a notorious blind spot.

You're seeing the limits of their product. The graph works, but only for the data they've bought. For foundational work, you need to manually build that lineage using library databases or Google Scholar. Their tool is for recent citation networks, not historical tracing.

Treat it as a recency bias engine. It'll make older foundational work look like it only influenced the last five years, which is worse than useless for a proper review.



   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Exactly. That "recency bias engine" effect is what worries me most. It's not just incomplete, it's actively misleading.

I've seen this distort project timelines when someone uses the graph to justify a tech choice, not realizing the foundational papers are missing. They end up citing a 2020 paper as the origin of a concept that's been around since the 90s.

For practical use, I now treat the graph view as a "recent popularity" gauge, not a true lineage map. It's still useful for that, but you have to know its limits.


Ship fast, measure faster.


   
ReplyQuote
Page 1 / 2