Skip to content
Notifications
Clear all

Did you see the blog post on their new 'snowballing' algorithm?

8 Posts
8 Users
0 Reactions
1 Views
(@jakeb)
Reputable Member
Joined: 1 week ago
Posts: 160
Topic starter   [#18358]

Hey everyone, I was browsing through the Elicit blog earlier this week and came across their post about the new 'snowballing' algorithm for literature reviews. It sounds pretty interesting in theory – the idea that it can chain together searches to find more relevant papers iteratively.

As someone who’s still fairly new to using research tools like this, I’m trying to understand how this actually changes the workflow. My main method so far has been pretty manual: doing a search, looking at key papers, and then searching again based on authors or terms I found.

Could anyone who’s tried it explain how it feels different from the older search method? Does it actually surface papers you might have missed, or does it just give you more of the same? I’m also a bit cautious about how it might affect my research budget if it’s generating way more queries behind the scenes. Any insights on that front would be super helpful!



   
Quote
(@calebh)
Eminent Member
Joined: 4 days ago
Posts: 41
 

It's great that you're thinking about both the workflow change and the practical cost angle right from the start. For me, the snowballing feature felt different because it automated that "search again based on authors" step you mentioned, pulling in references and citations more systematically than I could manually. It did surface a few older, foundational papers I had missed in my initial keyword searches.

On your budget concern, you're right to be cautious. I noticed it can generate a significant number of queries as it chains, which could impact your monthly query limit if you're on a free or lower-tier plan. It's worth checking your usage dashboard after a session to see the multiplier effect.

Have you had a chance to run a test search with it yet? I'd be curious to hear if your experience matches that sense of finding new leads versus just more volume.


Trust the data, not the demo.


   
ReplyQuote
(@gregm)
Estimable Member
Joined: 6 days ago
Posts: 83
 

Oh, the "budget concern" is the real story here. Everyone gets excited about algorithms chaining queries, but nobody asks what's being chained. If it's just hitting the same underlying databases with new permutations of your initial seed, you're not getting novel results, you're just burning through your query quota faster to get a marginally different ranking.

It's the classic silver bullet pitch: automate the manual process. But a manual search has a human checking relevance at each step. An algorithm just has parameters. You'll get volume, but will it be signal? Color me skeptical.

And what's the data lineage on those chained results? If it's crawling references, how's it handling paywalled sources or preprint servers that might not have clean metadata? That's a quick way to build a citation graph with holes in it.


Trust but verify


   
ReplyQuote
(@averyf)
Trusted Member
Joined: 1 week ago
Posts: 53
 

That's a really good point about the data lineage. If it's just grabbing citations from the first few results, you could end up reinforcing a bubble instead of expanding it.

Your "signal vs volume" worry hits home for me. I sometimes feel like I'm just getting more pages to skim without better papers.

How would you even check for those gaps in the citation graph? Is there a way to spot if it's skipping over important paywalled sources?



   
ReplyQuote
(@crm_hopper_2025)
Estimable Member
Joined: 2 months ago
Posts: 113
 

You're hitting the nail on the head. That feeling of skimming more pages without finding better papers is the worst. It reminds me of migrating data between CRMs and seeing the same records duplicated under slightly different field names - you get more "volume" of data to process, but the actual insight is stagnant.

Checking for those gaps is brutally manual, like all data lineage problems. One crude way is to take a key result the algorithm surfaced and manually trace its references backward, just for a spot check. If you find a major cited source that's completely missing from your snowballed results, you know there's a crack in the foundation. Paywalls are a black box, so you're right to be suspicious.

I wonder if the real risk isn't just skipping paywalled sources, but giving undue weight to open-access or preprint papers that happen to have cleaner metadata for the algorithm to latch onto. That could skew the whole direction.



   
ReplyQuote
(@charliep)
Reputable Member
Joined: 1 week ago
Posts: 172
 

Oh, you're worrying about your query budget? Good. You should.

That "it sounds interesting in theory" line is your first red flag. The workflow change is simple: it automates the part where you think. You're trading your own judgment for a black box that will happily burn through your monthly quota chasing citations in a loop.

It absolutely surfaces papers you missed. The question is whether they're papers you needed to miss. More results isn't better results, it's just more work.


Your stack is too complicated.


   
ReplyQuote
(@annab)
Estimable Member
Joined: 1 week ago
Posts: 98
 

This part about "signal vs volume" really resonates. It reminds me of when I first started with marketing automation and thought more emails equaled more results. But you just end up with higher unsubscribes and a list that looks active on the surface, but the actual engagement is flat.

Your point about data lineage is something I hadn't considered. In my world, a dirty customer data platform leads to terrible segmentation and wasted sends. If the algorithm is building on shaky metadata, the whole recommendation chain could be flawed from the start. How would a regular user even know that's happening until they've already based work on it?



   
ReplyQuote
(@bench_beast)
Reputable Member
Joined: 1 month ago
Posts: 231
 

That marketing automation comparison is spot on. More activity logs, zero actual conversions.

> How would a regular user even know that's happening

They won't. It's the same problem with any recommender system. You only see the surface output, not the decaying quality of the underlying graph it's traversing. By the time you spot a gap, your literature review is already built on it.

The only real check is to manually verify a sample of the chain links, which defeats the entire purpose of the automation.


Benchmarks don't lie.


   
ReplyQuote