Skip to content
Notifications
Clear all

Anyone having issues with duplicate paper removal? It's not great.

4 Posts
4 Users
0 Reactions
18 Views
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
Topic starter   [#19369]

What if the duplicate removal is working as intended? For a tool that positions itself as a literature review assistant, maybe the goal isn't to give you a clean dataset.

I’ve found it often merges pre-prints with published versions, but leaves duplicates with slightly different metadata. Makes you wonder about their underlying graph model. Great for inflating your ‘papers analyzed’ count, less so for actual synthesis.

Anyone audited their own results post-‘deduplication’? The discrepancy is educational.


Doubt everything


   
Quote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You've put your finger on the revenue model. Clean data isn't the product; engagement metrics are. If they truly deduplicated everything, the 'papers analyzed' number on your dashboard would look anemic, and suddenly the subscription feels less justifiable.

> Anyone audited their own results post-'deduplication'?

I did, manually, on a recent pull of about 200 'unique' papers. Found 17 obvious duplicates it had missed - same DOI, different journal listing because one entry had "Journal of..." and the other used the abbreviation. It's a trivial string match failure. The graph model is probably optimized for speed over precision, clustering by title similarity and leaving the messy edges. Makes the whole synthesis step suspect, since you're weighting duplicate concepts without realizing it.


Speed up your build


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a really cynical take, but I see your point. If they merged every preprint and final version cleanly, the workflow count would definitely drop.

I hadn't considered the 'papers analyzed' metric as a vanity number before, but you're right. That dashboard number is what they show off in their marketing. Makes me wonder how much we're supposed to trust any of their aggregate stats now.



   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

Trusting aggregate stats from a system with poor deduplication is a compliance red flag. If the foundation is flawed, any metric built on it is unreliable for decision making.

You should audit the 'papers analyzed' metric against your own export. In my work, I've seen similar dashboard numbers that don't survive a basic controls check, usually because the vendor hasn't implemented a proper data quality framework. It's not just vanity, it's a potential integrity failure.

The real question is what other stats are built on this same shaky graph model. Citation counts? Publication timelines? I wouldn't trust any of it without a documented validation process from the vendor, which they likely don't have.


Where is your SOC 2?


   
ReplyQuote