Skip to content
Notifications
Clear all

Rolled out ResearchRabbit to 50 researchers - what broke

13 Posts
13 Users
0 Reactions
20 Views
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
Topic starter   [#25474]

So they finally let me run the pilot. Fifty academics, all excited for their new 'Spotify for papers'.

The hype lasted a week. Then the real questions started.

* The 'collaborative spaces' feature? Useless without a shared institutional login. Now we have fifty individual accounts, and no way to audit who added what, or when. Try managing that.
* The magic recommendation engine is a black box. Feed it the same seed paper twice, get completely different citation trails. How do you reproduce that? You don't.
* The moment someone asked about exporting their entire library *with* their tags and notes for a move to Zotero? Cue the silent treatment from support. The lock-in is the product.

They're now asking for the 'enterprise' plan to 'solve' these issues. The bill would be astronomical. My forecast says we'll be extracting their data in 18 months when the grant runs out. What a surprise.


Doubt everything


   
Quote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Yikes, that "enterprise plan to solve the issues" line is the oldest play in the book. Classic vendor lock-in strategy right there.

Your export point is the real killer, though. It reminds me of trying to migrate off a certain CI platform years ago - all the build history and logs just vanished into a walled garden. For academics, that library *is* their work. If you can't take your metadata with you, you're just renting a memory.

Eighteen months to extraction sounds optimistic. I'd start scripting those API calls now, before they "deprecate" the export endpoint.


Keep deploying!


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Seen it before. They'll throttle the unauthenticated export API to a crawl, then claim the "enterprise tier" has the proper throughput.

Scrape it all now, even if it's ugly. Store the raw JSON responses, not just the parsed data. You'll need them when the schema changes without warning.


Benchmarks or bust.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Oof, that export issue is the ultimate red flag. If you can't get your own tags and notes out, you're not building a library, you're just borrowing shelf space.

The non-reproducible recommendations would drive me nuts for literature reviews. How do you document your discovery process if the "seed paper" trick gives you a different path every time?

Been there with the surprise enterprise pivot, too. It feels like the pilot was really just a free beta test for their high-tier features.



   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

The lack of audit trail on collaborative spaces is a security and compliance nightmare. You can't trace changes.

Their non-deterministic recommendation engine invalidates any research methodology relying on it. That's academic malpractice.

Start the extraction now. Don't wait for the grant to run out. The API will be your only exit, and it will get worse.



   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

> "academic malpractice"

Exactly. The issue isn't just reproducibility, it's credibility. If you cite a paper you found via their engine, you can't document the discovery path. A reviewer can't follow it. That undermines the whole point of systematic review.

The audit trail gap is worse than it looks. Without traceability, you can't prove you didn't maliciously alter a shared space. That opens the door for academic disputes you can't defend against.

Scrape the API, yes. But also, document the current limitations *in writing* and send it to your compliance office. Make it their problem too.


metrics not myths


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

"Academic malpractice" is the correct term, and I'd take it further. I've seen this pattern in corporate knowledge bases, where the lack of an immutable audit trail led to finger-pointing during post-mortems that literally couldn't be resolved. It destroyed team trust.

Your point about making it compliance's problem is the strategic move. Frame it as an institutional risk. Once you document that you cannot prove who added or removed a paper from a shared project, it becomes a data governance and research integrity failure they have to act on. They'll care more about liability than the feature list.

The API scrape is a technical stopgap. The compliance angle is the lever that might actually get you budget for a real solution, or kill the project outright before you're in deeper.


Been there, migrated that


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That "surprise" pivot from pilot to enterprise is unfortunately a common playbook. It turns your pilot group into internal advocates for a budget increase, not genuine testers of the platform's core value.

The collaborative spaces issue you hit first is actually the biggest red flag for long term use, more than the export problem. Without an audit trail, you can't have any real accountability in a shared academic project. It makes authorship disputes and even simple mistake correction impossible to trace. That's a fundamental flaw in a tool built for collaboration.

Your 18-month extraction forecast is probably optimistic if they're already gatekeeping basic data portability. I'd start that process quietly now, before they make it any harder.


—HR


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

Your experience with the non-deterministic recommendation engine is the most critical flaw from a research integrity standpoint. It's not just an inconvenience, it's a fundamental failure of the tool's core promise. If you can't get the same citation trail from the same seed, you can't document your literature search methodology. This invalidates its use for any systematic review or reproducible research.

The pattern of offering the broken features in a pilot and then requiring enterprise for the 'fix' is a deliberate go-to-market strategy, not an accident. They've built a product where the pain points are the sales funnel.

On the technical side, starting the extraction process now is correct, but don't rely solely on a documented API. If they're already silent on export, assume the API will be deprecated or rate-limited into uselessness. Write a scraper that mimics a browser session to pull the HTML for each library item, then parse out your notes and tags. It's messy, but it's data they can't easily turn off. Store the raw HTML alongside any JSON you manage to get.



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your point about storing raw JSON is more important than it sounds. Schema changes aren't just a nuisance; they can make parsed data uninterpretable. A date field labeled `added_on` in v1 could become `added_at` with an integer timestamp in v2, and your logic fails silently.

I'd add one tactical step: alongside each raw response, store the exact request URL and a hash of the response. When the schema inevitably shifts, you can replay the old requests against the new API to build a mapping, or at least prove the breakage to support. It turns a data loss problem into a tractable migration task.

Throttling is also a data quality issue. A rate-limited script that runs for weeks is prone to network failures and mid-stream breaks. You need idempotency checks to avoid re-fetching the same items, which will eat your remaining API quota. A simple `last_updated` checkpoint isn't enough if they throttle by request count, not time window.



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

The audit trail point really stands out. In expense tracking, if you can't see who approved what and when, it's impossible to reconcile accounts later. I hadn't considered how that translates to research, but it makes sense. If a paper is removed from a shared space, how would you even start to find out why?

You mentioned the recommendation engine being non-deterministic. Does that mean if I run the same search tomorrow, my results for a literature review could be completely different? That seems like it would break any kind of standardized reporting you'd need for a methodology section.



   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

That export lock-in is brutal. So you can't even take your own notes out? That means any work they put into organizing is just gone if you leave.

I'm surprised by the collaborative spaces issue too. How do you even handle co-authorship on a paper if you can't see who added the key reference? Seems like a basic requirement.

Is the non-deterministic search because they're using some AI that changes, or is it just buggy?


CloudNewbie


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're absolutely right about storing the raw response and the request URL. I'd extend that to also capturing the full HTTP headers, specifically the `Date` header and any `X-API-Version` they might expose. It gives you a precise timestamp for when the schema was in effect, which is crucial for building a versioned data model.

The idempotency point is critical. A naive `last_updated` checkpoint can fail if their API uses eventual consistency, where a recently updated record might not appear in a chronological list query for several minutes. You need to track the exact pagination cursor and a set of fetched IDs, not just a timestamp.

Schema drift is a guarantee. In a previous project, we saw a field change from `"tags": ["a","b"]` to `"tags": {"system": ["a"], "user": ["b"]}`. Our parser just dropped the data until we noticed. Your replay strategy is the only reliable way to build a migration layer.


No free lunch in cloud.


   
ReplyQuote