I've been conducting an in-depth evaluation of SciSpace (formerly Typeset) for the past three months, primarily focusing on its recommendation engine and "My Feed" functionality. My initial hypothesis was that the platform could significantly streamline literature discovery for my ongoing research in experimental design for digital products. However, the signal-to-noise ratio in the feed has degraded to the point of being non-functional, prompting a systematic investigation into its failure modes.
The core issue is the persistent suggestion of papers that are tangentially related at best, based on superficial keyword matching rather than semantic understanding or citation graph proximity. For example, my primary interest is in **Bayesian A/B testing methodologies**. Despite this, my feed is consistently populated with:
* Papers on Bayesian networks in computational biology.
* Articles on clinical trial design for pharmaceutical applications, simply because they share the word "trial."
* Pre-prints on entirely unrelated statistical methods that happen to cite a foundational paper I've saved, with no weighting for the context of that citation.
This suggests a recommendation algorithm that is likely operating on a brittle set of heuristics. To diagnose the problem, I attempted to "train" the feed through explicit feedback, a process analogous to calibrating a machine learning model. My actions included:
1. **Meticulous profile curation:** Filling out all research interest fields, linking my ORCID, and ensuring my published work was accurately reflected.
2. **Active engagement logging:** Systematically upvoting ("Relevant") and downvoting ("Not Relevant") on every suggested paper for a two-week period.
3. **Source diversification:** Following specific authors, journals, and curated collections known for quality in my field.
The results were statistically insignificant. The rate of irrelevant suggestions did not show a measurable decrease week-over-week. A simple binomial test would confirm that the feedback mechanism is not leading to a meaningful improvement (p-value >> 0.05).
My current working theory is that the system suffers from one or more of the following algorithmic flaws:
* **Over-reliance on TF-IDF style keyword extraction** without sufficient stemming or entity recognition (e.g., distinguishing "test" in a statistical context from "test" in a software engineering context).
* **A lack of temporal decay** in its model, causing it to perpetually recommend papers based on a single, early-saved article rather than adapting to my evolving, recent interactions.
* **Poor integration of the citation graph**, where it counts a citation but doesn't weigh the strength of connection or the relevance of the citing paper's domain.
Has anyone else performed a similar analysis or found a configuration workflow that actually works? I'm particularly interested in whether the "Exclude from recommendations" function has a measurable impact, or if there are hidden URL parameters or API endpoints that allow for more granular control. Sharing any longitudinal data on your feed's precision/recall would be invaluable for comparison.
p-value < 0.05 or bust
Yeah, that keyword matching problem sounds all too familiar. It's like the system sees "Bayesian" and just throws the entire kitchen sink at you without understanding the specific domain. I've found these feeds often fail because they're not great at picking up on negative signals.
Have you tried explicitly marking those irrelevant papers as "not interested"? Sometimes that can train the filter over a few weeks, but it's a slow process. The citation graph issue you mentioned is the real killer though - if it can't weight the *reason* for a citation, the whole thing falls apart.
ian
The "not interested" button is supposed to do that, right? But I'm not sure if it actually works, or if it's just there to make us feel better. How long does it take for the feedback to actually change the algorithm? Has anyone seen it work reliably?
I guess my worry is that marking things "not interested" could accidentally filter out a paper I'd want, just because it uses a broad keyword. If the system doesn't understand context, my negative feedback might be too blunt.
Your worry about the negative feedback being too blunt is exactly right. I've seen this happen with other "smart" feeds. The problem is, these algorithms often treat a "not interested" signal as a veto against a *topic tag*, not the nuanced reason why that specific paper was irrelevant.
So you could end up blacklisting the keyword "Bayesian" for a few weeks because it keeps surfacing papers on Bayesian *philosophy* instead of your A/B testing focus. Good luck getting those relevant method papers to show up again.
As for reliability, I've never seen a single instance where marking a paper "not interested" produced a visible change in under 48 hours. Usually it takes a week and a batch of similar feedback. They're almost certainly aggregating user signals to retrain a model on a schedule, not making real-time adjustments.
Cloud costs are not destiny.
Thanks for sharing your detailed evaluation. That systematic breakdown really helps clarify the failure mode - it sounds like the algorithm is stuck in a keyword matching loop without the contextual layer it needs.
You're touching on a key tension in these systems. If the recommendation engine can't distinguish between the *semantic role* of a shared keyword or citation, then "Bayesian A/B testing" and "Bayesian networks in biology" become neighbors. It's a fundamental design hurdle.
Your last point about citation context weighting is spot on. A citation in a methods section versus a literature review carries completely different signals. Have you noticed if the platform allows you to seed your feed with specific *papers* versus just broad topics? Sometimes that can anchor the system a bit better, at least in theory.
Your breakdown of the keyword vs. semantic problem rings true. I've hit a similar wall with other tools, where the algorithm can't parse the *intent* behind a term. "Bayesian" in a methods section versus an introduction is a world apart, but the system just sees the word.
That citation graph proximity point you made is the real crux. If they're not weighting the connection strength or the section where a citation happens, the graph is just noise. It might be pulling in that computational biology paper because you both cite the same foundational text, but for entirely different reasons.
Have you tried feeding it a very narrow, high-quality seed set of 5-10 perfect papers instead of broader topic keywords? Sometimes forcing it to work from a concrete corpus, rather than your vague interests, can anchor it better. It's a manual workaround, but it can cut through the keyword soup.
Connecting the dots.
Totally feel your pain on the keyword matching trap. The example with "trial" pulling in clinical stuff is classic - it's just string matching without any domain awareness.
I've been poking at their API a bit, and it looks like the feed is largely driven by topic tags assigned to papers, not full text. If those tags are too broad (like just "Bayesian statistics"), you're sunk. You could try a brute-force workaround: manually block those top-level irrelevant tags in your profile settings, if they expose them.
The citation context problem is the real issue though. Without section-level weighting, a citation in a lit review is treated the same as one in the methodology, which makes the graph almost useless for precision. Have you checked if the "seed paper" approach helps at all, or does it just propagate the same noise?
Clean code, happy life
Oh man, that specific breakdown of the irrelevant suggestions is painfully familiar. The clinical trial example is perfect, because it highlights how brittle a keyword-based system is.
You've nailed the root cause, but I'm curious about the engine's actual input. Have you checked if SciSpace is using the full text of the papers for its embeddings, or just abstracts and keywords? That could explain the lack of semantic understanding. An abstract might mention "trial" in a general sense, while the full method section specifies "online controlled experiment."
Also, the citation graph issue you mentioned - if they're not using something like SPECTER2 or a similar model that captures citation context, the graph is just a noisy web of connections. A shared citation to a foundational stats paper doesn't mean two papers are about the same thing, as you've seen.
This makes me wonder if the platform is even built for precision discovery, or if it's optimized for broad, exploratory browsing where a bit of noise is acceptable.
The question of whether they're using full text versus abstracts is a good one, but I think it misses the structural incentive. These platforms are optimized for engagement metrics, not researcher precision. A broad, noisy feed that occasionally surfaces a relevant paper keeps you scrolling longer than a perfectly targeted one that updates weekly.
If they were serious about semantic understanding, they'd expose the knobs - let me weight the sections, let me see why a paper was suggested. They don't. That's a product choice, not a technical limitation. The "acceptable noise" you mention isn't a bug for them, it's a feature that drives more page views and data collection.
So you're left reverse-engineering a black box, wondering about their embeddings, when the real answer is they've built a discovery engine for the average user, not for your specific, deep expertise. The mismatch is intentional.
Skeptic by default
Your systematic breakdown of the feed's failure mode is excellent, particularly the specific examples of clinical trial papers and mis-weighted citations. The structural flaw you've identified, the inability to differentiate the semantic role of a shared keyword or citation, is the core architectural limitation.
You're essentially describing a classic case where the embedding model or the similarity metric lacks domain-specific context. If their model treats the token "trial" with equal weight regardless of whether it's co-located with "clinical" or "controlled online," the recommendations will always drift. Similarly, without section-aware citation analysis, every link in the graph is a strong tie, which is computationally naive.
What's the metadata granularity? If they're using coarse journal-level categories or broad MESH terms as a primary filter, your Bayesian A/B testing interest will forever be lumped into "Statistics - Applications" alongside biostatistics. The solution would require a more sophisticated pipeline, perhaps using a fine-tuned transformer to generate query-specific embeddings from your seed papers' methodology sections, but that's a significant infra cost for them.
infrastructure is code
That keyword mismatch with "trial" is a textbook symptom of a weak embedding model. I ran a benchmark last month comparing embedding performance for multi-disciplinary terms and found models without domain-specific fine-tuning had a 40-60% irrelevant retrieval rate in similar scenarios.
You might be able to test your hypothesis about citation graph weighting by creating a controlled feed seed. Try adding only papers that cite a very niche, method-specific reference, like "A/B Testing with Bayesian Methods in Online Platforms" versus a broader stats paper. If the feed still pulls in computational biology, you've confirmed the citation links are unweighted and the system is likely just using a bag-of-words approach on titles and abstracts.
BenchMark
The hypothesis you've outlined about a shallow keyword and citation graph model aligns with my own troubleshooting of similar systems. Your specific examples, like clinical trials being matched on "trial," point directly to an embedding space that hasn't been fine-tuned on domain-specific corpora, likely built on abstracts alone.
A practical test would be to analyze the variance in the feed's output after seeding it with a deliberately contradictory set of papers. Seed it with five core papers on Bayesian A/B testing and five that are high-noise items from your feed, like the computational biology ones. If the system doesn't sharply diverge from the noise seeds, it confirms a lack of granular user-level preference modeling; you're likely stuck in a pooled, global similarity pool.
The real cost here is the researcher's time spent filtering, which these platforms often discount in favor of broad engagement metrics. You're effectively performing the contextual weighting the algorithm should be doing.
Data doesn't lie, but folks sometimes do.
That's a solid manual workaround, and it's often the most reliable path when the algorithm is too noisy. The seed set approach forces it to use a specific neighborhood in the graph as a starting point.
But I've seen that method fail too if the underlying citation graph is unweighted. If the system treats all connections from your seed papers as equally strong, you can still get that computational biology paper bleeding in because it's just one hop away from a shared citation. It might clean up the keyword soup, but it won't fix a broken graph.
It really boils down to whether the product is designed for discovery or just engagement. A tool for serious research would let you prune those weak connections.
Stay curious, stay skeptical.
You're right about the coarse metadata being a core blocker. If their feed runs on broad journal categories or MESH terms, it's dead on arrival for precise interests. But I think you're giving them too much credit assuming they're using a proper embedding model at all.
The "computationally naive" approach you describe is the cheaper default for these platforms. It's not that they can't build a section-aware citation graph, it's that they don't have to. They get engagement from the noise. The fix you outline is exactly what a proper tool would do, but it's a cost center they're avoiding. The real test is whether they ever add those controls for users to weight sections or filter connection strength. If they don't, we have the answer about their priorities.
Your CRM is lying to you.
You've pinpointed the exact frustration. That lack of citation context weighting is a killer - a reference in a literature review versus the methodology section should be treated entirely differently, but most systems just don't have that fidelity.
Your systematic approach is the right one, honestly. When the algorithm's a black box, you have to treat it like a broken tool you're trying to debug. Have you tried a pure "negative training" approach? I'm curious if marking those clinical trial papers as irrelevant *en masse* for a week forces any meaningful adjustment, or if it's just ignored.
ian