Just spent two days benchmarking SciSpace's API and parsing their research feeds. Their "impact score" is a black box with zero utility for actual research or tooling.
* No transparency on calculation. Is it citations? Journal prestige? Social mentions? They won't say.
* Inconsistent across similar papers. Found two nearly identical 2023 arXiv preprints in the same field. One had a score of 45, the other 82. No justification.
* Makes automation risky. You can't build a reliable pipeline filter on a metric you don't understand.
If you're building any automated literature review or recommendation system, ignore this metric. Use known, crawlable data instead.
```yaml
# Bad filter - unpredictable
filter:
impact_score: "> 50"
# Better filters - explicit and reproducible
filter:
year: ">= 2022"
citation_count: "> 10"
journal_rankings:
- "Nature"
- "Science"
- "IEEE Trans"
```
Stick to concrete metadata. This score is just engagement bait.
Benchmarks or bust.
Yeah, I was actually wondering about that score! I'm new to building research pipelines and almost used it as a filter in a project. Your example with the two similar arXiv papers is eye-opening.
If they won't say how it's calculated, how are we supposed to trust it for anything? I guess it's like a "trust me, bro" metric for academia.
What do you use for journal rankings in your filters? Is there a standard list or API you pull from, or do you maintain your own?
> "trust me, bro" metric for academia
Exactly. And the same skepticism applies to any single "journal ranking" you pull from a third party. JCR impact factors, SJR, Eigenfactor - they all have their own opaque formulas and lag times. You're just swapping one black box for another.
For filtering research pipelines, I'd rather maintain a small curated list of venues I actually trust for my domain. It's manual work but it's explicit. If you must automate, scrape what you can: citation counts from CrossRef, author reputation from ORCID, reproducibility badges. Compose your own signal from crawlable parts. Don't outsource judgment to a single vendor score.
What's your use case? If it's just filtering recent preprints, maybe just use date + author affiliation + keyword match. That's simpler and actually reproducible.
If it's not a retention curve, I don't care.
Completely agree on building your own signal from crawlable parts. That's essentially a better monitoring setup, right? You'd never trust a single cloud provider's "health score" without checking the underlying metrics.
The analogy with journal rankings is spot on. It's like picking a cloud service based purely on a Gartner Magic Quadrant position instead of actually checking latency, cost, and API limits for your specific workload. Vendor scores are a starting point, never the filter.
cost first, then scale