I'm planning a project to automate parts of systematic mapping for my literature review. The goal is to extract key data from PDFs (like methods, sample size, outcomes) and organize it into a structured database.
I've narrowed it down to using either Scholarcy's API or Semantic Scholar's API. For those who have tried both, which would be better for this kind of automated extraction? I'm especially curious about the quality of the metadata and the ability to pull specific, detailed fields from academic papers reliably. My stack is Python, and I'll be handling a few hundred papers.
I'm a security engineer at a mid-sized biotech firm, and we run a custom literature mining pipeline for threat intelligence and compliance mapping, processing several thousand PDFs a month. Our stack is mostly Python with some Go services.
1. **Extraction Depth vs. Extraction Breadth:** Scholarcy's API is built to pull specific, detailed fields from individual PDFs - methods, sample sizes, outcomes - which is exactly what you listed. Semantic Scholar's API provides richer, cross-paper metadata (citations, influential citations, TLDR summaries) but its entity extraction from an arbitrary PDF's full text is less granular. For systematic mapping where you need a structured database from the paper's content, Scholarcy is the dedicated tool.
2. **Pricing and Scale Reality:** Scholarcy's API is priced per document, around $0.10 to $0.25 per PDF for the extraction you need, which for a few hundred papers is manageable. Semantic Scholar's core API is free with rate limits (100 requests every 5 minutes). The hidden cost is that to get the deep, reliable data extraction you want from Semantic Scholar, you'd likely need to build and maintain your own NLP layer on top of their raw text or summary endpoints, which is not trivial.
3. **Integration and Reliability:** Scholarcy's API expects you to upload a PDF and returns a structured JSON blob with the sections you need. It works consistently if the PDF text is selectable. Semantic Scholar's API is primarily query-based: you give it a DOI, arXiv ID, or title, and it returns its pre-computed metadata for that paper. If your PDF isn't in their corpus (like a preprint or a lesser-known journal), you get nothing, and you can't force-process your PDF through their deep extractors.
4. **Where Each Breaks:** Scholarcy can struggle with older PDFs that are scanned images or have complex layouts, though their markup parsing is decent. Semantic Scholar breaks when you need data from a paper they haven't ingested, which for a systematic review could be a significant gap. Their "paper" endpoint might not have the specific field like "sample size" neatly extracted and waiting for you.
I'd recommend Scholarcy's API for your specific case of pulling defined fields from a known set of a few hundred PDFs into a structured database. If your project instead required analyzing citation networks or trends across a vast corpus where most papers are from major journals, Semantic Scholar's free API would be a starting point. To make the call clean, tell us the source of your PDFs (major publishers vs. gray literature) and whether you have the bandwidth to build additional text processing if the API's output is incomplete.
For your specific goal of pulling structured data like methods and sample sizes directly from PDFs, Scholarcy is the clear choice. Semantic Scholar is great for metadata about the paper itself, but Scholarcy's entire purpose is to tear apart the PDF's contents into those discrete fields you need.
Be ready to write some cleaning logic, though. Their extractions are good, not perfect. For a few hundred papers, it'll be fine. Just don't expect 100% accuracy on the first run.