I've been testing Scholarcy's summarization and key point extraction against a batch of recent philosophy and critical theory papers. The results are inconsistent, bordering on unreliable for dense, argument-driven texts.
The main issues I'm seeing:
* It often misidentifies the core thesis, especially when presented dialectically.
* Key methodological terms or niche concepts get glossed over or omitted.
* "Findings" extraction is a mess for non-empirical work, pulling out tangential statements instead of central claims.
This feels like a model trained primarily on STEM or social science literature. For anyone using it in humanities research, what's your process?
Are you pre-processing the PDFs in a specific way? Using the browser extension vs. the library yields different results for me. Or are you just using it for initial citation and reference extraction, and ignoring the summary logic?
I'm looking for practical workflow adjustments, not promises of future updates. The tool is marketed for systematic reviews across disciplines, but the accuracy gap is significant.
-- CRM Surfer
Your CRM is lying to you.
Interesting, I've been wrestling with this too for some history papers. It really does seem to miss nuanced arguments.
> Are you pre-processing the PDFs in a specific way?
That's a good question. I haven't tried pre-processing, but I've found the browser extension on the publisher's site gives a slightly better output than uploading a raw PDF to the library. Still not great though.
Maybe it's better for just pulling out the references and bibliography first, and using that as a starting point? Have you found any other tools that handle theory-heavy text better?
CloudNewbie
The browser extension's advantage is it's often working with cleaner HTML, not the OCR nightmare some publisher PDFs become. But you're right, it's marginal.
> using that as a starting point
That's the real workaround. Treat these tools strictly as pre-processors for the bibliography and the raw text extraction. Their "analysis" layer for humanities is noise. I feed the extracted clean text into a local semantic search setup to find concept frequency and connections myself.
For theory, you're still reading. The tool just gets the text off the page. Any claim of understanding is marketing.
Trust but verify – and audit
Your diagnosis about the training bias is almost certainly correct. Most summarization models are optimized for structured abstracts and empirical findings, where claims are often explicitly signaled. Dialectical arguments in philosophy don't fit that pattern.
I've developed a workflow that treats the tool's output as raw, often flawed, data. First, I always use the browser extension on the publisher's site for the cleanest text capture, as you noted. Then, I run the same paper through the library upload; the discrepancies between the two generated summaries are themselves informative. Where they disagree on a "key point" is frequently where the model is guessing, and that's often precisely the nuanced argument I need to examine manually.
My practical adjustment is to ignore the summary and "findings" entirely. I use Scholarcy's reference extraction as a starting point for Zotero, and I use its highlighted "concepts" not as accurate tags, but as a quick map of term density. This flags where a niche methodological term like "parataxis" or "ontoepistemology" appears frequently, telling me where to focus my own reading, even if the tool's definition is wrong. The tool isn't parsing the argument, but it can crudely index the text.