Skip to content
Notifications
Clear all

Help: Parsing fails on older scanned PDFs, any OCR settings?

3 Posts
3 Users
0 Reactions
1 Views
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
Topic starter   [#29513]

Hi everyone,

I've been seeing a recurring issue come up in a few threads and wanted to consolidate the discussion. Several users, myself included, have run into problems when using Scholarcy to parse older scanned PDFs—think pre-2000 journal articles or book chapters that are essentially just images of the page. The parsing often fails, returns garbled text, or misses entire sections.

From what I understand, Scholarcy does have some built-in OCR capabilities, but the settings aren't exactly front-and-center. I'm wondering if anyone has found a reliable workflow or specific settings tweak for these tougher documents. For instance, are you pre-processing the scans with another dedicated OCR tool (like Adobe Scan or ABBYY) before feeding them into Scholarcy? If so, what format and settings give you the best results for Scholarcy to then summarize correctly?

Also, if the Scholarcy team is listening, some clarity on the OCR engine's limits and any planned improvements would be incredibly helpful for the community. These older documents are a common pain point in academic workflows.

Let's share what's working and what isn't.


Keep it civil, keep it real.


   
Quote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Great question. I've definitely wrestled with this exact problem when trying to summarize older technical reports from my field. While I appreciate Scholarcy's built-in features, for these really stubborn scans, I've found a dedicated pre-processing step is almost mandatory for reliable results.

My current workflow is to run the PDF through ABBYY FineReader first. I export as a searchable PDF, but I make sure to enable the option to keep the original page image alongside the text layer. That format seems to play much nicer with Scholarcy's parser than a plain text export. The key setting in ABBYY, I've found, is to manually set the language dictionary to match the document - it cuts down on those garbled word issues significantly.

It's a bit of a hassle, but the accuracy jump is worth it for critical materials. I'd also love to hear if the Scholarcy team has any guidance on optimal DPI or image cleanup steps before upload.


Happy testing!


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

That's a really helpful tip about keeping the original image alongside the text layer. I hadn't considered that. My follow-up question is about ABBYY FineReader itself - it's quite an investment. Have you found the free version sufficient for this, or did you need the full paid version to get that specific export option? Trying to gauge the total cost of this workflow before I commit.



   
ReplyQuote