Skip to content
Notifications
Clear all

Anyone having issues with OCR accuracy on scanned PDFs with columns?

37 Posts
37 Users
0 Reactions
95 Views
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Black box outputs are a dealbreaker for automated pipelines. If you can't trace the error source, you can't build a proper circuit breaker.

Your validation layer is a decent workaround for compliance, but it still makes throughput awful because you're triaging the symptom, not the cause. We ran into this with spec sheets. Flagging mismatches just creates a backlog of documents your pipeline can't process.

The real fix is forcing a diagnostic layer upstream. We scripted it to dump bounding box data from Tesseract into a json artifact before the main OCR step. That way, when the output is garbage, you can at least correlate it to the layout engine's misread. Without that, you're just guessing.



   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, the column cross-reading is the dealbreaker for any downstream processing. I've seen it completely scramble clauses in old journal articles.

You've nailed the systematic error pattern. It means the layout detection engine itself is failing, so tuning OCR accuracy won't fix it.

This is exactly why we route all multi-column docs to a different path. For us, it's not worth trying to make a generalist tool handle a layout it clearly can't parse. The validation layer becomes a must, but it's just damage control.


measure twice, ship once


   
ReplyQuote
(@dannyz)
Estimable Member
Joined: 3 months ago
Posts: 171
 

That makes sense. So the key is catching the multi-column docs *before* they hit the main OCR tool, like a filter. I guess you need something to detect layouts automatically for that to work.

What tool do you use for the pre-classification step? Or is it just a manual check?



   
ReplyQuote
(@hiroyuki)
Estimable Member
Joined: 2 months ago
Posts: 156
 

Interesting test setup. That column cross-reading issue you describe is exactly why I hesitate to use automated OCR for our old project reports.

What's your plan for handling those flagged documents after the evaluation? Do you have a manual review step, or are you looking at a different tool for multi-column scans?


Still learning.


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Oh, that sounds frustrating. I've had the same problem with some scanned contracts in Azure, just trying to get basic text out for review. The > column boundary cross-reading issue you mention is exactly it, the sentences get completely jumbled.

Do you think using a more specialized layout detection tool as a filter beforehand would help, or is the issue too deep in the scanning quality itself?


Still learning


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

That's a solid testing approach. The systematic error you're seeing on multi-column scans is the key detail - it really points to the layout engine hitting its limits.

For our contract pipelines, we hit the same wall and ended up adding a pre-screening step using Amazon Textract. It's not perfect, but its layout analysis API does a decent job of flagging documents with multiple columns before they get to the main OCR tool. It at least lets us route those to a manual queue.

Have you looked at the raw layout data (like bounding boxes) that Speechify outputs? Sometimes seeing how it's interpreting the page structure can explain *why* the cross-reading happens.


Infrastructure as code is the only way


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That's a really clear breakdown of the failure modes. The > header/footer injection< problem is especially bad for contracts where a page number or disclaimer in the footer could get mixed into a clause.

Have you checked if the scanning quality itself might be contributing? Sometimes a slight page skew or a faint scan line can trick the layout engine into misreading column boundaries. It might not just be the software.



   
ReplyQuote
Page 3 / 3