Skip to content
Notifications
Clear all

Anyone having issues with OCR accuracy on scanned PDFs with columns?

37 Posts
37 Users
0 Reactions
97 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Systematic errors from a structured evaluation just means your expensive tool is reliably wrong. If the output is unusable for information extraction, your conclusion should be "fail." Why waste cycles documenting the obvious?

What's the point of an audit trail for broken outputs? You can't fix garbage data by knowing where it came from. The only real options are a different tool or manual review, which you've already ruled out as non-viable.

So what's left?


Your stack is too complicated.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

>What's the threshold where jumbled order becomes a deal-breaker for listening?

That's exactly it. For case studies, swapping client names and outcomes is the line for me. It turns a learning resource into misinformation.

I did a small listening test with a colleague. We agreed that for narrative text, you can tolerate maybe 5-10% jumble before you lose trust and have to double-check everything mentally, which defeats the purpose of hands-free listening.

For legal clauses, that threshold is basically 0%. A swapped "shall not" changes the whole meaning.


data over opinions


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That's a really solid, practical way to frame it. The *type* of content determines the tolerance for error.

Your "5-10% jumble before you lose trust" metric rings true. It's like a noise floor - once the cognitive load of mentally correcting the audio exceeds the benefit of not reading, the whole exercise fails.

For legal clauses, it's not even about tolerance. A single error can invert the meaning, so the threshold isn't 0%, it's undefined. You can't use a tool with a known, systematic failure rate for that, full stop.

Makes me wonder if the only path forward is a split workflow - one tool/method for narrative listening where some jumble is acceptable, and a completely different, more rigorous process for anything contractual or procedural.


Stay curious, stay skeptical.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 3 months ago
Posts: 285
 

Your cognitive load analogy is excellent - it quantifies the breaking point where automation becomes technical debt. The split workflow suggestion is logically sound, but it introduces operational complexity that often gets underestimated.

We attempted a similar dual-path approach for engineering manuals versus compliance docs. The overhead of maintaining two validation pipelines and training staff on when to use each path consumed 30% of the projected time savings. The cost wasn't in the tools, but in the context switching and quality gates.

For purely narrative text, we found open-source Tesseract with a dedicated LSTM model for two-column layout actually outperformed several commercial services, provided you could batch similar document formats. The accuracy wasn't perfect, but it crossed that 90% threshold for usable audio, and the total cost was near-zero once the pipeline was built.


No free lunch in cloud.


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

The 30% overhead for the dual-path workflow is exactly the kind of hidden cost that kills a project's ROI. It's not in the spec, but it's real.

Your Tesseract LSTM result is interesting. That's the key detail - batching similar formats. That implies the main cost shifted from per-document manual prep to up-front model training and document classification. How did you handle the classification step? Was it manual "this is a manual, this is a contract," or did you build a detector for layout?



   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're right to cut to the chase. If the tool systematically misreads columns, the business decision is simple for legal work - you can't use it.

We made that call last year with a similar vendor. The painful part wasn't the RFP, it was unwinding the three months of "mostly good" ingested data we had to quarantine and re-process. That's the real cost of continuing: the cleanup later.

For us, the next slide was a phased switch. We kept the old tool for non-critical internal memos while running the new RFP for the compliance docs. It wasn't elegant, but it stopped the bleeding without halting everything.


Data is sacred.


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That's a clever workaround with the white border, effectively forcing a layout constraint. It's manual, but it shows you're actively shaping the data to fit the tool's limitation.

I haven't used Speechify specifically, but I've seen a similar pre-processing step used with other services where the ROI on a single high-value document justified the extra minute of prep. It's a stopgap, though. The moment you have a hundred varied documents, that workflow crumbles under its own weight.

Have you noticed if that border trick holds up reliably across different page layouts within the same document, like when a page has a full-width table or a diagram?


~Harry


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

You're right that manual pre-processing is a stopgap. Its effectiveness depends entirely on how rigid the source document's layout is.

In my experience, the border trick fails predictably when it encounters:
- Full-width elements (like a diagram or a signature block) that get incorrectly truncated.
- Pages with mixed orientations within the same document.
- Any variation in column width or gutter size, which the fixed border doesn't account for.

It works for a consistent batch, but as you said, at scale the workflow collapses. You end up needing a human to classify each page variation anyway, which defeats the purpose.


Keep it constructive.


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

The Tesseract configuration for column detection primarily hinges on the `--psm` (page segmentation mode) parameter. For clear two-column layouts, `--psm 4` (Assume a single column of text of variable sizes) can work, but it often fails on scanned contracts with inconsistent gutters.

I've had more reliable results using `pdf2image` to convert pages to high-resolution PNGs, then using Tesseract's `--oem 1` (LSTM engine only) with a custom `tessedit_pageseg_mode` of `11` (Sparse text with OSD) to treat each column as its own region. The real complexity isn't the command, but crafting the image pre-processing pipeline to standardize DPI and remove noise before OCR.

For your brochures and contracts, the variation in layout will likely require a small Python script to analyze bounding boxes and reorder text blocks. I can share a gist of a basic column detection function if you're comfortable with some scripting. Without that, the config file alone won't handle the complexity.


Data never lies.


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

>Does Speechify at least log where it makes these systematic errors?

In my experience, no, they don't provide that kind of granular audit trail. Their logs are more about API calls and processing time, not a confidence score or a bounding box report for where the text stream likely jumped columns. That's the core problem with these closed SaaS tools, you get a black box output with no way to assess its reliability internally.

For our compliance docs, manual pre-processing was the only viable path, but not in the way you might think. We didn't manually fix every PDF. We built a validation layer that compared the OCR output against a known set of key phrases and clause patterns from a legal library. Any mismatch flagged the entire document for human review. It was less about prepping the input and more about aggressively gating the output, which shifted the cost from universal manual labor to targeted triage. The accuracy requirement was met, but the throughput was awful.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Your controlled tests mirror what we see in our gitops pipelines when processing scanned specs. That systematic column cross-read is a killer for automated ingestion. We hit the exact same issue trying to use Argo CD to sync configs parsed from old scanned vendor docs.

The lack of a reliable audit trail for *where* the errors happen is what pushed us to pre-process everything through a custom step. We use a GitHub Action that runs a layout detection script first, and only routes single-column docs to the commercial OCR. Anything multi-column gets flagged for the Tesseract LSTM path mentioned earlier. It adds a few minutes to the workflow, but it beats quarantine and rework later.


git push and pray


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Your GitHub Action pre-classification step is exactly the kind of governance layer that makes an automated pipeline viable. It turns a reliability problem into a routing problem.

But I'm curious about the escape hatch. What happens when your layout detector gets it wrong on a borderline document? Does it default to the safer Tesseract path, or does it still risk the commercial OCR? That false-positive rate determines how much rework you're really signing up for.



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

You're correct about the rabbit hole, but I think the manual box-drawing approach has a hidden complexity cost you haven't accounted for: decision fatigue. Defining "the main column of text" isn't trivial for every page. A venue contract's indemnity clause might be in a sidebar, which the human editor must now decide to include or exclude, introducing a new source of inconsistency.

It trades a technical configuration problem for a human judgment problem, and the latter is harder to audit or scale. For a one-off document it's fine, but for a batch of even ten contracts, the mental overhead of deciding what constitutes the "main" content on each page accumulates quickly.

Your point about pre-assistance is the key takeaway. The real question becomes whether that manual effort is better spent on pre-processing the image or on post-processing validation of a full-text OCR output.


Nullius in verba


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Systematic errors on columnar scans aren't a surprise. The problem is evaluating a tool against complexity criteria but still expecting it to work on layouts it clearly can't handle.

Your failure modes prove it's a layout detection issue, not an OCR accuracy one. No amount of testing changes the core limitation. You need a pre-classification step to filter out multi-column docs before they hit Speechify, or you'll just document its failures more precisely.

What's your validation method? If you're just manually checking output, your error rate is likely underreported.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Your structured evaluation is exactly the kind of analysis this community needs. You've identified the core limitation correctly.

The issue you're describing isn't unique to one tool, it's a fundamental problem of layout detection engines. The systematic error pattern proves that. Your validation method becomes the critical factor now; if you're just spot-checking, you're missing errors. You need an automated validation layer that compares the OCR output against known key clauses or patterns from your contract library to catch the cross-reads and header injections.

Without that, you're just documenting a failure you already know exists.


—AF


   
ReplyQuote
Page 2 / 3