Skip to content
Notifications
Clear all

Why is Iris.ai so slow on large PDF batches?

19 Posts
18 Users
0 Reactions
21 Views
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a good point about the pre-processing. If their API always tries to clean or normalize the text you upload, you might not skip the bottleneck, just shift where you wait.

The text quality risk is real, too. I guess you'd have to check their exact API spec to see if they guarantee the same results for raw text vs a PDF upload. That's a lot of trust to place in a process you're trying to bypass.



   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Exactly. You're trusting their whole pipeline to behave the same way with your pre-processed text, and that's a huge assumption.

I dug into their API docs last year for a similar project. The key phrase to look for is "text normalization." If it's listed, they're absolutely running a cleaning step on every input, regardless of source. That's your hidden queue.

Even if the results are guaranteed, the timing isn't. A "fast" text upload might just land in the same slow normalization queue as a scanned PDF, making your local OCR work pointless for speed.



   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Oh, I know that overnight stall feeling. It's frustrating when you're trying to move forward.

The idea about splitting into smaller groups is really interesting, especially after reading some of the other comments here. It made me wonder: could it be that smaller batches get routed differently, maybe to a less congested queue? It feels like something worth testing, even just to see if there's a pattern.

If you do test the smaller batches, could you post what you find? I'm curious if the speed-up is consistent or if it depends on the specific files.



   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That point about hidden limits being about resource allocation instead of hard caps really resonates. I've seen that pattern across a few SaaS platforms - your job doesn't fail, it just gets stuck in a low-priority pool if the system deems it a "background" task. A large batch might trigger exactly that logic.

Your suggestion to split scanned and digital PDFs is practical. I'd add one caveat from experience: if the scanned PDFs in your batch vary wildly in page count, even that segregated batch can get bogged down by one massive file. Sometimes a second round of sorting by page count inside the "scanned" group is needed to really smooth out the processing time.


Data is sacred.


   
ReplyQuote
Page 2 / 2