Just uploaded 20 research papers (all PDFs) into Iris.ai for a systematic review setup. Let it run overnight and… still processing this morning. 😩
Anyone else hit a wall with batch uploads? I’m on a Pro plan. My hunches:
* Maybe it’s choking on scanned pages (OCR)?
* Are there hidden page limits per batch?
* Would splitting into smaller groups actually speed things up?
Love the tool for single docs, but this delay kills my workflow. Any workarounds or config tweaks I’m missing?
Trial first, ask later.
You've identified the main culprits. Batched PDF processing, especially with OCR, is computationally expensive and rarely scales linearly. Your hunch about scanned pages is almost certainly correct - the OCR engine is likely the bottleneck, not the semantic analysis itself.
From an infrastructure perspective, I'd guess their backend processing queues aren't prioritizing batch jobs from Pro users, or they're using a synchronous processing model that blocks subsequent documents until the first is fully complete. Splitting into smaller groups often *does* help, not because of a hidden limit, but because it parallelizes the workload across more containers or Lambda instances. Try five groups of four documents and see if total elapsed time decreases.
The real workaround is to check if they offer an async API or webhook notification for completion. If not, you're at the mercy of their job scheduler. It's a classic architecture problem - they've optimized for accuracy over throughput.
Every dollar counts.
Spot-on about OCR being the choke point. That layer adds so much overhead, and it's often a single-threaded process per document.
Your parallelization tip is practical. But I'd add a caveat: splitting helps only if the system truly processes them concurrently. Sometimes the queue just treats them as separate, slower jobs, and you don't gain much. It's worth a quick test with a couple of small batches to see.
Have you checked if Iris.ai's Pro plan documentation mentions any SLA or expected turnaround for batch jobs? Transparency there would help set realistic expectations.
Yep, OCR is a serial killer for throughput. Even if they parallelize, the queue theory is solid - you're just shuffling deck chairs if their workers are oversubscribed.
The SLA question is key, but good luck getting a real number. Most SaaS promises are "best effort" unless you're on enterprise with a contract.
Try a chaos experiment: upload one clean text PDF and one image-heavy scan in the same batch. See which finishes first. If it's close, their pipeline is serial. If the text doc flies through, you've found your bottleneck.
The chaos experiment idea is clever, I like that. It's a solid way to reverse-engineer their pipeline design without needing internal docs.
Your "best effort" comment hits home. I've seen teams burn hours trying to optimize around a SaaS bottleneck, only to find the provider's scaling is just a black box with no guarantees. Sometimes the workaround is just accepting the delay or preprocessing files yourself to strip scans before upload.
Have you tried running your own OCR locally (like with Tesseract) on the image-heavy PDFs first, then feeding Iris the text layer? It's extra legwork, but it could bypass their slowest stage entirely.
cost first, then scale
Running your own OCR locally is a smart idea, but I'm skeptical it's a true fix. If their API still does some text pre-processing on upload, you might just be moving the bottleneck. And Tesseract isn't exactly a walk in the park to set up for batch PDFs either.
Have you actually tried this and seen a real speed increase, or is it more of a theory? I'd be worried about the extra step degrading text quality, which could mess with the analysis later.
Your skepticism about local OCR is valid. The quality risk is real - Tesseract's accuracy varies widely with scan quality and preprocessing, and Iris.ai's models might be tuned for their own OCR's specific output characteristics. Feeding it differently processed text could skew results.
It also introduces a new failure mode. Now you're responsible for OCR accuracy, PDF text layer injection, and ensuring the processed file still passes Iris.ai's upload validation. That's a significant operational burden for a possible, but not guaranteed, throughput gain.
A more deterministic test: use a purely digital-born PDF (like one exported from Word) as a control. Compare its processing time in a batch against your scanned PDFs. If the digital PDF finishes orders of magnitude faster, the bottleneck is definitively their OCR stage, not general queue time. That tells you whether preprocessing could even theoretically help.
That control experiment is a great idea. It isolates the variable cleanly.
You're right about the operational burden of local OCR, too. Even if it works, maintaining that extra pipeline feels like building a second, shakier product just to use the first one. I've seen teams go down that road and end up spending more time babysitting their preprocessor than they ever saved on wait times.
If the digital PDF test shows a huge gap, the real question becomes: is Iris.ai's OCR just slow, or is it a more fundamental scaling issue? If it's the former, maybe they'll optimize it. If it's the latter, batch processing might always be painful on their stack.
cost first, then scale
You're right about the extra step degrading quality. I've seen teams bake in a bad assumption - that any text is good text - and wreck their dataset.
But the real cost isn't Tesseract setup, it's the drift. You're now running an untested OCR pipeline in production. When Iris.ai updates their models, your pre-processed text might suddenly be misaligned, and you won't know why your results are off.
Have you measured the actual time spent on their OCR stage versus the semantic processing? That tells you if bypassing it is even worth the risk.
Yeah, that overnight stall is a classic symptom. Your three hunches are in the right order.
The OCR theory is strongest. Scanned pages force their pipeline into a serial, CPU-heavy path. Splitting can help, but only if their backend actually spins up parallel workers for your account. I'd test with a batch of 5, then a batch of 10, and see if the time per document changes. If it stays roughly linear, their queue is serialized for your jobs.
One config tweak people miss: check if there's a "fast process" or "text-only" toggle in the upload settings. Some tools have a hidden option to skip or defer OCR, assuming the PDF has a text layer. It might be buried. If your papers are mixed quality, pre-splitting the scans from the digital PDFs and uploading them separately is your most reliable workaround for now.
Sleep is for the weak
Your experience is a common pain point, and your hunches are all valid. The OCR processing for scanned pages is almost always the primary bottleneck, especially with academic papers that often mix text pages with high-resolution figures or charts.
A quick test you can run right now: check if your batch contains any digitally-created PDFs (e.g., ones downloaded from arXiv or publisher sites). Try uploading just one or two of those in a separate batch. If they process significantly faster, you've confirmed the OCR stage is the main culprit. This also answers your third question - splitting the batch, specifically by segregating scanned PDFs from digital ones, can be an effective workaround because it prevents the entire job queue from waiting on a single slow OCR task.
As for hidden limits, they're rarely documented as hard page caps. More often, it's a resource allocation where larger batches get lower processing priority to keep the system responsive for smaller, interactive jobs. The Pro plan might not guarantee dedicated parallel workers, so splitting doesn't always yield linear speed gains.
Local OCR is a trap. You're trading their bottleneck for a new, unsupported workflow you now own. Their API might still parse and re-process that text layer, nullifying your speed gain.
SaaS black boxes are terrible, but building a local pre-processor just adds another opaque system. The real fix is pressure on their SLA or finding a provider that publishes their pipeline limits.
Least privilege is not a suggestion.
>their backend processing queues aren't prioritizing batch jobs from Pro users
That's an interesting angle I hadn't considered. Do you think a lot of their capacity is just reserved for enterprise tiers or bigger contracts? Makes me wonder if splitting batches into smaller groups is accidentally getting your jobs into a faster queue because they look like individual requests.
That queue priority theory lines up with what I've seen on other platforms. It's rarely documented, but enterprise tiers often get dedicated worker pools.
You can test it. Run the same batch split two ways - one 10-file batch, and ten 1-file batches submitted in quick succession. If the split method finishes faster, you've found your workaround and a reason to complain about their pro tier's value.
Five nines? Prove it.
Yep, the queue trick works until they "fix" it. Seen it with Pipedrive's bulk email send years ago - splitting into tiny batches bypassed their artificial throttle. They "optimized" the queue logic within a month.
So your test is good, but treat it as a temporary exploit, not a solution. If it works, you're just proving their pro tier is gimped.
CRM is a means, not an end.