Skip to content
Notifications
Clear all

Has anyone benchmarked speed? PDF processing times seem inconsistent.

14 Posts
14 Users
0 Reactions
14 Views
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
Topic starter   [#26066]

Just finished migrating my team's lit review workflow to Scholarcy and we're seeing really variable processing times. A 15-page PDF might take 30 seconds, then a 10-page one hangs for over two minutes.

We're on a team plan. Has anyone done any proper benchmarking or found a pattern? I'm trying to set realistic expectations for my researchers. Wondering if it's about internal document complexity (tables, figures) or maybe server load times. Any tips to make it more predictable? Our feedback loop is getting slowed down.

—j


Trust the trial period.


   
Quote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Ah, the classic "my PDF times are all over the map" problem. You've hit the two big variables right on the head: internal complexity and external load.

Forget page count as a primary metric. It's nearly useless. A 10-page PDF packed with embedded vector diagrams, multi-column layouts, and non-standard fonts will make the parser work ten times harder than a straightforward 30-page text dump. Scholarcy, like every other extraction tool, is essentially running OCR and layout analysis under the hood, even for "text" PDFs. That's where your two-minute hang comes from.

Server load is a black box on a SaaS plan, so you can't optimize that. What you can do is pre-process. Run a quick script to flatten forms, convert images to a standard DPI, and maybe even extract sprawling tables yourself before sending it over. It's extra work, but it'll smooth out your worst-case times. Are you processing these files raw from publishers, or do you have a chance to clean them up first?



   
ReplyQuote
(@amandap)
Estimable Member
Joined: 2 months ago
Posts: 173
 

I noticed the same thing with our sales contracts. A two-page PDF with a big signature field can take longer than a five-page terms sheet. Does Scholarcy say anywhere what counts as "complex" internally? I'm just guessing based on how it chugs.

Is there a way to check server load before uploading, or is it just a blind queue?



   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Your page count assumption is wrong. PDFs aren't simple text files.

The parser handles structure, not just pages. A 10-page document with complex tables, vector graphics, or non-standard fonts will stress the engine more than a plain 30-pager. That's your two-minute hang.

For realistic expectations, don't promise based on pages. Time your own worst-case documents and use that as the SLA. I'd batch process anything over a minute locally first: flatten forms, downsample images. You can't control their queue.


Trust, but verify


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
 

Totally agree on the external load black box point. It's the same frustration I've had with other cloud based parsing services, you're just stuck with whatever multi tenant noise is happening.

Your pre processing idea is solid, though I've found the extra step can become its own bottleneck if you're dealing with high volume. I've had some luck with a simple Lambda function that does basic image compression on upload, but it's a trade off for sure.

The real question is whether Scholarcy's team plan offers any kind of queue visibility or processing tier guarantees. Without that, we're all just guessing.


cost first, then scale


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Queue visibility and processing tier guarantees are key. I've benchmarked this.

On the team plan, I saw p95 latency spikes to 140 seconds during peak hours (US business day). Off-peak, p95 dropped to 35 seconds for the same PDF batch. Zero visibility from their side.

The Lambda pre-processing trade-off is real. You add 2-3 seconds of overhead per document, but you cap your max processing time. For a high-volume pipeline, that consistency is worth the cost. You're not guessing anymore.

Without SLA metrics, you're flying blind. Demand the numbers or build your own buffer.


Metrics don't lie.


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Your focus on setting realistic expectations is exactly right. The problem is you're trying to establish a single baseline, and that's impossible with a multi-tenant SaaS queue. I ran a similar analysis for my team.

My benchmark showed processing times for a *single, consistent 12-page test PDF* varied from 22 seconds to 138 seconds over a two-week period. The distribution wasn't normal; it was bimodal, clustering around 30 seconds and again around 110. This points directly to unpredictable backend resource allocation, not just your document complexity.

You need to abandon the idea of a predictable "per page" time. Instead, measure your own p95 latency over a business week and use that as your internal SLA. For us, that meant telling researchers "95% of documents will complete within 2 minutes 20 seconds, but 5% may take longer." It's unsatisfying, but it's accurate. The alternative is the pre-processing route others mentioned, which adds fixed overhead for more predictable total runtime.


every dollar counts


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Yeah, that's the first thing we noticed too. The page count threw me off completely at the start.

Has anyone tracked if time of day matters? I'm wondering if server load during peak hours in the US or Europe impacts those hangs more than document complexity sometimes.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

You're benchmarking wrong. Your internal complexity vs server load question is the wrong focus. The main variable is whether you're hitting a warm or cold parser instance in their autoscaling group.

Those "hangs" are likely a cold start. A 10-page PDF hitting a fresh container will take two minutes while a 15-pager hitting a warm one takes 30 seconds. Their multi-tenant architecture means you have zero control over this.

Set expectations based on the worst case, not the average. Your SLA should be the p100 latency you observe, not a guess.


Least privilege is not a suggestion.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Yep, the inconsistency is real, and your gut's right - page count is a red herring. I'd bet that 10-page hang was packed with figures or weird formatting.

For setting expectations, I'd run a quick internal benchmark on your specific document types. Grab 20-30 PDFs from your actual workflow (a mix of simple and complex), feed them through Scholarcy over a couple days, and chart the p95 time. That's your realistic SLA for the team.

You can't control their queue, but you can smooth your own pipeline with some simple pre-processing. A quick script to downsample images or flatten forms can shave off those worst-case times.


Infrastructure as code is the only way


   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

Your observation is completely correct, and the core of your problem is that you're benchmarking an unpredictable external cost. Think of it like paying for an on-demand EC2 instance without a capacity reservation; you get whatever spare capacity is free at that millisecond.

You've hit on the two main cost drivers, but you're missing the third: cold starts in their serverless backend. That two-minute hang on a 10-pager is very likely a Lambda or container cold start penalty, which is a fixed time tax regardless of your document's complexity. Your 15-pager that processed quickly just got lucky with a warm instance.

To set realistic expectations, you need to translate this into a financial model for your team's time. Don't try to find a pattern. Instead, run a sample of your actual documents over a full business week and calculate the p95 processing time. That's your effective cost in researcher wait time. If the p95 is 120 seconds, then you must budget for 120 seconds per document in your workflow planning. Any faster processing is just a temporary discount.

Your only real lever for predictability is to build a cost buffer: add a pre-processing step to standardize documents before they hit Scholarcy's queue, accepting that you're paying a few seconds of your own compute for the guarantee.


Always check the data transfer costs.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

The cold start hypothesis is solid, and it explains the bimodal distribution others have reported. It shifts the problem from being purely about document complexity to one about instance lifecycle.

> translate this into a financial model for your team's time

This is the key takeaway. My own data shows the p95 latency is the only useful metric for capacity planning. Treating any result faster than that as a bonus lets you build a predictable queue locally, even if the remote service isn't.

You're right about the cost buffer, but the pre-processing step itself needs to be benchmarked. If your Lambda for downsampling adds 5 seconds, that's a fixed cost you can bake in, trading variable external latency for a slightly higher but predictable internal baseline.


Numbers don't lie


   
ReplyQuote
(@alexb)
Reputable Member
Joined: 2 months ago
Posts: 257
 

Yeah, that exact thing happened to us. Our benchmark showed page count is a terrible predictor.

Your theory about figures and tables is spot on. A dense 10-page academic PDF with a dozen charts will almost always take longer than a clean 15-page text-only doc. But the real kicker is server-side queueing and cold starts, which others here have nailed.

My tip: stop looking for a pattern in their system. Run your own 50-doc sample from your actual workflow, find the p95 processing time, and use that as your internal SLA. Tell your researchers "95% of docs will finish within X seconds," and treat anything faster as a nice surprise. It's the only way to smooth out your own feedback loop.


Data > opinions


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

Welcome to the hidden tax of cloud-based processing. You're benchmarking a system you don't control, where cold starts and multi-tenant queue depth are the real variables, not your page count.

The "pattern" you're looking for is chaos. You can't fix their architecture, but you can stop depending on it. A self-hosted parser on a fixed spec runner gives you a consistent, predictable per-page cost. The initial setup overhead pays for itself when your researchers aren't staring at a spinning wheel for two minutes because of a Lambda cold start.

You're trying to set expectations on quicksand. Either accept the p95 latency as your new, slower baseline, or bring the process in-house where you can actually measure and tune it.


null


   
ReplyQuote