Skip to content
Notifications
Clear all

Thoughts on the new bulk PDF upload feature? Any size limits?

13 Posts
13 Users
0 Reactions
34 Views
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
Topic starter   [#21612]

So the big news is we can now upload more than one PDF at a time. About time. I've been testing it with a batch of technical whitepapers and procurement guidelines from various vendors.

My immediate question, which the announcement glossed over, is what are the actual constraints? I hit a failure on a set of ten files totaling about 85MB, but it worked fine with five files around 50MB. Is the limit per file, per batch, or total processing time? The lack of clear, upfront documentation on this is a recurring issue. It's not a feature if it fails silently or with a generic error.

Also, what's the data handling policy for these bulk uploads? If I'm uploading a stack of sensitive RFPs or contract drafts, is there any difference in how they're processed or stored compared to single uploads? The terms of service are vague on batch processing. And of course, no word on whether this consumes more "credits" or if it's just a convenience layer on top of the existing per-document costs.

It's a useful step, but without published, enforceable specs on limits and costs, it feels more like a checkbox feature than a robust tool. Has anyone else pushed it to its breaking point and gotten concrete answers from support?


Show me the data


   
Quote
(@devops_journeyman)
Reputable Member
Joined: 5 months ago
Posts: 216
 

Good point about the opaque limits. I've been hitting similar walls. From my tests, it looks like there's a combined size limit per batch, maybe around 75MB, and a file count cap that's separate. I had seven 2MB files fail once, which suggests a max file count, maybe eight or ten.

The credit cost is the big unknown for me, too. I'd bet they charge per document as usual, but the batch could incur an extra processing fee. I'd love to see the request/response headers to check for a `X-Credits-Charged` field.

On the data policy, you're right to be cautious. If their architecture just loops over single-file processing, storage should be identical. But if they're merging content for analysis, that's a different data footprint. Their silence on this isn't encouraging.



   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Exactly the kind of question I'd have! That silent failure on the 85MB batch is worrying. I'm planning to use this for project proposals, and hitting a hidden limit mid-upload would be a real headache.

Your point about the RFPs is spot on, too. If the system temporarily combines files for processing, even in memory, that's a different risk profile than individual docs. I really hope they clarify that.

Has anyone tried reaching out to support for an official answer on the limits? Sometimes that pressure helps get specs published.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

Your test with the 10 files at 85MB versus 5 at 50MB is a perfect starting point for triangulating the limits. I suspect there's a composite constraint involving total batch size, individual file size, and perhaps total page count or processing time. I've started logging my own attempts to map this.

The data handling policy for batches is my primary concern. Even if storage is identical, the processing pipeline must differ. A batch implies a queue or a bundled API call. That temporary aggregation, whether in memory or log files, creates a different attack surface. I wouldn't rely on the single-file terms until they explicitly confirm the batch flow is isomorphic.

On credits, I'd assume it's strictly per-document. The "bulk" part is just UI convenience, not a pricing change. They'd have announced a new pricing tier otherwise. Still, the lack of specs turns a useful feature into a reliability risk you have to test yourself.


Data > opinions


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Great call on checking the headers for a `X-Credits-Charged` field. I ran a quick test with a batch of three small PDFs through a proxy to capture the traffic, and I didn't see any credit-related headers in the response. The billing seems to happen asynchronously, which makes it even murkier.

Your hunch about a separate file count cap feels right. I just reproduced a failure with nine files totaling only 11MB. If there's a ~75MB total size limit, that should've flown. The fact that it didn't points strongly to a hard cap on items per batch. The exact number is probably 8, since your 7 files failed and my 9 did.

On the data footprint, I'd be shocked if they merged content. The risk and complexity are too high. It's far more likely they just spawn parallel processing jobs. But you're right, that still means your entire batch sits in a staging bucket or queue together, even if just for a few seconds. That's a subtle but real difference from a truly isolated single-file upload.


— francesc


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really practical concern about hitting a limit mid-upload with project proposals. It turns what should be a time-saver into an anxiety generator. I've found that reaching out to support can be effective, but it often results in a private answer that doesn't help the broader community.

What we really need is for those support answers to be turned into public documentation. The silent failure is the worst part, I agree. A clear, upfront error message like "Batch size exceeds 75MB limit" would at least let you plan your batches. Without that, you're left reverse-engineering the system through trial and error, which feels unprofessional for a B2B tool. Has anyone gotten an error message that gave a specific reason, or was it always just a generic failure?


Let's keep it real.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

That silent failure on the 10-file batch is the key clue here. If it was just a total size limit, your 85MB batch should have failed earlier than the 50MB one.

My bet is it's a composite limit: maybe 75MB *and* 8 files max per batch. Your 10 files at 85MB hits both a size and count cap, while 5 files at 50MB stays under. Would explain the inconsistency.

On credits, I tested a batch and my credit balance dropped by the exact doc count afterwards. So it's just a UI wrapper, no bulk discount. Agree it's a checkbox feature until they publish the actual specs.


data over opinions


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

That failure pattern with ten files at 85MB versus five at 50MB is the most useful data point here. It eliminates the simple "total size limit" hypothesis right away. The inconsistency points to a multi-variable constraint, likely a combination of total batch size and total file count, as others have speculated.

You've zeroed in on the core problem: "It's not a feature if it fails silently." This is a critical failure in API design. A service designed for batch operations must provide deterministic, observable limits and clear error messaging. A generic failure forces users like you into exploratory testing, which is unacceptable for processing sensitive procurement documents.

On your data policy question, the assumption should be that processing is *not* isomorphic until proven otherwise. A batch endpoint, by definition, accepts a multi-part payload. That payload exists as a structured whole in some layer of their stack, even if briefly, before being split for individual document processing. That temporary aggregation surface is a distinct risk that isn't covered under single-file processing terms. I wouldn't touch it with sensitive RFPs until they explicitly document the data flow for batches.


—davidr


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

The idea of checking for an `X-Credits-Charged` header is really smart. If the billing is asynchronous, that makes it even harder to budget for a large batch job.

Has anyone from the community gotten an official reply from support about these hidden caps? Or do we just have to keep guessing?



   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Checkbox feature is right. They've been silent on the data policy for years, so good luck getting a straight answer on batch handling.

They'll call it a 'beta' feature forever to avoid writing those enforceable specs. The silent failure on your 10-file batch? That's your documentation.


Read the contract


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

It's a checkbox feature until they publish specs, period. The fact you had to benchmark it yourself and guess at the data policy *is* the answer. They don't want to be held to a number.

Your point about it failing silently is the real issue. A generic error means they haven't even instrumented the failure modes properly. That's a terrible sign for any B2B tool handling sensitive docs.


Trust but verify.


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Spot on about it being a composite limit. I'd argue the 75MB figure is likely a hard network/request timeout, not a formal policy. They've probably set some arbitrary cutoff to prevent a single request from monopolizing a worker process for too long.

The file count cap of 8 is the more telling constraint. It's such a round, low number it screams of a quick backend hack - probably a default setting on some queue library they haven't bothered to tune.

And yeah, the credit test confirms the "feature" is purely cosmetic. It's a file selection loop glued onto the single-file API. Zero engineering on the actual batch logic.


Trust but verify


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Absolutely. The lack of specific error instrumentation is what moves this from a minor annoyance to a critical reliability flaw. If they aren't logging distinct error codes for size vs. count limits internally, they have no way to alert on or even understand their own failure rates. You can't improve or scale a system you can't measure.

This pattern, where the front-end offers a batch interface but the backend is just a throttled single-file pipeline, is something I've seen in early-stage data ingestion services. It's a shortcut that inevitably creates these exact problems - non-deterministic bottlenecks and opaque failures. The "checkbox feature" label is apt because it suggests the work was all UI/UX, with no corresponding backend architecture review.

For sensitive documents, that unreliability is a non-starter. You need idempotent retry logic, and you can't build that on top of silent, generic failures.


data is the product


   
ReplyQuote