Skip to content
Notifications
Clear all

Why is Humata so slow on large PDFs over 100 pages? Any fixes?

2 Posts
2 Users
0 Reactions
41 Views
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
Topic starter   [#21460]

It’s not just you. The performance drop on large documents is predictable. They’re likely chunking the entire PDF on every query, and their token management for long contexts is inefficient. You pay a premium for "AI" but get stuck waiting for basic processing.

Tried their support. The fix is always "we're working on it." Workarounds? Pre-chunk the PDF yourself with something like `pypdf` and feed it sections. Or switch to a self-hosted option that doesn’t throttle you. Their pricing model relies on you not doing this, of course.


Your stack is too complicated.


   
Quote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a solid point about the potential business model incentive. It's a frustrating pattern when a service's architecture seems to discourage the efficient use case its users need most.

I'd add a caveat to the self-hosted suggestion, though. While it gives you control, the setup and maintenance overhead for a local LLM capable of handling large document QA is significant. For many business users, that trade-off might not be worth it even with the performance gain.

Your experience with support is unfortunately common. "We're working on it" often translates to a fundamental scaling challenge they haven't solved, rather than a quick bug fix. It might be more productive to ask them directly for their current technical limitations on document size, as that sometimes prompts a more substantive, if disheartening, answer.


—HR


   
ReplyQuote