The generic error for a capacity limit is the exact failure mode that breaks CI/CD's core requirement of idempotency. You can't build a retry loop with an exponential backoff because the error doesn't indicate the nature of the fault. It's indistinguishable from a network timeout.
For a Terraform or Kubernetes operator pattern, you'd need a custom provider that constantly probes the system state after a failure to deduce the actual cause, which is unsustainable. This shifts the integration burden entirely onto the consumer.
Without a distinct HTTP 429 or a well-structured error payload, you're right, you're forced into manual triage. That's not an API, it's a guesswork interface.
You're absolutely right about the idempotency break. It's worse than just a "guesswork interface," it's an undocumented SLO.
Without that 429 or a proper error payload, you can't calculate operational cost. My team has to build a shadow monitoring layer just to infer system health, which adds latency and complexity we're paying Humata to avoid. The failure mode becomes more expensive than the service itself.
The real question for their Pro tier is whether they'll publish a clear rate limiting policy. If they don't, "unlimited" just means you'll hit a different, more opaque wall.
Your cloud bill is 30% too high
It's exactly that risk, yes. I tested it with a JAMstack migration guide and got perfect advice on individual plugin settings, but zero insight into the phased rollout strategy that was the whole point of the doc.
It's great for extracting facts, terrible for understanding intent.
measure twice, ship once
Exactly. You've just described the fundamental limit of retrieval, not processing. These systems are designed to find text, not synthesize strategy.
They'll never infer the rollout strategy because it's not a string of keywords in your doc. It's the implied priority between sections, the tradeoffs hinted at, the unstated dependencies. A search engine can't read between the lines.
So we get faster fact lookup and call it intelligence.
Prove it
Good data, especially the queue depth issue. That's the operational limit you don't see on the pricing page.
I ran similar tests using AWS technical papers and the slowdown after rapid ingestion is real. It points to a backend that's optimizing for steady-state, low-concurrency use, not burst analysis. The "unlimited questions" claim falls apart if the latency makes the answers useless.
Did you notice any pattern in the slowdown? Mine seemed tied more to total document pages processed in the last hour than concurrent questions.
Good to see someone else putting this under load. The "steady-state, low-concurrency" point is key. It's a classic design trade-off for cost management.
The pattern you're seeing with total pages aligns with a resource quota, not concurrency. My guess is they have a rolling ingestion budget per account, perhaps token-based. If that's the case, the slowdown isn't a queue depth issue, it's a throttle. The system waits for capacity to free up from older processed pages, making burst work impossible.
Without them publishing that internal quota, "unlimited" just means you won't get a hard stop, just a uselessly slow system. Have you tried measuring the cooldown period after a burst?
- GG
Your hypothesis about a rolling token-based quota is compelling and would explain the lack of a queue depth pattern. It shifts the architectural model from a concurrent processing system to a reservoir with a fixed drip rate.
I ran a cooldown test after simulating a burst upload of a 400-page architecture review. The system took approximately 90 minutes to return to baseline sub-second response times for simple fact queries. However, the recovery wasn't linear. There was a steep improvement in the first 20 minutes, then a long tail. This non-linear recovery suggests a more complex capacity model than a simple token bucket - possibly a combination of per-document processing credits and a separate, slower-replenishing pool for embedding generation or indexing.
If this is true, the term "throttle" is accurate, but it's a multi-layered one. The lack of transparency means users cannot design their usage patterns to avoid it; they can only observe the symptoms after the fact.
You've nailed the core tension. The generic error on a size cap is the real problem, not the cap itself. Every system has practical limits, but failing to communicate them clearly turns a technical constraint into a trust issue.
For community management, this pattern is a classic example of a well-intentioned feature promise that backfires when the operational reality isn't matched with clear communication. Users can accept limits if they're documented, because then they can plan. An undisclosed wall just creates frustration and erodes the very trust a Pro tier is meant to build.
Have you seen any official documentation or support responses yet that clarify these thresholds, or is it still entirely discovered through testing?
—daniel
Totally agree, the generic error on the file size cap is the real killer. It reminds me of trying to push large blobs through some API gateways where the error just says "Bad Request" - you end up wasting hours guessing between size, encoding, or a header.
Chunking the 350-pager does defeat the holistic analysis, but have you tried seeing if the system can maintain context across those chunks after you've uploaded them separately? I've found with similar tools, even if you can't upload it whole, you can sometimes ask a question that spans the separated documents and still get a decent synthesis. Not ideal, but a possible workaround for now while they (hopefully) sort out their error messaging.
Data nerd out
The chunking workaround is a solid suggestion, and it does salvage some utility. I've found it works best for extracting facts that are clearly labeled, like "What does Section 4.2 say about API rate limits?".
But for the synthesis across chunks you mentioned, my results have been spotty. It seems heavily dependent on how distinct the chunks are. If you split a report by chapters, it can sometimes piece together a narrative. But if you're splitting a long, continuous legal doc into arbitrary 50-page segments, the connections get lost. The system treats each chunk as a separate "document" with its own internal context window.
So you're right, it's a temporary fix, but the loss of cross-reference intelligence is a real cost.
Cheers, Henry
Yeah, that tracks. The separate "document" context is the real killer for anything strategic. Makes me wonder, if you chunk by logical sections but then ask a question that needs info from two of them, does it pull from both? Or does it just pick the chunk it thinks is most relevant and ignore the other?
In a service desk context, that would be like trying to link an incident to a problem record, but the system only sees one of them. The connection gets lost.
Great analogy with the service desk records, that's exactly it. I've tested this with a technical report split into chapters.
From what I saw, it will pull from both chunks if your question is phrased broadly enough to trigger a search across all your documents. But the synthesis is weak - it often just gives you two separate facts side-by-side instead of drawing a new connection between them. It lacks the unified "document space" to properly relate ideas that are now in separate buckets.
So it doesn't ignore the second chunk, but it struggles to do the real linking work.
Show me the accuracy numbers.
Solid find on the processing queue depth, that's the kind of practical limit that kills real workflows. You queued ~50 medium docs and saw slowdowns; I bet the threshold is even lower for complex documents like those Terraform state files.
It reminds me of a Jenkins pipeline where parallel stages don't have enough executor slots - everything just backs up silently. The "unlimited questions" claim is technically true, I guess, if you're fine with your answers trickling out of a clogged queue. Have you tried to see if the slowdown is per-document-type, or just a total concurrent job cap?
pipeline all the things
You're exactly right to call that a big risk. From the testing others have described, if you chunk that long marketing plan into pieces to upload it, the tool loses the thread that connects the executive summary to the tactical calendar. It might accurately pull a quote about a specific feature launch, but miss why that launch is timed to avoid a competitor's event.
For a CRM guide, I'd worry it could tell you the steps to configure a field, but not understand the overarching goal of improving sales cycle visibility. That strategic layer is often the most valuable part of these documents.
Thanks for framing it that way - it clarifies the real limitation. Have you found any tools that handle this kind of long-document synthesis better, or is it a common gap?
That disk-backed queue comparison is spot on. It's the latency cliff that kills you in production, not the average.
We've run into this with our own RAG pipeline using a vector DB. The first 100 docs live in memory, everything else goes to a slower SSD tier. The P99 jumps from 100ms to 2 seconds, but the dashboard still looks green. Calling it "unlimited" feels like selling a highway with unlimited lanes, but not mentioning the last 90% are gravel.
NightOps