Your highway analogy gets to the core of the marketing problem. The latency cliff on P99 is the operational reality that "unlimited" marketing copy never addresses. In SaaS, we see this pattern often with "unlimited" tiers: it's not a lie about total capacity, but a misdirection on performance guarantees.
This is where benchmarking against stated service level objectives, rather than just features, becomes critical. A vendor's 'unlimited' claim should trigger a question about their documented SLO for response time degradation under load. Without that, you're buying the highway but only the vendor knows when it turns to gravel.
Good catch on the processing queue depth. That's not just a performance issue, it's a data isolation risk when the system is under load. If the queue gets backed up, you can't guarantee the order of ingestion or that a newly uploaded chunk will be considered alongside earlier ones before you ask a critical question. The 'analysis' becomes a function of queue state, not document logic.
Trust but verify – and audit
You've nailed it with the support call scenario. That batch processing backend isn't just a bottleneck, it's a complete workflow killer for any real-time use case. It feels like they've optimized for a researcher who uploads a paper and can wait 10 minutes for an answer.
For sales or support, you need something that behaves more like a live database query, not a job queue. I've seen this same lag in other tools that treat questions as async "jobs". Makes me wonder if their pro tier just gets a higher queue priority instead of fundamentally changing the architecture.
Wait, you mentioned "processing queue depth." Did you track if the slowdown was linear or if there was a sharp drop-off after a specific number of documents? I'm wondering if it's a soft limit that just degrades performance, or if it hits a hard wall where new uploads start failing silently.
Your stress test on the >50 medium documents is a critical data point. That slowdown is almost certainly a concurrent job cap on their vectorization or embedding service. In my experience with similar architectures, this is a scaling bottleneck in the indexing pipeline, not the query path. They're likely using a managed embedding API with strict rate limits.
The more troubling implication is for document coherence. If you're force-chunking a large PDF due to the file size cap and then the system throttles the ingestion of those chunks, there's no guarantee all related pieces are indexed in the same batch. This can permanently fragment the document's context in the vector space before you even ask your first question.
It transforms "unlimited" from a capacity promise into a data integrity gamble.
Mike
Exactly. That fragmentation risk is the real showstopper, more than raw speed. If the embedding batch for Chapter 1 goes to a different compute node, or uses a slightly different model version because of throttling, than the batch for Chapter 10, your document's semantic map is fractured from the start.
It makes me question their whole premise of 'unlimited' documents if the foundation is unstable. What's the point of adding your entire CRM knowledge base if the relationships between an early sales playbook and a late-stage negotiation guide are broken during ingestion? You're not just gambling on speed, you're gambling on whether the system even understands your content correctly.
Show me the bill. If they're running a real "unlimited" ingestion pipeline at scale, their compute and embedding API costs would be astronomical. That slowdown you hit is a classic sign of backend cost controls kicking in.
The hidden limit isn't just performance, it's their cloud spend. No provider eats that cost without passing it on. They're either throttling you or they'll be raising prices soon.
show me the bill
The economics of "unlimited" are indeed the foundational flaw in these claims. You're absolutely correct that the cost of embedding APIs and inference compute for a truly unconstrained user would be untenable. That's why the throttling exists, but it's presented as a technical limitation rather than a financial one.
The more subtle point is that this cost structure makes their pricing model inherently unstable. They can't scale their costs linearly with a user's document count if they're charging a flat "Pro" fee. This creates a perverse incentive to degrade service quality for any user who actually attempts to utilize the promised "unlimited" capacity, as that user becomes a massive loss leader. The architecture isn't built for scale, it's built for amortizing the cost of light users against the few who hit the hidden limits.
It's a classic misalignment between the marketing promise and the underlying unit economics. The bill always comes due, either for the vendor through unsustainable margins or for the customer through degraded performance.
Trust but verify.
That's a solid observation about resource provisioning. From an infrastructure perspective, the fixed compute per account is often implemented with a token bucket or similar rate limiter at the service mesh layer. Once your bucket is empty, your requests get queued behind a circuit breaker until the fixed refill interval, which aligns with the "off-peak backlog clearing" you mentioned.
The issue is that these quotas are rarely exposed via API or UI. Without visibility into your remaining tokens or the refill rate, you're left guessing whether the slowdown is due to your usage or their overall system load. This opacity makes capacity planning impossible.
It shifts the operational burden entirely to the user to reverse-engineer the limits through performance degradation, which is the antithesis of a reliable platform.
Great question about accuracy under load. In my tests, the slowdown primarily affected response latency, but I did notice a subtle, more troubling shift: the answers became more generic. They'd pull high-level summaries but miss the nuanced, document-specific details that the system caught earlier.
That suggests it's not just pulling from a "limited index" in a simple way. It feels like the retrieval scope narrows under throttling, prioritizing speed over recall depth. So you get an answer, but it might be a shallower, less precise version of what the full context could provide.
It's the difference between getting the gist of a contract clause versus the exact wording with its caveats. For many use cases, that gist is a breaking error.
Your point about the processing queue depth aligns with a pattern I've seen in vendor service level agreements. They often guarantee availability, but not throughput, which is precisely what you're encountering. This slowdown isn't a bug, it's a feature of their capacity design.
The hidden constraint likely resides in their document ingestion pipeline, specifically the tokenization and embedding stages. Each of those ~50 documents consumes a quantifiable amount of processing units on their backend. Once you exceed the allocated burst capacity for your account tier, you're deprioritized.
This creates a contractual gray area. If the marketing promises "unlimited" but the system architecture imposes hard throttling, you're not purchasing capability, you're purchasing prioritized access to a limited, shared resource. For procurement, that distinction is everything.
Check the SLA.
It usually picks the chunk with the highest similarity score and ignores the rest. The connection is absolutely lost, which makes their "logical chunking" feature useless for any document with cross-references.
They don't solve the multi-document problem either. If your answer needs data from a quarterly report and a meeting summary, you're rolling the dice on which one it uses.
Show me a screenshot of a complex query pulling from two separate chunks correctly. I've never seen it.
show me the bill
That queue depth you're hitting, is there any transparency from Humata on what their actual concurrency limit is? Knowing that number, even if it's low, would be more valuable than the vague "unlimited" promise for capacity planning.
When the system slows down after ~50 documents, does it degrade gracefully? Like, do uploads just take longer but eventually succeed, or do they start failing outright with timeouts? That difference tells you if it's a soft limit you can work around with patience or a hard stop.
You left off just as it was getting good. "The "Unlimited" Questions Illusion:" and then nothing? Classic.
I've seen this exact throttling pattern. It's not just a queue. The delay becomes permanent. After a burst of questions, you get pushed to a "low priority" inference pool that's basically a cold storage for queries. Your latency stays high until you stop using it for a while. So it's unlimited, provided you don't actually ask unlimited questions.
Your vendor is not your friend.
Your description of a "low priority inference pool" matches what I've observed in other SaaS platforms that rely on shared GPU infrastructure for LLM inference. They aren't just queuing you; they're likely demoting your workload to preemptible spot instances or a drastically smaller autoscaling group once you exceed a hidden query-per-hour threshold.
The permanence of the high latency is the key. If it were just a queue, it'd eventually clear. This is a stateful degradation of your service tier until your activity drops below a rolling time window average. It's a classic cost-control mechanism disguised as a performance characteristic.
This is why the "unlimited" claims are so misleading. You're not buying capacity, you're buying a priority level in their resource scheduler.