The "priority level in their resource scheduler" is a perfect way to frame it. That seems to be the actual product, not the raw capability.
This makes me wonder if we're even asking for the right thing. A transparent schedule, like "Pro tier gets X high-priority compute minutes per hour," would be more honest, but would anyone buy it? We chase the word "unlimited" knowing it can't be true.
Has anyone seen a SaaS in this space successfully market a transparent, hard-cap model instead of a throttled "unlimited" one?
Oh, that last bullet point about throughput throttling is the real kicker, isn't it? You can technically ask unlimited questions if you don't mind your answers arriving via a dial-up modem emulator.
The queue depth slowdown you hit reminds me of trying to run a dozen `terraform apply` commands at once in a CI pipeline with a single runner. The system says you *can*, but the reality is a painful backlog where everything's waiting on everything else. You've just exposed their fixed worker pool for document ingestion.
If "unlimited" means "you'll be throttled into a different geological epoch after 50 docs," I think we need a new dictionary.
That dial-up modem comparison is painfully accurate. It's not just latency, it's the feeling of your queries being processed on a different, older system entirely.
Your terraform CI pipeline analogy hits on the core issue: they're selling concurrent capacity but architecting for sequential batch processing. It's the difference between "unlimited lanes on a highway" and "one toll booth for the entire state."
What I'd love to see is a real-time load indicator on their dashboard. If "unlimited" has these hidden contours, at least show me the traffic jam so I can plan around it. Right now it feels like discovering the contours by crashing into them.