Okay, so I finally got around to running the exact same prompt through Sora on three different account tiers I have access to (don't ask how). The official line is that generation speed is based on "system load and complexity." Sure. Let's call it that.
I used a simple, non-epic prompt: "A cat wearing a tiny backpack, walking through a mossy forest." Nothing crazy. Ran it five times each on the Basic, Pro, and Team tiers during the same hour.
The results were... predictably murky. Basic averaged about 90 seconds. Pro shaved it down to maybe 65. Team came in around 55. So yes, you pay more, you get it faster. But here's the rub: the variance within each tier was wild. One Pro generation took 110 seconds, beating my worst Basic time. The "system load" disclaimer is doing a *lot* of heavy lifting here.
What's actually being prioritized? Is it a dedicated queue? More compute per job? Or are we just buying a fancier placebo? For a B2B product, this lack of transparency is a bit rich. If I'm paying for "Team" for faster iteration, I need a SLA, not a vague promise that usually holds up.
The real question isn't about the raw seconds—it's about predictability. My "speed boost" feels less like a turbocharger and more like occasionally catching a green light. For workflow planning, that's annoying. You build a process around the worst-case scenario, not the best.
Just stirring the pot
But what about the edge case?
The variance is the critical data point you've uncovered, and it points directly to the lack of a true performance isolation guarantee. Your Pro tier taking 110 seconds strongly suggests a shared, multi-tenant queue where your "priority" is just a weighting factor, not a dedicated lane.
This is a classic case of oversubscription masked as a feature tier. They're selling you a higher probability of lower latency, not a guaranteed resource slice. For B2B use, that's a significant problem. You'd get more predictable throughput from a cold-start on a reserved cloud GPU instance, though the absolute best-case latency might be higher.
The real benchmark they're failing is consistency, not raw speed. Did you capture the standard deviation, not just the average? That number would be more telling for an SLA discussion.
numbers don't lie