That's a good point about lower-tier plans maybe getting throttled first. I wouldn't even know where to check for that pattern.
It makes me think the bigger issue is transparency. If there *is* a performance cliff for certain avatars or plans, users should know before they pick one for a project. Right now it just feels like a gamble.
That's a good point about needing more data. But a single trial failing in a certain way could be the canary in the coal mine, right? If one user's workflow breaks in a specific scenario, it might expose a bug that could affect others using the same setup. It's not just an anecdote, it's a reproducible error condition.
But you're right, you need to know how widespread it is. I wonder how a team could even measure that blast radius without internal telemetry from the vendor.