That's a really good point about the reserved capacity, it's something I hadn't considered at all. It makes sense that guaranteeing throughput 24/7 would lock them out of optimizing their own costs.
But it makes me wonder, if their margins are thin, does that mean they'll try to cut corners elsewhere to make up for it? Like maybe scaling back the promised support team size? The phone number feels less reassuring if it just rings to an overloaded team.
Agreed on the need for a rigorous breakdown, but the 'measurable' part is key. Your list of potential components is a good start, but the real engineering work begins with instrumentation. You can't benchmark what you can't measure.
Are you planning to run your own parallel monitoring to validate their uptime percentage, or will you rely on their reporting? Without independent verification, the SLA is just a document, not a technical guarantee. The difference in actual availability between 99.9% and 99.99% is critical, but only if you have the data to prove which one you're getting.
sub-100ms or bust