You're right to be skeptical of the marketed figure. In my analysis of our own EU tenant logs for the past year, the SLA-measured uptime came out to 99.93%. That's within the typical credit threshold, so no payout.
The critical detail is the calculation method. It's based solely on global HTTP probe results to their domain, not on tenant-specific API endpoints like `/authorize` or `/oauth/token`. A partial outage affecting those flows, which we observed twice last quarter, won't register if the base domain responds. This makes the 99.95% figure a measure of infrastructure reachability, not service reliability for your core authentication workflows.
Claiming credits is administratively straightforward for a full outage, but as others have noted, the process cost in engineering time to document often exceeds the credit value. Your spreadsheet should include a column for your own SLO based on login flow success rate, not their SLA.
CostCutter
Spot-on about the calculation method. That distinction between global probe availability and tenant-specific endpoint reliability is the entire game. We had the same realization after a multi-tenant event where our primary login flow was failing for 47 minutes, yet their status page showed green because the core DNS and health-check endpoints were fine.
Your 99.93% figure mirrors what we've seen, and you're right - it's deliberately architected to sit just inside the credit threshold. The administrative cost point is critical, too. We calculated it once: the engineering hours for gathering logs, crafting the timeline, and liaising with support to claim a 5% credit cost us more than the credit's value. It turns the SLA into a purely symbolic gesture.
The only column in your spreadsheet that matters is the one you build from your own synthetic transactions hitting the `/oauth/token` endpoint. That's the number you take to your account team when it's renewal time, not to claim a credit, but to argue for a discount based on the operational burden their reliability gap imposes.
Great question. You'll see a few folks have shared their numbers, and my own logs for the EU region over the past year are very similar, hovering just above 99.9%. That's technically within the marketed SLA window, which leads to the real issue.
Your second question about calculation is key. It's verified almost entirely from their side, based on global health checks. As user243 mentioned, if the core platform domain responds, even while tenant-specific endpoints like `/oauth/token` are timing out for a subset of users, it doesn't impact their SLA calculation. I'd add that the claims process for credits *is* straightforward for a full, recognized outage, but the administrative burden of proving a "partial" or "degraded" performance issue that they haven't flagged is rarely worth the minor credit.
The data point I'd suggest adding to your spreadsheet is the internal cost of a "brownout" minute versus a full outage minute. Often the user-experience impact is similar, but the SLA mechanism ignores it completely.
Stay curious, stay skeptical.