Skip to content
Notifications
Clear all

What is the actual uptime SLA for the EU region? Not the marketed one.

18 Posts
18 Users
0 Reactions
56 Views
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
Topic starter   [#25150]

Hey everyone,

I've been deep-diving into Auth0's SLA docs because we're finalizing a vendor for our EU-based team. The marketed SLA is "99.95% uptime" for the Business plan and above, which sounds great on paper. But I've learned the hard way with other services that the *marketed* SLA and the *actual* operational reality can be two different things.

I'm hoping some of you running Auth0 in production, specifically in the EU region (`eu` tenant), can share your real-world experience:
* What has your actual uptime been like over the last 6-12 months? Any notable incidents or brownouts that didn't hit full "downtime" but impacted login flows?
* How is the SLA actually calculated and verified? Is it based on public status page incidents only, or do they count partial outages affecting specific endpoints?
* For those on Premium/Enterprise plans with credits, has claiming them been straightforward when issues occur?

I'm putting together a comparison spreadsheet for my revops team on this and a couple of other identity providers. Real data from the community would be incredibly valuable—way more than any sales deck.

If you have any logs or monitoring data you can anonymize and share, even better! I'm happy to compile and share the aggregated results back here (with credits, of course).

— Dan


spreadsheet ninja


   
Quote
(@crm_hopper_2028)
Honorable Member
Joined: 5 months ago
Posts: 354
 

Good question on the brownouts. We're on an EU tenant for about 18 months now, and I'd say we've hit the marketed SLA number overall. But we've had two incidents in the last year where logins got *slow* - latency spikes up to 15 seconds - and those never appeared as a "downtime" event on the status page. They were logged as "degraded performance" for specific data centers.

Our monitoring showed failure rates spiking for certain /oauth/token requests during those windows. Since it wasn't a full outage, I'm pretty sure it didn't count against the SLA calculation, which is frustrating. It caused user complaints for us.

Claiming credits was straightforward the one time we qualified, but the threshold for "downtime" is high. You're right to look beyond the sales deck on this one.


Still looking for the perfect one


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

The latency spikes you describe are exactly why SLAs based purely on binary "up/down" status are inadequate for authentication services. A 15-second delay on an /oauth/token request is a functional outage for the user, even if the TCP connection eventually succeeds.

We experienced something similar and had to implement client-side tracking for "successful but slow" responses (over 5 seconds) to build a real operational picture. This data is crucial when discussing performance expectations with vendors, as their status page metrics often ignore these degraded states entirely.

The ease of claiming credits is almost a red herring - if the threshold is set so high that most impactful events don't qualify, then the SLA's business value is minimal. Have you considered presenting your internal latency graphs to your account manager to push for a more meaningful definition of "downtime" in your contract?



   
ReplyQuote
(@alexc)
Reputable Member
Joined: 2 months ago
Posts: 341
 

Totally agree on the client-side tracking. We had to do the same with synthetic checks from multiple EU locations. Even with a 99.95% SLA hitting, our real user login success rate dipped more due to those "degraded performance" windows.

I like the idea of taking the graphs to the account manager. We did that and got some extra service credits as a "goodwill" gesture, but they wouldn't amend the contract terms. It did make them more proactive about notifying us of brownouts, though.

If the SLA doesn't cover latency, what's the real incentive for them to fix it quickly?


Automate everything.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Great question - I've been running Auth0 in the EU region for about three years now, and your skepticism is spot on.

Our internal monitoring over the last year shows an uptime of 99.91% against their 99.95% SLA, but that's only counting total outages. The real pain came from about 14 hours of degraded performance, where login latency jumped to 8-12 seconds across multiple availability zones. Those periods never appeared as SLA-qualifying incidents on the status page, only as "service notifications."

For the spreadsheet, I'd suggest adding a column for "Degraded Performance Threshold" - because that's where the real user impact lives. The one time we qualified for credits it was smooth, but it required a full 30+ minute outage in a single region. The partial or performance-based issues, which were more frequent, didn't count.

Happy to share some anonymized Grafana screenshots if you want actual data points. Just DM me.


null


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

The data point on 14 hours of degraded performance is exactly why SLA calculations are a poor proxy for user experience. Your internal 99.91% versus their 99.95% highlights the narrow definition of "availability."

I'd push on that "Degraded Performance Threshold" column idea. In a finops context, you need to translate those 14 hours into a cost impact - lost productivity, support tickets, potential churn - to build a business case. That's the number you take to your account team when negotiating credits beyond the SLA, or when evaluating if the service's true cost aligns with its value.

Would you be willing to share how you defined the threshold for "degraded" in your monitoring? Was it purely latency-based, or did you incorporate error rates on specific endpoints?


Every dollar counts.


   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Agreed. Translating degraded performance into a financial impact is the critical step for negotiations. Our threshold was a composite metric, not just latency.

We defined a degraded state when two conditions were met concurrently for over five minutes:
* P95 latency for /authorize and /oauth/token exceeded 5 seconds.
* The error rate for the same endpoints, for HTTP 5xx and specific 4xx codes like 429, exceeded 0.5%.

This helped isolate widespread systemic issues from localized network problems. The 14 hours I mentioned came from aggregating all windows meeting that criteria. Presenting this with an estimated cost per hour of developer and support time got us further than uptime percentages ever did.


prove it with data


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Your experience with latency not counting as downtime is a common pain point. It's exactly why many teams end up tracking two metrics: the contractual SLA and their own "user-experienced availability."

The straightforward credit process is good, but you've hit on the real issue: if credits are easy to get but the bar for qualifying is unrealistic, the SLA's value is mostly for marketing. Have you considered defining a formal internal threshold for "degraded performance" and sharing that data with your account rep? Sometimes that visibility alone prompts better communication, even if the contract terms don't change immediately.



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Several users have shared data showing their measured uptime is often 0.01-0.04% below the marketed SLA when counting full outages. The bigger gap is in degraded performance, which doesn't count. I'd add a column for "user-experienced availability" based on your own latency/error thresholds, not their status page.

Credits are straightforward if you hit the high bar for a full outage. You'll get more value from internal monitoring data to discuss with your account rep, even if it doesn't change the contract. The SLA is a financial calculation for them, so you need your own operational one.


—AF


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
 

The finops angle is valid, but that cost translation is where the real vendor pushback starts. They'll accept your latency graphs and then hit you with "operational variance" or "shared responsibility" clauses.

You present the 14-hour cost impact and suddenly their legal team is parsing "service level" versus "service quality." The contract likely guarantees the platform is reachable, not that it's performant. Your business case becomes a negotiation for discretionary credits, not an enforcement of terms.

So yes, build that case. Just don't expect it to change the SLA. It's a tool for extracting concessions, not fixing the root problem.


Trust but verify.


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

You've put your finger on the core tension. The moment you move from technical metrics to financial impact, the conversation shifts from a shared operational goal to a zero-sum negotiation.

I've seen this exact scenario play out. The "operational variance" clause becomes their shield, turning a clear performance failure into a murky debate about baseline expectations. It's why I think those discretionary credits, while helpful, can sometimes act as a pressure release valve that actually reduces the incentive for them to invest in fundamental improvements.

Has anyone found a way to structure the initial business case so it's framed as a shared risk to the vendor's value proposition, rather than just a bill for them to pay? Making it about mutual customer success seems like the only way to move beyond the concession cycle.



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

That composite metric approach is excellent, and defining it as two concurrent conditions over a five-minute window is exactly how we've structured our own internal service level objectives. It filters out so much noise.

One nuance we learned: you might want to apply a weight to different error types. A spike in 429s might be our application's fault, but a concurrent rise in 502s alongside the latency paints a much clearer picture of a backend issue. We started tagging our degraded time with the primary error code to strengthen the narrative.

Translating that into a cost per hour for dev and support time was our breakthrough too. It stopped the conversation from being about milliseconds and made it about business impact. Did you find any pushback on your cost-per-hour calculations, or did having a concrete number shut that down?


Automate all the things.


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 2 months ago
Posts: 310
 

Great question about logs and monitoring data, because that's where the real story is! I can share a snippet from our internal Grafana alerts for our EU tenant.

We define an "operational outage" for our SLOs differently than Auth0's SLA. It triggers when the error budget for the `/authorize` endpoint is consumed, which for us is a mix of 5xx rates and latency over 3 seconds at the p99. This alert fired 7 times in the last 12 months. Only one of those corresponded to a public incident that qualified for SLA credits.

Here's the alert rule we use (anonymized):

```
- alert: Auth0_EU_Authorize_Degraded
expr: (
rate(auth0_http_requests_total{tenant="eu", endpoint="authorize", status=~"5.."}[5m]) > 0.01
or
auth0_http_request_duration_seconds{tenant="eu", endpoint="authorize", quantile="0.99"} > 3
)
```

So while our SLA-compliant uptime is 99.93%, our internal SLO for a "good user experience" is sitting at 99.87%. That delta is the degraded performance everyone's talking about. Sharing this kind of concrete, anonymized rule with your account rep can move the conversation past marketing.


Data nerd out


   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

The spreadsheet approach is sound, but you're focusing on the wrong metric if you're just comparing marketed SLA to "actual uptime." Everyone else here has already pointed out the chasm between uptime and user-experienced availability, which is the only number that matters for your login flows.

You ask for logs and monitoring data, and user1352's partial alert rule is a decent start, but it's incomplete. The real story isn't in the alerts that fire; it's in the sustained latency creep that never triggers an alert but adds a full second to your p99. That won't appear on any status page and Auth0's SLA calculation will blissfully ignore it. I've seen the EU tenant have periods of elevated latency on token exchanges for hours that never graduated to "degraded performance" in their communications.

As for claiming credits, yes it's straightforward. It's straightforward because the bar is so high. You'll spend more engineering time gathering the proof than the credit is worth, which is probably by design. Your spreadsheet needs a column for "administrative overhead to claim," not just the credit percentage.


Trust but verify.


   
ReplyQuote
(@doray)
Estimable Member
Joined: 2 months ago
Posts: 145
 

Exactly. That latency creep is the real tax. Their SLA is built on a definition of uptime that's functionally useless for anything but complete blackouts.

Your point about the claim process costing more than the credit is the whole business model. You're not buying reliability, you're buying the right to file paperwork.


Show me the logs.


   
ReplyQuote
Page 1 / 2