I keep seeing vendors tout their "robust" 99.5% uptime SLA as a major selling point. Let's do the basic math: 99.5% uptime means 0.5% downtime. Over a year (8,760 hours), that's 43.8 hours of acceptable, non-refundable downtime.
That's over a full business week where the system can be down without you getting a penny back in service credits. For a sales or support team, that's catastrophic.
The real contract trap isn't the percentage itself—it's what's attached:
* The credit you get for missing this SLA is usually a paltry percentage of your monthly fee, not your annual contract value.
* Scheduled maintenance is almost always excluded, and they define what "scheduled" means.
* "Downtime" often only counts if it's a complete outage. Severe performance degradation, where the UI is unusable but technically "up," doesn't count.
* You have to submit a claim within a narrow window, often 30 days, or you forfeit it.
I was reviewing a contract last week where the SLA had a 99.5% guarantee, but the credit calculation capped at 10% of one month's fee. For a $50k/year contract, that's maybe $400 back for a 40-hour outage. The cost of lost productivity dwarfs that.
Before you sign, check the actual SLA exhibit. Calculate the annual allowable downtime. Then read the credit schedule and the definitions of "excluded downtime." That's where the joke is really told.
-- CRM Surfer
Your CRM is lying to you.
You're hitting on the most important part. The 43 hours is just the shiny object they want you to focus on. The real magic trick is in the definition of an "hour of downtime."
I've seen contracts where the clock only starts after you open a severity-one ticket and they confirm it's their fault. So a 20-minute blip at 2 AM? That's zero SLA hours. They need 60 consecutive minutes of "confirmed, complete outage" to even log one unit against their guarantee. Suddenly those 43 annual hours become almost impossible to actually accrue.
It's a billing construct, not an engineering promise.
If it's free, you're the product. If it's expensive, you're still the product.
Exactly. That "confirmed, complete outage" clause is where the whole promise unravels. I had to shepherd our legal team through this last year. We pushed back and got them to include language for "material degradation" impacting core workflows, but it was a fight.
Even then, the burden of proof is on you. You need monitoring that matches their definition and the staff to compile reports. Most teams just don't have the cycles to play SLA accountant, so the credits go unclaimed.
ian
Oh wow, that math is eye-opening. A full business week of allowed downtime is crazy when you put it that way.
And you're so right about the credit being a tiny slice of the monthly fee. It's like they're insuring a house for the cost of a window. Has anyone ever actually found a vendor that offers credits based on the annual contract value instead?
That breakdown of the credit cap is really helpful. I'm new to evaluating these contracts and I hadn't thought about the 10% of one month being so small. Is it common to get that clause changed to a higher percentage or a different calculation, or do most vendors just refuse to budge?
Yeah, that credit cap is wild. I once saw a contract where the monthly fee was just $100, but the annual commitment was way higher. The credit was capped on that tiny monthly fee, so even a major breach paid back basically nothing. Is there any standard for how they pick that monthly reference amount?
Containers are magic, but I want to know how the magic works.
Absolutely. That "confirmed, complete outage" clause is the killer. Even if your internal monitoring screams it's down, the SLA clock hasn't started ticking until they say so. You're basically waiting on their support queue during the actual outage.
It gets worse when the platform is a bit flaky but not fully dead. Users are screaming, but because the login page still loads, it doesn't count. Those partial performance hits are where real productivity dies.
I've found the only way around it is to bake a specific, objective performance metric (like API response time >5 seconds) into the SLA definition itself. Otherwise, you're right - it's just a billing construct.
Trust the trial period.
Yep, that monthly fee trick is the real sleight of hand. You'll also see them bury a rolling measurement window in the fine print. So a 40-hour outage could happen over 60 days and they'd claim each month hit, say, 99% and therefore "no breach".
Forget the SLA number. Fight to define downtime by your own HTTP probe failing from three external regions. That's the only metric that matters. Good luck getting them to agree though.
You nailed the rolling window trick. They slice and average until the breach disappears.
I've seen them measure over a 30-day period, resetting at the start of each calendar month. A 43-hour outage straddling the end of one month and start of another gets halved, making each month's percentage look passable. Suddenly, no credits.
Pushing your own HTTP probe from three regions is the right goal, but you'll never get it. Their counter is always "our internal monitoring is the source of truth." The real fight is over the *frequency* and *location* of their own probes. Getting them to commit to probing from more than one cloud region is a small win.
show me the bill