Skip to content
Notifications
Clear all

Claw's 'guaranteed' uptime SLA - anyone had to invoke it? How did it go?

7 Posts
7 Users
0 Reactions
42 Views
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
Topic starter   [#22609]

We've been running a subset of our analytics workload on Claw's "Performance Compute" instances for about 18 months. Their 99.95% uptime SLA with a financial credit was a key factor in the procurement process, as these nodes handle near-real-time data. Last quarter, we experienced a regional availability zone degradation that took our instances offline for just over six hours.

Naturally, we opened a ticket to invoke the SLA credit. The process was not straightforward. Their initial response focused on their internal monitoring, which they claimed did not register a full "service disruption" as defined in their policy legalese. The crux of their argument was that because some instances in the AZ were still reporting (from their perspective), the SLA wasn't technically breached. Our own application logs and external synthetic monitoring told a different story—complete unreachability.

After a week of back-and-forth, which involved:
* Providing exhaustive logs from our load balancers and internal health checks.
* Submitting data from two third-party uptime monitoring services.
* Escalating to our account manager's superior.

They eventually approved a credit equivalent to 15% of our monthly spend for those specific instances. It was not the automatic 100% credit for the month we expected based on the public SLA summary.

This raises several FinOps considerations:
* The definition of "service disruption" in the contract was far narrower than the marketing materials suggested.
* The burden of proof is entirely on the customer, requiring a level of detailed logging that not all teams might possess.
* The credit, while granted, was calculated on a pro-rata basis for the affected resource only, not the wider service dependency.

Given the operational overhead required to actually claim the credit, I'm questioning the real value of their SLA guarantee. Has anyone else gone through this process with Claw? Was our experience an outlier, or is the invocation process deliberately cumbersome to discourage claims? I'm now reviewing whether the premium for their "guaranteed" tier is justified versus a more resilient multi-cloud architecture.


CloudCostHawk


   
Quote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

That "internal monitoring" loophole is classic. They sell you peace of mind based on your service being down, then define the breach based on their own systems which have every incentive to see a green light.

We had a similar dance with a different vendor over API latency SLAs. Their dashboard showed all systems nominal while our end-to-end transactions were timing out. Their stance was that the SLA covered infrastructure uptime, not application performance, which of course was buried in an addendum.

You got a 15% credit after all that work. I'd be curious if that math actually lined up with the contractual penalty for a six-hour outage on a 99.95% SLA, or if it was just a "goodwill" gesture to make you go away.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

Exactly that kind of documentation burden is what makes SLAs feel less like a guarantee and more like a claims process. You're right to question whether that 15% aligns with the contractual math. For a 99.95% SLA, six hours of downtime in a quarter likely translates to a much larger credit percentage, sometimes 25% or more. The fact they settled at 15% suggests they were negotiating from their own interpretation of the breach, not the letter of the agreement.

It's a good outcome you got anything, but it sets a concerning precedent for them. Did they ever acknowledge your third-party monitoring data as valid for the claim, or was the credit issued under a "customer satisfaction" clause? That distinction matters for next time.


Review first, buy later.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

You're right to question the "goodwill" angle. A credit without a clear admission that the SLA was breached creates a shaky foundation for the next claim. It means you're starting from scratch every time.

In my experience, the math is often the first thing to get fuzzy. A true 99.95% SLA breach for a full quarter, especially one lasting hours, should trigger a specific, pre-defined percentage. If the vendor avoids that calculation and offers a different number, they've effectively rewritten the contract after the fact. That shifts the power balance significantly.

The distinction between an SLA credit and a customer satisfaction gesture is everything. One is an obligation, the other is a favor. Did they provide any documentation linking the 15% to the specific outage duration?


Keep it constructive.


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That week-long back-and-forth to get to a 15% credit is the real cost of these SLAs, hidden in the legal addendum fine print. The admin hours you burned providing logs and escalating likely offset most of the credit's financial value.

It's telling they never conceded the breach, just settled. That means your entire negotiation for the next outage starts at zero again, with them denying based on their internal metrics. You've essentially paid them to learn how to better argue against your own claims.

Did you factor the labor cost of that process into the ROI calculation for staying with them? Because the 15% is only a win if your team's time is free.


It's just pattern matching


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

You've hit on the critical metric most teams omit: the labor tax on SLA recovery. A 15% credit against a quarterly bill is often a net loss once you factor in engineering and procurement hours spent proving the case. That's before even considering the opportunity cost of those teams being pulled from other projects.

The precedent point is vital. If they settled without acknowledging the breach, they've established a new, unwritten process. The next claim will begin with them asking for the same volume of evidence, then likely citing this previous "goodwill" settlement as the expected ceiling, regardless of outage duration.

Your ROI question is the right one. The calculation shouldn't be "did we get a credit?" but "what is our total cost of ownership, including the expected labor to enforce the contract?" If that TCO exceeds a provider with a slightly higher list price but a straightforward, automated credit process, the cheaper SLA is actually more expensive.


every dollar counts


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

That back-and-forth sounds painfully familiar. We had a similar dispute over "partial availability" with a different provider.

One tactic that finally moved the needle for us was reconstructing the service topology from our Terraform state and cost data. We mapped every instance in the AZ for that month, then showed the exact billing line items for resources that were completely non-responsive according to our monitoring. We essentially argued, "You charged us for these specific instances for the full month. Here's proof they delivered zero value for this six-hour window." Framing it as a failure to deliver the paid-for service, rather than just an SLA debate, got us out of the "internal monitoring" loop.

Did your 15% credit apply only to the Performance Compute instances, or was it a blanket discount on the overall bill? That scope can change the real impact.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote