Skip to content
Notifications
Clear all

What's the real uptime SLA? We've had two outages during critical audit weeks.

2 Posts
2 Users
0 Reactions
0 Views
(@grafana_guy_night)
Reputable Member
Joined: 4 months ago
Posts: 170
Topic starter   [#22859]

Hey everyone, new here. I switched to a DevOps role last year and we use Tugboat Logic for managing our SOC 2. My team is really relying on it, especially during audit periods.

We've had two major outages in the last quarter, both during critical audit weeks. The dashboard just shows "Service Unavailable". Our auditors were waiting on evidence pulls and we couldn't access anything. 😓

What's the real, historical uptime SLA you've experienced? Is this common? Our sales rep promised "five-nines" but our reality feels more like 99%. I'm trying to build a case for either pushing them hard or looking at alternatives.

Here's a quick Prometheus query I set up to check our external probe for their login page (simple up/down, not full functionality):
```promql
sum(probe_success{instance="https://app.tugboatlogic.com/login"}) / count(probe_success{instance="https://app.tugboatlogic.com/login"}) * 100
```
My scraped data shows 99.2% over the last 90 days. Would love to see if others are tracking this too.



   
Quote
(@chloeh)
Estimable Member
Joined: 2 weeks ago
Posts: 68
 

Oof, that's brutal during audit week. Your 99.2% tracks with what I've heard from other teams - the "five-nines" promise is marketing, not reality for most platforms.

Smart move pulling your own metrics. I'd take that data straight to your CSM and ask for a credit on your contract. Outages during your most critical periods should have real financial consequences for them.

Have you looked at Vanta or Drata as alternatives? Their reliability has been better in my experience, though no vendor is perfect.



   
ReplyQuote