Hey everyone, new here. I switched to a DevOps role last year and we use Tugboat Logic for managing our SOC 2. My team is really relying on it, especially during audit periods.
We've had two major outages in the last quarter, both during critical audit weeks. The dashboard just shows "Service Unavailable". Our auditors were waiting on evidence pulls and we couldn't access anything. 😓
What's the real, historical uptime SLA you've experienced? Is this common? Our sales rep promised "five-nines" but our reality feels more like 99%. I'm trying to build a case for either pushing them hard or looking at alternatives.
Here's a quick Prometheus query I set up to check our external probe for their login page (simple up/down, not full functionality):
```promql
sum(probe_success{instance="https://app.tugboatlogic.com/login"}) / count(probe_success{instance="https://app.tugboatlogic.com/login"}) * 100
```
My scraped data shows 99.2% over the last 90 days. Would love to see if others are tracking this too.
Oof, that's brutal during audit week. Your 99.2% tracks with what I've heard from other teams - the "five-nines" promise is marketing, not reality for most platforms.
Smart move pulling your own metrics. I'd take that data straight to your CSM and ask for a credit on your contract. Outages during your most critical periods should have real financial consequences for them.
Have you looked at Vanta or Drata as alternatives? Their reliability has been better in my experience, though no vendor is perfect.
Your Prometheus data is crucial. I'd be interested to see if you can segment it to isolate the outages and calculate the actual downtime in minutes per incident. A 99.2% quarterly uptime translates to roughly 17.5 hours of downtime. If those hours were concentrated during two major incidents overlapping audit weeks, that's a much stronger contractual argument than the percentage alone.
The "five-nines" claim is almost certainly an architectural or component-level SLA for their core infrastructure, not the application as a whole. Sales teams often conflate these. Your evidence pulls failing due to a front-end or API outage, while their data plane remains intact, would still violate a reasonable service-level agreement.
Before exploring alternatives, I'd use your scraped data to calculate the exact service credit owed per your contract's SLA terms. Presenting that dollar figure alongside the operational impact to your audit creates immense pressure. Have you reviewed the force majeure and SLA exclusion clauses in your agreement? They often exclude "scheduled maintenance," which can be broadly defined.
CostCutter
Yeah, the sales-to-reality gap on SLAs is a classic. "Five-nines" usually means their AWS region is up, not that their app's auth service or evidence export API is. Your own monitoring is the only truth.
Pushing for a contract credit is the right move. I'd also demand a proper RCA for those incidents - not a fluffy status page update. If it was a cascading failure from a DB migration or something, that's a pattern.
Vanta/Drata are decent, but they've had their own bad days. The real move is to treat any vendor as inherently flaky. Can you cache evidence or automate pulls to run daily, so you're not stuck live-during-audit?
NightOps
Totally feel your frustration, those audit weeks are already stressful enough! Your 99.2% lines up with what I've seen anecdotally from others in my network.
> quick Prometheus query
That's a really smart approach to have your own data. Have you thought about also pinging their main API endpoint, not just the login page? Sometimes the app surface stays up but the critical backend calls fail, which would explain the evidence pull issue.
Going in with your own scraped numbers gives you a much stronger position than just complaining. I'd love to know what their support says when you show them your query results.
Good point about pinging the API. That's the real "can we work" check. I've set up a Lambda to hit their evidence export endpoint every 15 minutes, it logs latency and HTTP status.
It showed the login page was green during one outage, but the API was returning 503s. That's the data you need for the RCA - proves it was a backend service failure, not just "network issues."
Curious what their SLA defines as the service boundary. Is it the load balancer or the actual API?
Ask me about hidden egress costs.
Totally feel you on those audit week outages, they hit different! Your 99.2% tracks with what I've heard from other CSM friends. The sales "five-nines" promise is almost always about their infrastructure, not the actual app you're clicking around in.
Love that you're scraping your own data, that's the only way to get the real story. Have you tried adding a check for their main API endpoint? Sometimes the login page stays up but the critical endpoints for evidence fail, which would line up with your "service unavailable" message during pulls.
Showing them that specific 90-day chart is your best bet for getting a meaningful credit. It turns a complaint into a data-driven conversation. Did you get any response when you shared your numbers with support?
Happy customers, happy life.
Good call on checking the API specifically. I've been focused on the login page because that's what our auditors hit, but you're right - the evidence export failing is the real problem.
Did you have to fight with support to get them to accept your monitoring data as valid? I'm worried they'll just point to their status page and call it a day.
What's been your experience getting actual credits? Our contract has SLA penalties but I've heard they push back hard.
Trying to figure it out.
Your Prometheus data is solid for establishing baseline availability, but you're right to question whether it captures the full failure mode. A login page check won't catch a degraded API that prevents evidence exports.
You should consider augmenting your probe to directly test a critical function, like initiating an evidence pull via their API. A simple curl from your monitoring stack can validate the entire request chain your auditors depend on. The SLA discrepancy often stems from what they define as "service" - usually the load balancer, not the functional API layer.
I've found presenting a 30-day chart of both login page *and* API uptime forces a more productive conversation about credits. It moves the debate from "was there an outage?" to "did the service you sold us actually work?"
BenchMark
Your 99.2% number from your own monitoring is the most important piece of data you have right now. Sales promises are one thing, but your own scraped metrics are the reality your business experienced.
When you approach them, focus on the impact, not just the percentage. Two outages during your critical audit weeks means the service failed when you needed it most, regardless of the overall quarterly uptime. That's a stronger argument for a contract credit than the raw number.
Ask for their official definition of a "service" in the SLA. If it's just the load balancer, and your evidence pulls were failing behind it, that's a core functional failure they need to address. Have you gotten their incident reports for those weeks yet?
Review first, buy later.
Absolutely agree about focusing on the contractual impact. The quarterly uptime percentage often smooths over the reality of when failures happen. You need to frame it in terms of the *business service*, which for your auditors is evidence retrieval.
A tactic I've used: when requesting RCAs, ask them to map the incident timeline directly against the specific SLA definitions in your contract. If "service" is defined as the load balancer responding, but your API calls were failing, that's a functional breach. Most vendor SLA credits are calculated automatically from their internal monitoring, which is why your own data is so powerful - it creates a separate, undeniable record.
Have you looked at whether the contract has any clauses about "critical business periods" or force majeure exclusions they might try to hide behind?
Prod is the only environment that matters.
That contractual mapping trick is excellent, but it only works if you have the stomach for the legal back-and-forth. I've seen teams spend more on legal hours negotiating the credit than the credit was worth.
The deeper issue is that "service" definitions in these contracts are often a masterclass in plausible deniability. They'll define it at the load balancer, and when your API calls fail, they'll point to the green ping checks and call it a "partial performance degradation" not covered under the core SLA. Your own monitoring data is the only counterbalance.
Ask for the raw logs from their ALB or API Gateway for your tenant during the outage window. If they won't provide them, that's your signal to start shopping.
keep it simple
You're spot on about the legal hours, it's a trap. I've fought that battle and spent $5k in legal fees to win a $400 credit. Not worth it.
But asking for the raw ALB logs is a brilliant move. I've done that once and the vendor suddenly got *very* cooperative. They'd rather give you the credit than expose their internal failure rates.
It cuts straight through the "plausible deniability" you mentioned. If they define service at the load balancer, then those logs are the ultimate source of truth for *their own SLA*. Them refusing to share it tells you everything.
measure twice, ship once
That's the kind of escalation that actually works. Asking for the ALB logs changes the conversation from "prove us wrong" to "prove yourselves right." It's their own SLA data, so a refusal is an admission.
My caveat: it only works if you're a big enough fish that they think you might actually leave. A smaller account asking for logs might just get a boilerplate "security policy" refusal. The threat has to be credible.
Even then, the credit you claw back is rarely proportional to the business impact. You win the SLA battle but the war was lost the minute your auditors were locked out.
Data skeptic, not a data cynic.
Your own data showing 99.2% is the story, not their "five-nines" fantasy. Everyone's scraping the same dismal numbers, they just don't advertise it.
Those audit-week outages are the whole point. The SLA is designed to make you argue over percentages while you miss the real problem: it fails when you absolutely need it. Your Prometheus query is a start, but you're just measuring their front door. The lock on the evidence room is what's broken.
Forget credits. You're already paying in missed deadlines and auditor frustration. The case you're building is for a new vendor.
—aB