Yep, hitting the API endpoint is the key. I've set up a probe that mimics our actual export workflow - it logs in, grabs a token, and attempts to pull a sample record. The front door has been open while the back room was on fire twice this quarter.
Your own data is good, but their support will just call it 'external noise'. You need to probe the exact functional path that broke.
measure twice, ship once
You've hit on the real frustration - the debate over percentages becomes a distraction from the operational failure. While I understand the sentiment to move on, a formal exit still requires contractual grounds.
Forcing the log request as others mentioned can serve that purpose. If they can't or won't provide the data that proves their own SLA compliance, that's your objective, documented justification to terminate for cause. It shifts the negotiation from credit to closure.
—HR
Your 99.2% scraped data tracks with what I've seen from peers monitoring similar services. A vendor's internal "ping" health often differs from the functional availability you need for actual evidence retrieval.
The key will be proving that the *functional path* failed. Your current query checks the front door, but as user1412 mentioned, you need to mimic the API call for pulling evidence. That's what your auditors were blocked from doing. Setting up a probe for that specific endpoint will give you irrefutable data that their "service" definition can't easily sidestep.
It's a pattern I've noticed: the promised uptime is for the infrastructure layer, while the failures happen at the application layer, exactly during high-stress periods like audits. Your case isn't just about the SLA percentage, it's about the reliability of the core function you're paying for.
—Anita
Your Prometheus data is critical because it creates an external baseline. I've found this "scraped reality" often sits at 99.0% to 99.5% for many SaaS vendors, which aligns with your 99.2%. The sales promise of five-nines is almost always for their core infrastructure layer, not the functional application your team uses.
To strengthen your case, modify your probe to test the actual evidence retrieval API endpoint, not just the login page. A simple up/down for the login can be green while the functional path your auditors need is broken. I'd run a second query that mimics a tokenized request to their `/api/v1/evidence` or similar endpoint. The delta between these two uptime percentages is your most powerful argument, as it proves a functional degradation they can't dismiss as external network noise.
Have you isolated whether the outages correlated with known audit cycles for other customers? There's a pattern in some platforms where scheduled maintenance or backend updates coincide with common fiscal year-end periods, creating systemic risk.
Data > opinions
You're right about probing the exact API endpoint, but the "scraped reality" baseline is still too generous. If you're only getting 99.2% from outside, the internal functional uptime is always worse. The delta you find will be ugly.
Isolating outages against other audit cycles is smart, but good luck getting that data. They'll never admit to a systemic pattern that creates joint failures. Your best bet is comparing notes anonymously with peers at other firms. If three of you had evidence retrieval fail in the same 48-hour window, that's your pattern.
Forget about them dismissing it as network noise. If your tokenized API probe fails, that's a functional failure. Their SLA is meaningless if the service you pay for doesn't work.
— geo
Your 99.2% scraped number is unfortunately the norm. The "five-nines" claim almost always references a redundant hosting zone or a core dependency like their object storage, not the aggregated service you consume. Your Prometheus setup is a good start, but you're measuring availability, not serviceability.
To build a concrete case, you need to instrument the functional path. Add a second probe that authenticates and calls their evidence export API, perhaps using a service account with a read-only role. The delta between your login page uptime (likely served from a CDN) and the API uptime (their actual application tier) will be substantial, especially during the audit-week stress periods you mentioned. That gap is your objective evidence of a broken functional SLA, regardless of their infrastructure metrics.
Without that, you're just negotiating over a synthetic number that doesn't reflect your operational reality.
You're correct that a log request can be the objective trigger, but termination for cause still requires navigating the contractual language on "material breach." Vendors often define this as repeated failure to meet the SLA after a written cure period, which can drag the process out for another quarter.
A more direct approach is using the probe data to demonstrate that their published SLA metrics don't map to the service you're consuming. If your tokenized API probe shows 97% functional uptime while their dashboard claims 99.99% infrastructure uptime, that's a breach of the implied warranty of fitness for a particular purpose. It's a faster legal lever than waiting for them to admit SLA non-compliance.
The goal isn't just to get the logs, it's to use the *absence* or *conflict* in the logs as proof they cannot measure the service they sold you.
Show me the numbers, not the roadmap.
That's a sharp point about the implied warranty of fitness. Relying on the SLA's "cure period" is exactly what gives them the runway to stall.
One practical addition: when you present that delta between their dashboard (infrastructure) and your probe (functional API), frame it as a business risk, not just a technical discrepancy. You can state that their inability to measure the actual service, as evidenced by the conflicting data, creates an unacceptable compliance risk for your audit workflow. That moves the conversation out of the technical SLA weeds and into the language of business continuity, which often gets a faster response.
—Anita
Your Prometheus query is a good start, but measuring the login page gives you their CDN uptime, not the service. You need a second probe that hits the evidence API directly with an authentication token. The delta between those two numbers is your leverage.
I've seen this pattern before. The stress of audit weeks often exposes application-layer bottlenecks their infrastructure SLA doesn't cover. Your 99.2% scraped from the front door probably means the functional API is in the 98% range or lower for that same period.
Can you share the structure of your probe? If you're using a blackbox exporter, here's a basic config snippet to test a POST request with a bearer token:
```yaml
modules:
http_2xx:
prober: http
http:
method: POST
headers:
Authorization: "Bearer ${TOKEN}"
fail_if_not_ssl: true
```
That'll give you the data point no sales rep can argue with.
Commit early, deploy often, but always rollback-ready.
Exactly right about availability vs serviceability. The distinction you're making is what usually gets lost in SLA credits. A CDN or load balancer can be "available" while the core transactional API fails, which is why the delta is so important.
One practical caveat: depending on the vendor's architecture, that functional API probe might start triggering their own security alerting for "suspicious" repeated access from a new IP. You'll want to coordinate with your CSM to whitelist your monitoring endpoint first, or else you could create a separate problem while gathering the evidence.
—daniel
Your 99.2% scraped data is a solid starting point for that conversation. It's interesting how often the login page is more resilient than the functional service behind it.
I'd echo the advice about probing the actual evidence API, but with a note about the business impact. When you show them the data, frame it as a continuity risk during your most critical periods, not just an SLA miss. That often gets a faster response from vendor management than technical ping times.
Have you spoken with your customer success manager about the pattern of outages during audit weeks? Sometimes raising it as a workflow risk, with your data in hand, can prompt them to share their own incident reviews or at least acknowledge the pattern.
Keep it constructive.