Skip to content
Notifications
Clear all

First-time buyer - what SLA should I ask for in an enterprise contract?

1 Posts
1 Users
0 Reactions
25 Views
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
Topic starter   [#12391]

Having recently navigated the procurement process for a large-scale AI/ML inference service, I found the SLA terms to be the most critical and, often, the most nebulous part of enterprise negotiations. Many vendors present a standard "99.9% uptime" SLA, but this is insufficient for production data pipelines where latency, throughput, and consistency are paramount. Based on our team's experience integrating a similar service into our real-time analytics platform, I recommend focusing on the following specific SLA components beyond mere availability.

**Core Technical SLAs to Demand:**
* **Availability:** 99.9% is table stakes. Push for 99.95% or higher, with explicit definitions of "downtime." It must exclude your scheduled maintenance windows and clearly defined force majeure events.
* **Latency P99:** This is non-negotiable. You need a guarantee on the 99th percentile response time for a defined payload size (e.g., "P99 latency < 2s for a 1KB prompt under 1000 requests per second"). Without this, your downstream data consumers (dashboards, applications) will face unpredictable slowdowns.
```json
// Example SLA clause structure
{
"metric": "end_to_end_latency",
"threshold": "2000ms",
"percentile": 99,
"measurement_window": "5 minutes",
"exclusions": ["known_bugs", "customer-side network issues"]
}
```
* **Throughput:** Commit to a minimum sustained Requests Per Second (RPS) per model or endpoint. Ensure there are defined ramp-up periods and costs for bursting.
* **Error Rate:** A simple "successful requests" metric isn't enough. Define acceptable thresholds for both HTTP 5xx errors (e.g., < 0.1%) and model-specific errors (e.g., content filtering triggers).

**Commercial & Operational Protections:**
* **Credits:** The SLA credit must have teeth. A 10% service credit for missing the availability SLA is common, but aim for a sliding scale where more significant breaches trigger higher credits or even contract termination rights.
* **Reporting & Transparency:** Require a real-time, customer-accessible dashboard for all SLA metrics and a monthly detailed report. The burden of proof for SLA misses should not be on you.
* **Support Response Times:** Tiered severity levels (Sev-1 for complete outage, Sev-2 for degraded performance) with concrete response and resolution time commitments (e.g., "Sev-1 initial response within 15 minutes, 24/7").

The key is to bind the SLA to your actual use case. If you're using the service for batch processing of overnight ETL jobs, latency SLAs are less critical than batch job completion windows. For real-time feature generation, latency and error rate are everything. Never accept an SLA that only covers the API gateway; it must cover the entire inference path.

Our negotiation resulted in a 30% higher credit rate and the inclusion of a P99.5 latency metric after we demonstrated how tail latencies were causing cascading failures in our pipeline. What specific metrics and penalties have others here successfully codified in their contracts?

--DC


data is the product


   
Quote