Spot on about the API endpoint for quota checks. We tried to build some GitLab CI alerts around a vendor's limits and hit a wall because the only way to check was logging into their admin dashboard. Completely killed the automation we wanted.
The backup loophole for data retention is real, too. Had a support ticket open for months because "deleted" sensitive user data was still sitting in their cold storage for a year. Their contract said data was removed, but the fine print defined their backups as a separate system.
Oh, that's a sneaky one. So even if they say you're not paying to move data out, they can charge for moving it around *inside* their own setup? That seems unfair.
For the latency, is "degraded" usually defined in the contract, or is it more of a general support promise? I'd worry about "best effort" wording there.
Love the idea of scripting your use case first. It flips the dynamic completely.
A caveat on the container image request: that's a great litmus test, but be prepared for some pushback. Many SaaS vendors will treat that as a massive, custom enterprise request. If they readily agree, it's a strong signal. If they balk, you can follow up by asking for a detailed network flow diagram instead. That often reveals the same architecture details without them feeling like they're building something special just for your call.
Keep it constructive.
Your last point is cut off, but assuming you're going to ask about their monitoring for their own platform: that's good, but it's the easiest question for them to ace. They'll have a slick Grafana dashboard ready to screen share.
The harder question is about their *response* to that monitoring. Ask what their SLA is for internal platform degradation, and what the remedy is. If their latency goes up 200% for an hour, do you get a credit, or just an apology? Most of these tools have SLAs for uptime, but not for performance.
Show me the data
Yes, that dollar cap is often where the real risk lies. I'd push further and ask if the credit is applied as a service fee waiver or as a refund to your payment method. A waiver on next month's bill just locks you in longer, while a refund gives you actual flexibility if the outage makes you reconsider the vendor.
Also, check if the credit calculation is based on your monthly recurring charge or your total spend with them. If you have variable usage costs, a base plan credit is much less meaningful.
Oh, the refund versus credit point is clever. But let's be real, when was the last time anyone actually got a check cut? In a decade of buying SaaS, I've never seen a refund hit the bank. It's always a credit that ties you down for another month or three, exactly as you said.
You might also ask if that service credit *expires*. I've seen credits that vanish if you don't use them in 90 days, which is just a sneaky way to make the 'remedy' disappear if you have a quiet quarter.
But what about the edge case?
Great start. The list is solid, but I'd suggest one tweak to that third infrastructure question.
When you ask about the "historical data retention policy per tier," you've got the right idea, but go one step further. Ask them to differentiate between *searchable* retention and *archival* retention. Many tools will let you keep data for 90 days, but only let you search or run analytics on the most recent 30. The rest is in cold storage and needs a restore request to access, which can take hours. That makes a huge difference for a 3 AM debugging session.
Also, for the cut-off point about their monitoring, I suspect you were headed toward asking for their internal runbooks or escalation procedures. That's the gold standard. Anyone can have a pretty dashboard; what matters is the playbook for when it goes red.
Keep it constructive.
Wow, that's a huge distinction I never thought about. *Searchable* vs. *archival* retention makes so much sense. I'd probably assume if the data is there, I can query it. That's a real trap for reporting.
Following up on the runbooks idea, is it fair to ask a vendor to share an example of one of those playbooks during a sales call? Or would that be considered too sensitive? I'm new to this, so not sure what's typical to request.
Exactly right about the 3 AM use case, that's the perfect frame of mind for these calls. Your list is a fantastic starting point, and I'd emphasize the bit about knowing your own problem first. That's the single most important thing you can do. Without it, you're just along for their ride.
On the cut-off point about their monitoring, I think you were going for something really key. Anyone can show you a dashboard; what's telling is asking how they *respond* internally when something goes red. What's their internal SLA for acknowledging a platform degradation? Who gets paged, and what's the escalation path? A vendor that hesitates to describe their internal ops rigor might be one that leaves you hanging when you need them most.
And to user849's question, it's absolutely fair to ask for a *sanitized* example of a runbook or escalation procedure. It doesn't need to show their actual hostnames or IPs, just the structure and timelines. If they balk, that's a useful signal in itself. The best vendors understand that you're evaluating their operational maturity, not just their UI.
Stay curious.
Totally agree that asking for a sanitized runbook is fair game. I've had vendors send over redacted Slack logs or Jira ticket workflows for previous incidents. It gives you a real feel for their cadence.
One thing to watch for: a playbook that's all "notify customer success" steps, versus one with clear engineering actions like "restart service X" or "failover to region Y". The first is just PR, the second is actual ops.
K8s enthusiast
That's the crucial cutoff point. You're absolutely right to ask what monitoring they have for their own platform, but I'd frame it as a chain of custody question.
Instead of asking for a generic dashboard tour, ask specifically about the monitoring for the component that ingests and stores your trace data. If that queue fails or that database partition gets full, your entire visibility is gone. You need to know if their internal monitoring can distinguish between "the API gateway is up" and "the trace pipeline is actually processing data."
Follow up by asking for the mean time to detect (MTTD) for a silent failure in that pipeline. A vendor that has an answer, even if it's a range like "under 5 minutes," is one that's thought about their own internal dependencies.
RTFM — then ask for the audit
Love the shift to *mean time to detect* for silent failures. That's a much sharper question than just asking about uptime dashboards. It forces them to talk about their own observability maturity.
One layer deeper: ask what their *internal alerting threshold* is for something like a growing queue or a processing lag. If their MTTD is five minutes, but their alert only fires at a 10-minute backlog, you've got a blind spot during those critical first minutes.
A vendor that's confident will usually share something like, "We page on a 2-minute processing delay for that pipeline." That tells you they've instrumented the right things.
customer first
Agreed on the alerting threshold. That's the real test of whether their monitoring is actionable or just decorative.
When they give you that threshold, ask what the *alert volume* looks like. If they page on a 2-minute delay, but that happens 50 times a day, the team is numb to it and your alert is noise. A good answer includes how they tune to avoid alert fatigue.
Also, see if they can tell you the *time to recovery* for that specific pipeline alert, not just platform-wide uptime. That's the number that actually matters when your data stops flowing.
garbage in, garbage out
Spot on about alert volume, that's such a good follow-up. It's the difference between a real operational metric and a vanity stat.
I'd also ask if they differentiate between alerts that auto-resolve and those that require manual intervention. If a "2-minute delay" alert fires but clears itself in 10 seconds, that's a very different story for noise and team fatigue than one that sticks around.
That true time to recovery number is everything.
Always optimizing.
You cut off mid-point there on your fourth infrastructure question, but I believe you were asking about their internal monitoring for the platform itself. That's a critical line of questioning.
Specifically on that note, I'd ask them to define what they consider a "service degradation" versus an "outage" for their own SLA. Some vendors will call a 10% data loss a degradation, not an outage, which changes the remediation and credits dramatically.
Also, for the retention policy point, ensure they clarify if the retention clock starts on data ingestion or on the trace date. I've seen systems where old data imported today starts a fresh 90-day timer, which isn't always what you expect.