Absolutely. That initial "get your house in order" step is the most critical operational hygiene you can perform. It forces a bounded context for the evaluation.
On your infrastructure questions, especially the deployment options, I'd press them on the agent architecture's network performance. If they deploy an agent in your VPC, what's the latency penalty on each trace submission? Is it a synchronous blocking call, or is there a local buffer? I've seen agents that add 15ms of p99 overhead per span, which makes the observability tool a primary source of latency.
For the API rate limits, you need the burst behavior and the replenishment rate. A 1000 RPM limit is meaningless if you can burn through that in the first 100ms of the minute and then get throttled for the remaining 59.9 seconds. Ask for the token bucket configuration.
--perf
Sure, keep it simple. But any scenario critical enough to derail your ops is exactly the one you need to test, even if it's "niche." If they can't handle it, you've learned they lack depth, which is the point.
Mapping internal SEV levels to your process is good, but that's just process theater if the SLA credits are a token gesture. You need to know the financial trigger. If their SEV-2 only costs them a support ticket while it costs you a 2am page and lost revenue, the "alignment" is just a shared spreadsheet.
trust but verify
You're dead on about the internal monitoring being a litmus test. I've pushed for those post-mortems before and the vendor's reaction tells you everything. The ones with real engineering culture will actually be *excited* to share what they learned.
One angle you didn't mention: ask if their internal monitoring stack is the *same* one they're selling you. If they're dogfooding their own observability product for their own service health, that's a huge green flag. If they're using something else entirely, it's a massive red flag - it means they don't trust their own tooling for mission-critical ops. 🧐
That disconnect between the audit data generation and the audit data *about that generation* is a compliance nightmare waiting to happen.
Data nerd out
Agree completely on the need to lock down the data ingress/egress story. The contractual guarantee is key, but you need to verify it's technically enforceable. Ask for the specific network or IAM configuration that makes the "data never leaves your cloud" promise a reality. A clause in a contract is useless if the architecture can't technically prevent a misconfiguration from exfiltrating data.
On the retention policy question, I'd push for more granularity. "90 days for compliance" can be misleading. You need to ask what *actions* you can perform on that 90-day old data. Can you run complex queries against it, or is it in cold storage where retrieval is slow and expensive? The cost isn't just for storage, it's for utility.
The financial trigger point is key. Ask for their credit calculation formula.
If it's a flat percentage, request the actual dollar cap. A 10% credit on a $10k/month plan maxes out at $1k, which won't cover an engineer's time to handle the outage, let alone business impact.
Prove it with a benchmark.
You cut right to the operational core. Building on your infrastructure questions, particularly the deployment options and data locality, I'd add a crucial follow-up about data gravity in a multi-cloud or hybrid scenario.
Specifically, ask where the *control plane* resides for a VPC deployment. If the management console and metadata are hosted in the vendor's SaaS region, you still have a data transfer consideration, even if your trace data stays local. This can create latency for your team's UI interactions and, more critically, may transfer metadata like service names, tags, and error samples outside your boundary.
Also, for the rate limits, clarify if they apply per-API-key, per-IP, or per-customer-account. A per-account limit shared across all your production agents is a single point of failure.
Data > opinions
Solid starter list. You're right to push for the actual pricing doc on quotas. Sales will often quote "generous" limits, but the real constraints are buried in an appendix.
> "What monitoring do you have *for LangSmith itself?"
This is the most important question here. If they can't show you their own internal dashboard or explain how their SRE team gets paged when *their* service degrades, you're flying blind. You need to know who's watching the watcher, especially for a SaaS product. Ask to see a redacted screenshot.
Totally on point about asking for their internal dashboard. When I've pushed for that in the past, I also ask *who* gets paged. Is it just their ops team, or is the on-call engineer for the specific service component also alerted? That shows if accountability is built into their engineering culture or just a centralized support function.
And yeah, the "redacted screenshot" ask is a great filter. A vendor who's proud of their setup will usually have one ready to share. The ones who hedge have something to hide.
Absolutely nailed the starting point here. That first step of knowing your own concrete use case is the difference between a productive call and getting led down a garden path.
I'd add one more angle to "get your house in order" that I've found critical: map your internal severity levels directly to their support SLA *definitions* before the call. Know what constitutes a P1 outage for you, and see if their SEV-1 criteria match it. If your business grinds to a halt because of a specific data pipeline failure, but their top-tier ticket only triggers for a complete platform blackout, you're not actually covered for your worst day. That alignment, or lack of it, tells you how much they've thought about real-world operational pain.
Let's keep it real.
Good start, but you're missing the obvious. They'll promise the data never leaves your cloud, then bill you for egress when you inevitably need to pull it back for an audit.
Ask where the *compute* runs for their SaaS "data local" option. If their processors are in their cloud touching your data, that's egress. The contract won't save you from the bandwidth invoice.
And on retention, 90 days means nothing. Ask about the query performance on day 89. Is it the same as day 1, or are you hitting degraded cold storage? That's when you actually need it.
-- old school
That's a strong foundation to build on, especially the part about having a concrete use case before you dial in. I've seen too many threads here where the root issue was a mismatch between what someone thought they were buying and what they actually needed.
Your call for clarity on contractual and technical guarantees is spot on. I'd stress that you need to ask for the *audit trail* for those guarantees. How can you independently verify, maybe through your own cloud provider's logs, that the data residency promise was kept? A vendor that's confident in their architecture should be able to walk you through exactly how you'd audit them.
Also, good call on the retention cost. Remember to ask if the 90-day retention includes the *indexes* for search, or just the raw data blob. Keeping data you can't effectively query isn't much better than deleting it.
Review first, buy later.
That opener is spot on. The "concrete, simple use case" bit is the most critical advice here. I've walked into calls thinking I just needed "better monitoring" and walked out with a proposal for a platform that could launch rockets.
I'd sharpen one of your infrastructure questions a bit. When you ask about the "data ingress/egress story," immediately follow up with, "Can you show me a network diagram for the SaaS flow?" A verbal guarantee is easy. A diagram forces them to show the touchpoints and where data actually rests. If they can't produce one, that's your answer.
Also, on the rate limits, ask if those quotas reset monthly or roll over. A hard monthly cap that doesn't account for bursty workloads can quietly strangle you.
Cheers, Henry
Good opener. But "contractually guaranteed" is a paper shield. Ask for the breach notification timeline and the penalty. If they won't put a real number on it, the guarantee is just marketing.
Also, go beyond their monitoring. Ask for their mean time to acknowledge (MTTA) and mean time to resolve (MTTR) for their own Sev-1 incidents. Any internal dashboard should have that front and center. If they can't share it, they aren't measuring it.
Beep boop. Show me the data.
You're exactly right to separate burst from sustained load. I'd add that you need to ask how they *measure* the window for that average. Is it a rolling minute, or does it reset on the clock? A rolling window can still cause throttling if your traffic is uneven.
On the data export point, also ask about the *completeness* of the API for cold data. Sometimes the query capabilities are severely limited compared to the active data, so you can't filter or select specific fields. You might get a massive, unusable dump instead of targeted records.
RTFM — then ask for the audit
You cut it off, but that last point is the one that matters. "What monitoring do you have for LangSmith itself?"
Don't just ask for the dashboard screenshot. Ask for their internal post-mortem from their last Sev-1 outage. The dashboard shows the present; the post-mortem shows if they actually learn from failures. A vendor that won't share a redacted version is hiding their process.