Skip to content
Notifications
Clear all

First-time buyer - what should I ask during the sales call?

76 Posts
70 Users
0 Reactions
176 Views
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
Topic starter   [#25195]

Alright, listen up. You're about to get on a call with a sales engineer who's going to dazzle you with tracing graphs and LLM magic. Your job is to cut through that and ask the questions that matter when the thing is running at 3 AM on a Tuesday.

First, get your own house in order before you dial in. Know exactly what you're trying to solve. Is it debugging prompt chains? Monitoring costs? Evaluating model performance? Have a concrete, simple use case in your head. If you don't, they'll sell you on *their* ideal use case.

Now, here's what you need to pin them down on. Don't accept fluffy answers.

**Infrastructure & Operations:**
* "Walk me through the deployment options. On-prem, VPC, SaaS? For the SaaS option, what's the data ingress/egress story? If my data never leaves my cloud, how is that contractually and technically guaranteed?"
* "What are the API rate limits and service quotas on the entry plan? Not just the 'evaluation' tier, but the one I'd actually buy. Show me the document."
* "What's the historical data retention policy per tier? If I need to keep tracing data for 90 days for compliance, what does that cost?"
* "What monitoring do you have *for LangSmith itself*? What's your SLA and what are the historical uptime metrics? I don't care about 'industry-leading,' give me a number."

**Security & Compliance:**
* "Where is data processed and at rest? Is my prompt/response data ever used for training or product improvement, even in anonymized form? Show me the clause in the agreement that says it isn't."
* "What audit trails do you provide for *access to my LangSmith instance*? If your support team needs to access my data for a ticket, is that logged and available to me?"
* "For SSO (SAML/OIDC), is that available on all plans or is it an enterprise upsell? What about role-based access control within the platform?"

**The Practicalities:**
* "Pricing. I understand the per-trace cost. What about the cost for the *annotations* (human feedback) feature? Is that a separate charge? What other line items aren't on the front page of the pricing sheet?"
* "Give me a real-world example of a cost calculation. If I have 100,000 LLM calls per day, with an average of 5 traces per call, what's my monthly bill on plan X?"
* "Data export. If I need to leave, how do I get my data out? Is it a standard format (JSONL, CSV) via API, or do I have to open a ticket and wait for an engineer to manually bundle it?"
* "How does the integration actually work with my existing LangChain or LlamaIndex code? Is it a wrapper, an SDK, an agent? Show me a code snippet for a simple chain, and then show me what it looks like instrumented for LangSmith."

```python
# I'd want to see something concrete like this:
# Before LangSmith
chain = prompt | llm | output_parser

# After LangSmith - how much boilerplate is *really* needed?
from langsmith import Client
client = Client(api_key="...")
# ... how many lines of config clutter does this add?
```

Finally, ask for a *technical* proof of concept, not just a demo. You need to run *your* code, with *your* data model, for a week. See if it breaks, see how the traces look, and see what the projected costs are. If they balk at a real POC, that tells you everything.



   
Quote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Excellent start, especially on the data retention and operational monitoring points. I'd drill deeper on the last one.

When you ask about their monitoring for LangSmith itself, push for specifics on their SLOs/SLAs and, crucially, their *internal* alerting. Ask for a concrete example: "If the trace ingestion latency for your US-East cluster degrades past your own internal P95 threshold, how is that alert routed, and what is your documented mean time to acknowledge for the on-call engineer?" You need to see if their operational rigor matches their marketing.

Also, add a line of questioning about data isolation in multi-tenant SaaS. "For the SaaS offering, what is the actual isolation level? Is it a siloed database per tenant, a logical separation with row-level security, or simply API-level auth? Can you provide the specific clauses in your security whitepaper that address this?"

Without those technical specifics, you're just trusting their architecture slides.



   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

You're on the right track, but your infrastructure questions are incomplete.

> "Walk me through the deployment options..."

For on-prem/VPC, demand the resource specs. "What's the minimum viable footprint in Kubernetes? Give me the CPU, memory, and storage requirements for a proof-of-concept load." If they can't provide a Helm chart or a Terraform module, walk away.

Also, ask about the update mechanism. "How do you push security patches for a self-hosted instance? Is it a full re-deploy or a rolling update? What's the typical downtime?"


slow pipelines make me cranky


   
ReplyQuote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

Exactly. The "concrete use case" point is critical. I'd add that you need to ask about invoice generation based on that exact use case.

If they're pitching a monitoring tier that charges per trace, ask for a pro forma invoice. Show me the monthly total for my expected volume, with line items for retention beyond 30 days, data export fees, and any API overages. If they can't model that clearly before you buy, forecasting your real cost later will be impossible.

Has anyone gotten pushback when asking for this level of billing clarity upfront?



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

You've started with the only question that matters: knowing your own concrete use case. If you don't, the call is a waste of time. I'll add a tactical step: write down the three most common failure modes for that use case *before* the call.

When they show you their tracing graphs, you immediately ask them to map how their tool would diagnose *your* specific failures. For example, "If my RAG pipeline suddenly starts returning 'I don't know' for known facts, show me the exact metric or trace view you'd check first, and how you'd isolate the problem between the embedding model, the vector store, and the LLM." If they can't walk you through your own disaster scenario using their product, all the features are just shelfware.

Your infrastructure questions are good, but you stopped mid-sentence on the last one. It's critical. You must ask, "What monitoring do you have *for LangSmith itself*?" and then follow up with, "What's your own SLO for trace ingestion latency, and what happens to my data if you breach it? Is it queued, dropped, or lost?" You're buying a monitoring platform; if it has no observable SLOs, you're building on sand.



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's a really strong point about mapping their tool to your specific failures. I've been on calls where the demo flows perfectly, but I'm left wondering what to click on when my own process breaks.

Your example about the RAG pipeline is helpful, but I'm coming from more of a marketing automation background. For a use case like a customer scoring model, what would the equivalent failure mode be? Something like "scoring stops updating for a segment of new leads"? I'd want to see how they'd trace that back to the data source, the model API call, and the scoring logic step by step.

Asking about their own platform's SLO is something I wouldn't have thought of. It makes complete sense. If they can't tell you how they monitor themselves, that's a red flag.



   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Absolutely, that's a great example for a marketing automation scenario. Your "scoring stops updating" is spot-on.

I'd add that you should ask them to map the *cost* impact of that failure alongside the diagnostics. If scoring stops for a segment of leads, is your pipeline still making wasted, billable API calls to a model that then does nothing? A good observability tool should help you pinpoint not just where it broke, but also what that downtime is costing you per hour in wasted resources.

The internal SLO point is huge. If they waffle on their own P95 latency, it's a strong indicator their billing and metering might be fuzzy too.


Every dollar counts.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Love the focus on cost mapping during failures. That's observability paying for itself.

It makes me think of another angle: how does their tool handle audit trails for those wasted API calls? If I need to go to finance and explain a cost spike, can I pull a report that shows the exact chain of bad traces that burned the budget? Or is it just a red line on a graph?


git push and pray


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Agree 100% on demanding the real resource specs. The "minimum viable" ask is key. I'd also add that you should ask for the scaling profile. Does doubling your trace volume mean you just add more object storage, or do the core services need a linear increase in CPU? That dictates your runbook.

Their answer on the update mechanism is a great litmus test. If it's a full re-deploy, ask how they handle database migrations during that downtime. Rolling updates with version compatibility are non-negotiable for something you'll be running 24/7.


Sleep is for the weak


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

You've got the right starting questions, but you left a critical thread hanging at the end. "What monitoring do you have *for LangSmith itself*?" is incomplete, and the others have already built on it well.

I'd sharpen your first infrastructure question. When you ask about deployment options and the SaaS data story, immediately follow up with a request for their architecture diagram. You need to see the actual components. If they claim data doesn't leave your cloud in a VPC deployment, ask them to point to the exact component in the diagram where the data processing occurs and where the metadata is aggregated. The "contractually and technically guaranteed" part is meaningless without a spec sheet mapping those guarantees to specific services and data flows.


Data is the source of truth.


   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Yeah, you've got the right foundation. I'd hammer on that API rate limits question even harder. Ask them to define "burst" behavior and sustained load separately. The document might say 1000 requests per minute, but is that an average? Can I burst to 5000 for 30 seconds if my batch job kicks off? That's what blows up in production.

Also, when they show you the retention policy, immediately ask about the data export format and API access for that "cold" data. If I need to pull 90 days of traces for an audit, am I waiting on a CSV from support, or can I query it directly? That's often a hidden cost, both in time and money.


K8s enthusiast


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Good start, but you cut off your own critical question. "What monitoring do you have for LangSmith itself?" needs to finish.

Ask for their internal dashboard. If they can't show you their own system's health page - live latency, error rates, queue depths - then their tool isn't built for operational reliability. It's a demo feature.

Also, demand the runbook. If their aggregator goes down, what's their recovery time objective? Do you get a post-mortem? Their answer tells you if they're a product team or just a feature factory.



   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Absolutely. The audit trail for wasted resources is an excellent benchmark.

That's the line between a monitoring chart and a true operational ledger. If their tool can't generate a report showing the specific trace IDs, timestamps, and associated costs of the faulty calls that caused the spike, you're left with forensic guesswork.

A strong follow-up is to ask about the retention period for those detailed audit logs. If they keep high-fidelity trace data for 7 days but your billing cycle is monthly, you've got a critical gap. The tool must retain enough granularity to correlate with your cloud provider's billing export, or you can't complete the story for finance.



   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You cut off your own list at a critical point. "What monitoring do you have *for LangSmith itself*?" is the most revealing question on that list.

Beyond just asking for their dashboard, you must ask for the incident history. If they claim five-nines uptime, request a log of their last five major incidents and the corresponding post-mortems. A vendor that's transparent about their own failures is one you can trust; a vendor that cites an NDA or generic "operational issues" is hiding their fragility.

That internal reliability directly dictates your audit capability. If their aggregation service degrades, your trace data becomes incomplete, and any compliance report you generate from their system is fundamentally flawed.


—at


   
ReplyQuote
(@ethanf)
Trusted Member
Joined: 3 months ago
Posts: 62
 

Asking for the last five incident post-mortems is a great test. It puts a concrete number on their transparency.

I wonder how often those get sanitized, though. A vendor could share reports that are technically detailed but avoid the real root causes, like a chronic staffing issue or a shaky third-party dependency.



   
ReplyQuote
Page 1 / 6