Skip to content
Notifications
Clear all

Why is LangSmith so slow on large trace volumes?

4 Posts
4 Users
0 Reactions
37 Views
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
Topic starter   [#10573]

Anyone else hitting a wall with LangSmith performance when you scale past a few thousand traces? My team's ingestion pipeline is crawling. Latency spikes, UI becomes unusable.

* "Observability platform" that can't observe itself? Classic.
* Their hosted solution feels like it's running on a potato when you have real throughput.
* Tried their suggested "best practices" for batching. Marginal gains, at best.

Are we just using it wrong, or is the architecture fundamentally not built for production-scale LLM apps? The pricing is painful for this level of performance.



   
Quote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Been there. The "potato" description is painfully accurate once you cross a certain throughput threshold. Their batching advice is essentially a band-aid on a deeper indexing/ingestion bottleneck.

We ran our own benchmark and saw trace latency increase exponentially after about 5k concurrent traces, regardless of batch size. The UI isn't just slow - the underlying queries time out.

You're not using it wrong. The hosted service's scale ceiling is surprisingly low for something marketed at production. We ended up shifting to a different vendor for the high-volume pipelines and kept LangSmith only for lower-stakes dev testing. Their pricing model doesn't reflect that reality.


shift left or go home


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

That "running on a potato" line hits hard, because it reveals the core issue. Their pricing tiers are based on trace volume, not compute or IOPs for the underlying datastore. You're paying for the quantity of your observability data, but they're not provisioning resources proportionally to handle querying that volume.

It's a classic mismatch between a usage-based revenue model and a fixed-cost infrastructure model. They sell you the traces, but the cost to *index* and *serve* those traces grows non-linearly for them. The performance ceiling you're hitting is likely that inflection point.

We saw the same, and the batching advice only helps their ingestion costs, not your query latency. Have you checked if the slowdown is uniform, or is it specific to certain views like the trace list versus latency charts? That can point to whether it's an indexing or a query problem.


Every dollar counts.


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

"Observability platform that can't observe itself" is the perfect summary. It's the central irony they don't want to talk about.

You're not using it wrong. The pricing model is the tell. They charge per trace because that's easy to meter, but the cost of querying that data is what kills them. Their architecture is clearly optimized for the demo video, not for you actually using the data you paid them to store.

Marginal gains from batching is exactly the experience. It helps their ingest cost, not your query performance. At scale, you're just queuing up more data for a system that can't serve it back.


Your stack is too complicated.


   
ReplyQuote