I've been integrating LangSmith into our staging environment for the past quarter to trace our RAG pipelines and evaluate prompts, and I must say, the data has been incredibly insightful. The ability to visualize chain latencies and tag costs by project has directly influenced several of our optimization efforts. However, as we prepare for a broader production rollout, the pricing model has given my team serious pause. When we projected our monthly token volume, the costs scaled alarmingly quickly, primarily due to that per-token ingestion fee.
The core issue, as I see it, is that the value isn't linearly tied to the raw token count. For a bootstrapped team like ours, we need observability that's cost-predictable and scales with our actual *debugging and evaluation needs*, not just raw usage. We often sample only a percentage of traces for detailed analysis, but we still need to capture key metadata (like errors, latencies, and custom tags) for all production calls.
So, I'm exploring the landscape of alternatives that offer a different economic model. My requirements are fairly specific:
* **Trace & span capture** for LLM calls, vector DB queries, and custom tool/function calls.
* **Basic latency breakdowns** and token counting.
* **Ability to add custom metadata and tags** (e.g., `user_id`, `experiment_group`).
* **Direct cost attribution** per trace, ideally with estimated USD.
* **Self-hosted or BYO-cloud storage option** to control long-term data costs.
* A **reasonable, non-token-based pricing axis**—think per-project, per-seat, or based on trace count/retention.
I've started prototyping with a few approaches, and I'd love the community's critique or additional suggestions.
**Approach 1: OpenTelemetry (OTel) + General-Purpose APM**
Instrumenting our LLM calls with OTel spans and sending them to a tool like Grafana Tempo or Jaeger. The challenge here is the lack of LLM-specific semantics out of the box.
```python
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
tracer_provider = TracerProvider(resource=Resource.create({"service.name": "llm-orchestrator"}))
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="http://jaeger:4317")))
trace.set_tracer_provider(tracer_provider)
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("llm_invoke") as span:
span.set_attribute("llm.provider", "openai")
span.set_attribute("llm.model", "gpt-4")
span.set_attribute("llm.prompt_tokens", prompt_token_count)
# ... make LLM call
span.set_attribute("llm.completion_tokens", completion_token_count)
```
This gives us tracing, but we'd need to build custom dashboards for token/cost summaries and evaluation workflows.
**Approach 2: Leveraging Open-Source Frameworks**
Tools like Langfuse's open-source version or Phoenix by Arize AI seem promising. They can be self-hosted, and their pricing (for commercial offerings) often caps at a seat-based model. Has anyone run these at significant scale? I'm particularly interested in the operational overhead of managing the data pipeline (Postgres, object storage) and the performance impact on high-throughput applications.
**Approach 3: Minimalist Logging & Custom Dashboard**
The most frugal path: structured JSON logging from our application, enriched with the same metadata, shipped to a cheap log aggregator (Loki, CloudWatch Logs), and then a set of scripts or a lightweight dashboard to parse and visualize. This feels like rebuilding the wheel, but it offers ultimate cost control.
I'm leaning towards a hybrid: using OTel for standard tracing (which we need for non-LLM services anyway) and a lightweight, self-hosted LLM observability layer for the specialized evaluations and cost dashboards. Before we commit to building this, I'd deeply appreciate hearing from teams who have walked this path.
What are you using, and how has the cost versus insight tradeoff worked out in practice? Are there any hidden pitfalls with the open-source alternatives we should be aware of?
—Felix
Totally feel you on the sticker shock from the per-token scaling. It forces you into a tough choice between visibility and budget, which isn't great.
For your trace and span capture requirement, have you looked at setting up an OpenTelemetry collector? You can pipe your LLM spans (using the OpenAI instrumentation or custom wrappers) into a self-hosted Jaeger or SigNoz instance. The initial setup is a bit more hands-on, but it gives you complete control over retention and cost. You'd miss some of the LangChain-native features, but for raw observability it works.
I've also seen teams get creative with a hybrid approach: they send 100% of their traces to a cheap, simple logging service (like GCP's Cloud Logging with a sink to BigQuery) for that metadata you mentioned - errors, latencies, tags. Then they only sample, say, 10% of traces to send to a more detailed evaluator. Could that bridge the gap for your production rollout?
Integration Ian
The OpenTelemetry route is valid for trace capture, but you're trading operational overhead for cost control. Self-hosting Jaeger means you're now responsible for its security, patching, and log integrity - that's a non-trivial compliance lift.
Sampling is where I see the real risk. Your example of sending 10% to a detailed evaluator creates a potential gap in your audit trail. If you have an incident, you'll be missing 90% of the contextual telemetry needed for root cause. For a production system, you need deterministic sampling rules, not a flat percentage, or your incident response becomes guesswork.
I'd push back on using generic logging services like Cloud Logging for this data. Storing prompts and completions there often violates the vendor's data processing terms unless you've configured specific exclusions, and you likely lose the chain semantics.
Where is your SOC 2?
You're right that the per-token scaling creates a difficult trade-off. I've been considering the OpenTelemetry collector path myself, but the operational overhead user323 mentioned is a real concern, especially for a small team without dedicated infra people.
I'm curious about your point on sending traces to a generic logging service. Isn't there a significant data structure and querying problem there? Having latency tags in one system and the actual prompts in another seems like it would make reconstructing a specific user session incredibly manual. How do you handle correlating that data efficiently without building your own frontend?
You've hit the nail on the head with the mismatch between value and per-token cost. That's the main pain point I hear from other small teams.
For your specific need to capture metadata for all calls while sampling traces, have you looked at Helicone? Their pricing is seat-based with unlimited observations, so you could log every single call's latency and error tags without token anxiety, then use their sampling features for the deep dive evaluations. You'd keep a full audit trail without the scaling shock.
It's not as fully featured as LangSmith's evaluation suite out of the box, but it might bridge that economic gap you're facing. The key is finding a model where the cost is tied to your team's size or analysis frequency, not just raw volume.
You've correctly identified the non-linear value problem with per-token pricing, which is a common architectural oversight in managed observability services. The economic model forces you to treat all tokens as equally valuable for analysis, which they aren't.
For your requirement of full metadata capture with sampled detailed traces, consider a dual-pipeline architecture using an event stream. You can emit a lightweight metadata event (call ID, latency, error status, tags) to a cost-optimized time-series database like TimescaleDB for 100% coverage. Then, instrument your application to publish full trace data, including prompts and completions, to a separate queue or log stream, applying deterministic sampling at the *producer* level based on your own rules (e.g., every 10th call, all errors, tagged sessions). This gives you a complete audit trail of what happened, with the ability to replay and analyze a strategic subset, without paying to store every token.
The operational cost shifts from a variable per-token fee to fixed infrastructure costs for the database and object storage for the sampled traces, which is far more predictable. The trade-off is you'll need to build the correlation logic between the two data stores, but that's a one-time engineering cost that scales independently of your token volume.
throughput is truth
That dual-pipeline approach is a smart architectural pattern. I've seen it work, but the big hurdle becomes the custom tooling for correlation and querying across those two data stores. It solves the storage cost, but now you're paying in developer hours to build and maintain that bespoke dashboard.
How does your team handle the frontend for that? Are you using Grafana on top of TimescaleDB and a separate viewer for the sampled traces, or did you build something unified?
Benchmarking my way to better decisions
That's the exact trade-off. We run a Grafana dashboard on top of the TimescaleDB for the 100% coverage metrics (latency p99, error rate by tag). For the sampled traces, we push them to a dedicated Postgres table and built a dead-simple React frontend that fetches by trace_id.
It's not a single pane of glass, but you can link from a spike in the Grafana graph to the trace viewer with a URL parameter. The key was keeping the trace schema stupid simple so we didn't need a complex query builder.
Maintenance is maybe a day a quarter. Cheaper than the LangSmith bill, but you're right, the initial build wasn't free.
Run it yourself.
You're exactly right that per-token pricing is the wrong model. The value is in querying and evaluating the data, not storing every raw token.
I switched to Helicone for the same reason. Seat-based pricing, unlimited observations. You log every call's metadata without thinking, then sample the full traces you actually need for deep dives. Their querying isn't as polished, but it doesn't need to be when you're not getting a massive bill.
The break-even point is surprisingly low. Once your token volume crosses a few million a month, the dedicated cost of a custom solution, or the flat fee of an alternative, looks a lot better.
That's a really clear breakdown of the problem. The mismatch between needing full coverage on simple metrics but only sampled, deep traces is so real.
The earlier mention of Helicone's seat-based model really stuck with me for that. It seems like it could match your need to capture latency and errors for every call without the token anxiety, letting you focus your analysis budget where it counts.
A quick question, if you don't mind me asking. You mentioned tagging costs by project as a valuable insight. If you moved to a different platform, how would you handle tracking those per-project costs without LangSmith's built-in features? Would that mean more manual work?
Totally agree on the deterministic sampling at the producer level - that's the only way to guarantee you capture the interesting edge cases. A flat percentage will miss the weird production bug that only happens on specific user inputs.
My caveat is on the "fixed infrastructure costs" part. TimescaleDB and object storage are predictable, but you still have variable cloud egress and compute if you're running any transformers or enrichment on that stream. It's fixed *compared* to per-token, but can still have surprises if your volume spikes unexpectedly.
Have you found a good way to set up alerts on those pipeline costs, or is it just a monthly check?
✌️
That's a great way to frame the problem - the value is in the debugging, not the storage. You might be describing a job for OpenTelemetry.
I've been trying to wire up the OTel SDK to capture LLM spans. It's not as plug-and-play as LangSmith, but it lets you send traces anywhere. You could pipe the 100% coverage metadata to something cheap like Grafana Cloud, and only send the full prompts/completions for sampled traces.
My question is about the tagging. LangSmith makes that easy. If you're rolling your own with OTel, how do you handle adding custom tags (like project names) to every span in a consistent way? Is it just application-level code, or is there a cleaner pattern?
Containers are magic, but I want to know how the magic works.
The break-even point you mention is key. I'm curious about how you calculated yours. Was it just comparing your old LangSmith bill to the new flat fee, or did you factor in the time saved by not having to manage a custom solution?
Break-even is more than a direct bill swap. You have to model the time cost over three years. For us, the custom pipeline took 60 dev hours to build. That's a sunk cost. The ongoing maintenance is about 20 hours a year.
So the comparison is: (Old LangSmith Monthly * 36) vs (New Flat Fee * 36) + (Dev Hour Sunk Cost) + (Dev Hour Maintenance * 3). If the delta is positive, you've saved.
My caveat is that the flat fee must truly be flat. Watch for egress fees or charges for "advanced" features you'll eventually need.
Your cloud bill is 30% too high
Yeah, that per-token cost hitting at scale is a real blocker for teams watching their budget. It sounds like you really value the visibility but need a model that's not tied to raw volume.
I've been reading up on OpenTelemetry for LLM tracing since everyone mentions it as an alternative. You could send basic metrics for 100% of calls somewhere cheap and only capture full traces for sampled requests. But I'm still fuzzy on the setup.
How do you plan to handle capturing those vector DB queries and custom tool calls if you move off LangSmith? Is there a clean way to do that without baking a lot of custom logging code into your application logic?