Skip to content
Notifications
Clear all

Any good open source alternatives for tracing yet?

67 Posts
65 Users
0 Reactions
138 Views
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Totally agree on the volume point. I've seen spans from a single chain call balloon to over 100 for a complex agent, and that's a huge cost driver for a backend like Jaeger or Tempo.

You mentioned sampling for problematic endpoints, which is a lifesaver. We set up a rule to sample at 1% for general health, but automatically bump it to 100% if the chain's error rate spikes or latency goes above a threshold. It kept our data volume sane while still giving us full fidelity when things started to break.

And that "rough estimate" warning for cost is gospel. We built a whole reconciliation dashboard that compares our tracer's token estimates against the actual billing data from the vendor API. The variance can be 10-15% some months, which is scary when you're scaling up.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Yeah, that's the part I'm stuck on right now. Building that billing layer feels like a whole separate project.

> a simple cost aggregation layer on top of the open-source trace data

Has anyone actually done this? Even something basic that just reads the spans from a table and multiplies tokens by a static price. I'd love to see an example config.



   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

That pseudo-trace visualization is exactly what you'll get from a basic Phoenix or OpenLLMetry setup with LangChain. For a straightforward sequential chain, it's plug-and-play. The challenge, as you've seen in the thread, is that the visualization can collapse useful detail for complex agents.

On cost tracking, neither will give you accurate totals. The tracer provides the raw components - model names and token counts per span - but the mapping to your invoice requires that separate enrichment layer. I built a simple version that reads from a tracing backend's database.

```sql
-- Example query against a span table
SELECT
span_name,
span_attributes['gen_ai.system'] as model,
SUM(cast(span_attributes['llm.token_counts.completion_tokens'] as int)) as completion_tokens
FROM traces
WHERE span_kind = 'LLM'
GROUP BY 1, 2;
```

You'd then join this against a static pricing table. The variance against actual billing can be significant, so treat it as an estimate. Start with the visualization for debugging, then add this aggregation once you need to forecast spend.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

That weekly pull is such a smart move! Storing it as versioned config is the key part most people miss.

We tried a similar approach but ran into issues with provider pricing pages that aren't machine-readable or have special introductory rates. We ended up having to maintain a small scraper for each one, which became a chore. Now we use a hybrid model: automated weekly pulls, but with a manual validation step before the new rates go live. It's a bit more overhead, but it saved us from a nasty surprise when one provider changed their page layout.

Your point about the effective date is spot on. Did you build any tooling to automatically apply the correct historical rate to a trace, or is that a manual lookup?


Clean code, happy life


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 2 months ago
Posts: 240
 

That smart sampling rule you mentioned is gold. We've been doing something similar by also triggering 100% sampling on any execution that includes a specific error-prone tool. It's saved us so much debugging time.

And that 10-15% variance on cost estimates? We see the same. It's usually because the tracer's tokenizer doesn't always match the provider's. Our reconciliation dashboard was a wake-up call. We now use it to flag any span where the variance is over 5% for a manual check.



   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

> flag any span where the variance is over 5% for a manual check

That threshold is a good starting point, but we found it needed tuning by model family. The tokenizer mismatch is worse for some models than others, so we set tighter variance bounds on the models we spend the most on. It cut down the noise and let us focus manual checks where they actually mattered for the bill.

The bigger learning was that variance isn't always random. If you plot it over time, you can sometimes spot a new deployment where the tracer library is out of sync with the provider's API changes. It became an early warning system for silent instrumentation drift.



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That's a great and very practical question. The custom span processor typically lives in your application's initialization code, right alongside where you configure the tracer itself. It's a piece of middleware that runs in-process, inspecting and tagging spans before they get sent out.

For the Azure model variant issue specifically, you'd write a processor that checks the span attributes for the Azure endpoint and uses the request parameters to add a more specific `model_variant` attribute. The tricky part is ensuring those request parameters are actually captured in the span to begin with, which sometimes requires additional instrumentation hooks in your LLM client library.

Starting that billing layer is daunting, but you can begin incrementally. Just get the enriched span data into a database first. Even a simple daily query that groups by your new `model_variant` attribute and sums token counts will give you a starting point to build on.


—daniel


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Exactly. That's why the span processor pattern is critical. But you're fighting upstream instrumentation gaps.

If your LLM client library doesn't expose the Azure deployment name as a span attribute, your processor has nothing to read. You'll need to patch the instrumentation or switch libraries. I've had to fork the OpenTelemetry semantic conventions for this before.

On the billing layer, starting with a daily query is the right move. But don't just sum tokens. Correlate it with your cloud provider bill's line items from day one. The discrepancy will show you where your span attributes are wrong.


Metrics don't lie.


   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

It's exactly a separate project, and a tedious one. That basic SQL query you mocked up is the start of the pain, not the end of it.

Your static price multiplier breaks the moment you use more than one model, region, or provider. Then you're joining against a pricing table that needs its own pipeline to stay current. Don't forget pro-rated reserved capacity discounts if you have any.

Everyone wants the example config, but the config is the easy part. The operational burden of keeping the price data accurate and mapping model strings from spans to SKUs on an invoice is where the real cost hides. You're just building a shadow billing system.


-- cost first


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

That last point about it becoming a shadow billing system is spot on. It's a common trap teams fall into when they start chasing perfect cost visibility.

We ended up with two people spending half their week just maintaining the SKU mapping file and arguing about which rate card was the source of truth. The time spent reconciling often outweighed the savings we were trying to capture.

Maybe the real question isn't *can* we build it, but *should* we, when that effort could go into feature work a customer would actually pay for.



   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

The pseudo-trace visualization you're describing is what Phoenix's UI or an OpenTelemetry collector with a compatible backend (like Jaeger) will show you for a standard LangChain sequential chain. Hooking it in is straightforward with their callbacks.

On your second point about cost tracking, that's where the open source tools stop. They give you spans with token counts, but the mapping from a model name and token count to a dollar cost is a separate system. You'll need to maintain a price table per provider, which becomes its own data pipeline problem as rates change.

For debugging, the visualizations are useful for spotting latency outliers or which tool in an agent is failing, but they can get visually noisy for complex workflows. The real value for us was exporting the trace data to our data warehouse, where we could build custom dashboards for token usage and latency percentiles across different model deployments.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Yeah, the flattened "invoke" span for agents is a killer. We ended up writing a custom span processor to break those out based on the internal execution IDs in LangSmith's trace events, but it's brittle. Every framework update risks breaking it.

Your point about model name mapping is exactly why we gave up on a unified cost layer. The semantic conventions just aren't detailed enough. We now have a separate small service that scrapes the actual deployment names from our Azure ML workspaces and joins them to spans in our data warehouse. It's another ETL job, but at least it's using the real resource IDs, not guessing from strings.


Automate everything. Twice.


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

You're asking the right questions, but you're setting yourself up for the classic open source observability bait-and-switch. The tracing part? OpenLLMetry and Phoenix will hook into LangChain just fine with callbacks and give you a nice flame graph of your chain. That's the easy win they advertise.

The cost tracking across models? That's where the free lunch ends.

> mapping from a model name and token count to a dollar cost is a separate system

This is the understatement of the year. It's not just a separate system, it's a full-time job. You'll be maintaining a price table that needs to ingest rate cards from three different cloud providers and a half-dozen SaaS AI vendors, all of which change prices with zero notice. Your "simple example" of a pseudo-trace with token counts is the *input* to a monstrous reconciliation engine you now have to build and operate.

You'll spend more engineering hours maintaining your "open source" cost dashboard than you'd ever pay for LangSmith. The visualization is useful for debugging a stuck tool, but it tells you nothing about why your bill jumped 40% last month. That comes from the shadow billing system nobody has time to build correctly.


pay for what you use, not what you reserve


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

You've nailed the brutal reality. That "monstrous reconciliation engine" is exactly where teams get stuck. We tried building it internally and the project quietly died after three months because the mapping logic was impossible to keep current.

My caveat is that the visualization *can* help with cost spikes, but only in a reactive, forensic way. Last month's 40% bill jump? We used the flame graphs to trace it to a new feature that was unknowingly calling a premium model on every retry. The spans showed the pattern, but figuring out the actual dollar impact still took a manual join with that week's Azure rate card. So it helped diagnose, but it didn't prevent it.

Maybe the takeaway is to use OSS tracing for debugging latency and errors, but accept that for real cost governance, you need a vendor whose whole job is tracking those pricing APIs.


Test, measure, repeat


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

You're right to look at Phoenix and OpenLLMetry for the trace visualization piece. They both hook into LangChain callbacks pretty cleanly and will give you that pseudo-trace breakdown you sketched out.

But on your second point about cost tracking across models, that's the rabbit hole everyone in this thread is warning about. Those tools give you token counts in spans, but the dollar conversion is a manual lookup against a rate card you now own. If you're using Azure OpenAI, Anthropic, and maybe some open-weight models, you're maintaining three different pricing tables that change without warning.

The visualization is genuinely useful for spotting which tool in an agent is failing or why latency spiked. Just don't expect it to show you a dollar amount. For that, you're building a separate billing pipeline.



   
ReplyQuote
Page 4 / 5