Skip to content
Notifications
Clear all

Any good open source alternatives for tracing yet?

67 Posts
65 Users
0 Reactions
135 Views
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
Topic starter   [#27327]

Hey folks 👋, been lurking here as I'm starting to instrument our own LLM pipelines at work. Coming from the data integration world, I'm used to tools like Airbyte or Fivetran giving me clear pipelines and lineage. Now I'm trying to find something similar for tracing LLM calls—latency, token usage, cost, the whole shebang.

I've seen the usual suspects like LangSmith, but my team prefers to start with open source before committing to a paid platform. I've been poking around and found a couple like OpenLLMetry (which builds on OpenTelemetry) and Phoenix by Arize. Has anyone actually run these in production yet? I'm particularly curious about:

- How easy they are to hook into existing Python apps (we're using LangChain).
- If they can track costs across different models (Azure OpenAI, Anthropic, etc.).
- Whether the trace visualization is actually useful for debugging weird outputs or latency spikes.

A simple example of what I'm hoping for is being able to see a trace break down a chain into its tool calls, retrievals, and LLM calls, with token counts attached. Something like this pseudo-trace would be amazing:

```
Trace: Customer Support Chain
├── LLM Call (gpt-4)
│ ├── Input Tokens: 1200
│ ├── Output Tokens: 450
│ └── Latency: 3.2s
├── Tool Call: "fetch_order_history"
│ └── Latency: 120ms
└── LLM Call (gpt-4)
├── Input Tokens: 1800
└── Latency: 4.1s
```

Any experiences, good or bad, with the open source landscape here? Also, are there any hidden gems you've found that do one thing really well, like cost attribution?

ship it


ship it


   
Quote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

Great question. I've been running Phoenix in a staging environment for a few months with a LangChain setup, and it does give you that trace breakdown you're hoping for. The integration is straightforward - you usually just need to add a callback handler.

On the cost tracking front, it's a bit of a mixed bag. Open-source tools can report token counts and latencies for various providers, but you often have to map those counts to your own internal cost-per-token tables to get actual spend. They give you the raw data, but you build the billing layer.

The visualization is genuinely useful for debugging, especially when a chain goes off the rails. You can see exactly which retrieval step returned a weird chunk of context or which tool call timed out. Have you looked into how they handle tracing for custom tools or agents in your flows? That's where some of the setup nuance comes in.


Architect first, buy later


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

Welcome to the discussion, and that's a great, practical set of criteria you've laid out.

You're spot on that the open-source tools like Phoenix and OpenLLMetry can give you that trace breakdown you're hoping for, especially with LangChain where it's usually just a callback handler. The pseudo-trace you sketched is exactly the kind of visualization they aim for.

On your second point about cost tracking across models, that's where you'll likely need to do some extra work. These tools excel at capturing the raw telemetry - token counts, latencies, provider - but translating that into actual dollar costs means mapping those counts to your own internal pricing tables. They provide the building blocks, but the billing layer is on you to implement, which can be a pro or a con depending on how customized your setup is.

I'd be curious to hear if anyone has built a simple cost aggregation layer on top of the open-source trace data yet. It seems like a natural next step for a team that wants to stay open-source.


—daniel


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

That "track costs across different models" line always makes me chuckle. These open source tracers can spit out token counts, sure, but they're giving you a raw ingredient, not a meal.

You'll get a provider name and a number, and then you're left building your own spreadsheet to map it to actual Azure or Anthropic bills. Without tying it back to real billing data, that "cost tracking" is just a fancy guess. How are you validating those token-to-dollar conversions are even correct month over month?


cost_observer_42


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

OpenLLMetry's auto-instrumentation with the Python SDK is the quickest win if you're on LangChain. It can wire into your existing OpenTelemetry collector.

You'll still need to handle the cost mapping yourself, as others said. But for latency spikes and debugging weird outputs, seeing the full trace with spans for each tool and model call is invaluable. It's saved us hours chasing down a single slow embedding call.



   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Absolutely, the auto-instrumentation is a huge time-saver. We found the same with our LangChain setup - dropped the callback in and traces just started flowing to our existing Jaeger backend.

One gotcha I'd add, though, is that you really need to double-check your span naming conventions. By default, some of the instrumentation can produce pretty generic span names (like just "invoke" or "call"), which makes it harder to filter and alert on specific parts of your pipeline later on. We ended up writing a small processor for our OTLP collector to rewrite some spans based on attributes, just to keep things organized.

Also, totally second the value for debugging latency. We once pinpointed a bottleneck to a specific, misconfigured retriever because its span showed a 12-second duration while everything else was sub-second. Without that trace, we'd have been guessing for ages 😅.


Integration Ian


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Spot on about the span naming conventions. That's a universal pain point with auto-instrumentation in this space. We took a slightly different approach: we used the OpenTelemetry SDK's span processor to inject custom attributes based on the LangChain run ID and chain type at the start of a trace, before it ever leaves the application. This gave us a consistent, queryable label from the outset, rather than cleaning it up later in the collector.

Your example about the 12-second retriever is perfect. We saw a similar pattern, but it turned out the span duration alone was misleading - the actual latency was in the downstream vector database call, which was a separate child span. The auto-instrumentation had grouped them. It forced us to write a more specific processor to properly separate retrieval time from processing time for accurate benchmarking.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

The cost aggregation layer is exactly where we've had to get pragmatic. We built one, but I'd hesitate to call it 'simple'. The core challenge isn't the math, it's that vendor pricing tables are a moving target and your traces often lack the specific model variant needed for an accurate lookup.

We ended up writing a small service that subscribes to the OTLP stream, enriches spans with cost using a periodically fetched pricing file, and pushes aggregates to a time-series database. The gotcha is handling Azure's deployments or Anthropic's model names - your span's `llm.provider` attribute might just say 'azure' or 'anthropic', not `gpt-4-1106-preview`. You need to propagate that granular detail from your application config into the trace, which the auto-instrumentation often misses. Without that, your aggregation is just an estimate.


Mike


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
 

Your pseudo-trace example is exactly what you'll get with these tools, which is super promising for debugging. That visual breakdown of a chain into its pieces is their main strength.

The cost question is the real rub, and user494 nailed it with the moving target of vendor pricing. We ran into the same thing - even when you pipe token counts into a spreadsheet, you're missing the model granularity. Our spans would just say "azure" but we had three different deployments with different costs. The auto-instrumentation won't capture that unless you explicitly pass it as an attribute.

We started adding a custom span processor to inject details like `llm.model_variant` from our app config. It's a bit of extra code, but it makes the cost calculations later actually meaningful. Without that, you're just guessing.


cost first, then scale


   
ReplyQuote
(@amelia7k)
Estimable Member
Joined: 3 months ago
Posts: 120
 

Oh wow, thanks for laying this out so clearly. Coming from Airbyte, that pipeline view is exactly what I'm hoping to see too.

From what everyone's saying, the open source tools seem to get you close to that pseudo-trace diagram for debugging, which is really encouraging. But I'm a bit nervous about the cost part now. If the spans just say "azure" and not the specific model, how do you even start building that billing layer?

Can I ask a maybe obvious question? When you add a custom span processor to inject the model_variant, where does that code actually live? Is it in your main application, or somewhere else in the pipeline? Sorry if that's jargon-y, I'm still figuring this out. Thanks for all the patience



   
ReplyQuote
(@emilya)
Reputable Member
Joined: 2 months ago
Posts: 323
 

The custom processor lives in your main app code, where the tracing SDK initializes. You add it before your first LLM call.

The key is you need the model config at runtime. We pass it as a custom attribute from our LangChain chain constructor. Example: we tag the parent span with `deployment_id: "gpt-4-turbo-0125"` when we instantiate the chain.

Without that, you're right, the billing layer is guesswork. You can't map a generic 'azure' span to a price list.


Prove it with a benchmark.


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Open source tracing? Sure, if you enjoy a DIY project masquerading as a product.

You'll get the spans and traces. Visualizing them for debugging, like your pseudo example, is actually the one part that works. Tools like Phoenix give you that breakdown.

The "cost tracking" line is a joke, though. As others have hinted, you get token counts but no useful mapping. Your span might say "azure" and you'll need a crystal ball to know if it was GPT-3.5 or GPT-4. You have to build the billing layer yourself, and it's brittle.

Start with Phoenix if you want the picture. Accept that the cost part will be your own spreadsheet hell.


CRM is a necessary evil


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Totally. The visualization for debugging really is the killer feature. It's saved us from chasing ghosts more than once.

> custom tools or agents in your flows

We ran into this. Phoenix's default instrumentation often grouped our multi-step agent tool calls into a single span. We had to tweak the callback to log each decision as its own step, otherwise we'd lose the detail on *which* tool search was looping.


Trust the trial period.


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

You're right that token counts alone aren't a bill. The real trick is having a source of truth for pricing that you can version and audit.

We pull model rates directly from the provider's public pricing pages weekly and store them as configuration, tagging each price with an effective date. That way, when we calculate cost from a trace, we know we're using the rate that was actually in effect at the time of the call. Without that link to a validated rate sheet, you're just doing spreadsheet math on stale assumptions.


catdad


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

You've got the right starting point looking at OpenTelemetry-based options. For LangChain specifically, the auto-instrumentation hook is straightforward; you usually just import and initialize a tracer provider before your chain runs. The difficulty surfaces when you need custom attributes, like others have noted.

Regarding cost tracking across models: the open-source tools will give you token counts via the standard OpenTelemetry semantic conventions for LLMs (like `llm.token_count.prompt`). But the mapping from a span to a dollar cost is absent. You must build that enrichment layer yourself, and its accuracy hinges entirely on whether your traces contain the precise model identifier. From our benchmarks, if you don't explicitly set `llm.model` or a custom attribute like `deployment_id`, it defaults to a generic provider name, making cost aggregation meaningless.

On visualization, both OpenLLMetry and Phoenix can render the breakdown you sketched. The debug utility is real. We found Phoenix's UI slightly more polished for inspecting individual tool calls within a LangChain agent, but OpenLLMetry gives you raw control if you're already committed to an OpenTelemetry observability stack. The pseudo-trace you drew is achievable; just expect to write that custom span processor for model details.



   
ReplyQuote
Page 1 / 5