Skip to content
Notifications
Clear all

Step-by-step: Connecting LangSmith to Datadog for APM-style dashboards

31 Posts
30 Users
0 Reactions
157 Views
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
Topic starter   [#22545]

I'm trying to get better visibility into my LangChain app's performance and costs. Since my team already uses Datadog for monitoring everything else, I wanted to see if I could send LangSmith tracing data there to build unified dashboards.

Has anyone set this up? I saw the webhook exporter in the docs, but I'm unclear on the specific steps. Mainly:
- What format does the data need to be in for Datadog?
- Is there a specific webhook endpoint or do I need a small middleware service?
- Can I track things like token usage per run alongside latency?

Any basic guidance would be really helpful. I'm coming from an accounting background, so seeing all the metrics in one place would make reporting much easier.



   
Quote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Webhook to an HTTP intake endpoint is the easy part. The hard part is mapping LangSmith's trace structure to Datadog spans and metrics. You'll need that middleware service.

Datadog wants JSON formatted specifically for their API. You'll be pulling `trace.tags` like `token_usage` and `latency_ms` from the webhook payload to create custom metrics.

Just pipe the webhook to a small Flask/FastAPI service that transforms and POSTs to ` https://http-intake.logs.datadoghq.com/api/v2/logs`. Tag everything with `langsmith_run_id`. Your cost reporting will thank you.


Prove it.


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

Yeah, that middleware mapping is the crucial piece. One thing I'd watch for is that the structure of a LangSmith run can get pretty nested, especially with things like retrieval steps. When you're flattening that into Datadog tags for metrics, you'll want to be consistent about naming. Like, prefixing everything with `langsmith.` has helped us avoid tag collisions.

Also, tagging with `langsmith_run_id` is perfect for correlation, but you might also want to pull in `session_id` if you're tracking user conversations. That's been a lifesaver for us when we're trying to piece together a problematic interaction.

The webhook payload itself is pretty rich, you can definitely get token usage and latency from it. Just be ready to handle some null values gracefully.


~Harry


   
ReplyQuote
(@brianw)
Reputable Member
Joined: 3 months ago
Posts: 242
 

The prefix idea is smart, especially for cost allocation. Without a clear namespace like `langsmith.`, those custom tags can get folded into generic AWS or container labels during billing exports, making the spend attribution a real mess later.

> you might also want to pull in `session_id`
This is critical for us, too. We map `session_id` to a cost center tag in Datadog, which lets us roll up token consumption and latency costs per user interaction or department. The nested run structure can actually help here if you propagate a `cost_center` tag from the parent run down through the child spans.


Spreadsheets or it didn't happen.


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

That's a really sharp point about propagating the `cost_center` tag from parent to child runs. It's the difference between spotty, high-level attribution and a truly accurate breakdown.

A small caveat on using `session_id` for billing, though. If a session is long-lived, the final cost roll-up might represent multiple different conversations or users if your app's logic reuses the session. It's often worth double-checking the business logic that creates that ID to be sure it aligns with your financial reporting period, be it per user, per chat thread, or per day.


Keep it civil, keep it real


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 3 months ago
Posts: 546
 

Yes, the webhook-to-middleware approach works. Since you're coming from accounting, that mapping layer others mentioned is like building a good chart of accounts - it makes all downstream reporting possible.

Start with the webhook payload structure in LangSmith's docs, then translate each run into a Datadog log entry. For your cost question, you absolutely can track token usage alongside latency. I send both as separate metrics from the same run data. The key is extracting `total_tokens` and `latency_ms` from the webhook's `outputs` and `extra` fields.

One tip: Datadog's metric namespaces help keep things tidy. I use `langsmith.runs.token_usage` and `langsmith.runs.latency` so they group cleanly on our main dashboard. Makes that single-pane-of-glass view you want much easier to build.



   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Agreed that structuring metrics like a chart of accounts is the right mental model. I found the `extra` field in the webhook payload to be inconsistent, though. Sometimes `latency_ms` is there, sometimes it's nested under `metadata`. We ended up creating a small transform function that checks multiple locations before falling back to a default.

On naming, `langsmith.runs.token_usage` is perfect for grouping. We also added a `langsmith.runs.cost_estimate` by pulling the token counts and applying a simple per-model lookup table. That metric feeds directly into a weekly cost dashboard.



   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

That's a smart approach to handle the inconsistency. We had the same issue with `total_tokens` appearing in different places, like `outputs` versus `extra`. A central transform function that looks in a priority order of fields before logging is definitely the way to go.

The per-model lookup table for cost is essential, but I'd add one caveat: those model rates can change. We ended up storing the lookup table in a config file that's decoupled from the transform service, so we can update prices without a redeploy. It also lets us run different cost models for dev vs prod based on tags.



   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

Absolutely, the middleware service is the core of this. Since user911 mentioned pulling tags like token_usage, I'd add that the payload location for those values can be inconsistent between runs, as several have pointed out. Your transform logic needs to check a priority order: `outputs`, then `extra`, then `metadata`, before falling back to a default.

Tagging with `langsmith_run_id` is perfect for traceability, but for cost allocation, I'd also recommend extracting any `session_id` or `project` field from the run's tags and propagating it as a Datadog tag. That way, your cost metrics can be grouped by session or project from the start, which is cleaner than trying to reconstruct it later from the run ID alone.


CloudCostHawk


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Yes, you need a middleware service to transform the webhook data into Datadog's specific JSON format. The docs show the webhook structure, but Datadog expects it at their HTTP logs intake endpoint.

Token usage and latency are in the payload, but their location varies. Your transform logic should check multiple fields like `outputs` and `extra` before logging to Datadog. Tag everything with `langsmith_run_id` for correlation.

Coming from accounting, think of the middleware as building your chart of accounts. Consistent tagging, like adding a `langsmith.` prefix to all metrics, is what lets you roll costs up cleanly in those unified dashboards later.


—AF


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

Totally on board with the middleware-as-chart-of-accounts analogy. A practical addition: we started versioning that transform schema itself (like `v1.2` of our mapping rules) and including that as a `schema_version` tag in the Datadog log. It saved us when we had to debug why old dashboards broke after a logic update.

Also, for the HTTP logs intake, setting up a dedicated `source` attribute like `langsmith-webhook` in the Datadog JSON makes creating exclusion filters in Logs Explorer way easier later on.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

The middleware service advice is solid, but calling it a "small" service is a bit optimistic. It's the most critical, fiddly piece of the whole setup.

You'll absolutely need it, because the webhook format and Datadog's intake API speak different languages. Think of it as building a bespoke translator for your metrics - one wrong mapping and your dashboard is gibberish.

You can track token usage and latency, but the data's hiding in different places each time. Your transform logic needs to be a detective, checking `outputs`, `extra`, and `metadata` before giving up. Welcome to the real cost of visibility 😉


Deploy with love


   
ReplyQuote
(@henryw)
Estimable Member
Joined: 3 months ago
Posts: 74
 

I just started looking into this too. The middleware service part worries me a bit. If the data location is inconsistent, how do you test your transform function is working right before connecting it all to Datadog? Is there a safe way to run a dry-run?



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 3 months ago
Posts: 546
 

Great question on testing! We actually set up a "shadow" endpoint in our middleware that writes to a local log file instead of Datadog for exactly this reason. You can point the LangSmith webhook at it, capture a bunch of real runs, and inspect the transformed output before going live.

Also, mocking the full Datadog JSON structure and running unit tests against historical webhook samples (you can export them) was a lifesaver. It helped us catch those inconsistencies in the `extra` field that others mentioned.



   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 297
 

That's such a smart testing strategy. The "shadow" endpoint idea is perfect for avoiding metric pollution while you're figuring out the mapping. We did something similar, but also built a small visual diff tool that compares the transformed JSON side-by-side with a sample of what we *expected* Datadog to receive. It flagged a ton of nested field issues we'd have missed just reading logs.

Exporting historical webhooks for unit tests is the other key. I'd recommend anyone doing this to start by collecting webhooks from runs that include both successful and errored states, since the payload structure can shift dramatically on a failure. That's where our first transform logic fell apart.

One caveat: if you're using the shadow log method, just watch your disk space if you have a high volume of runs! We learned that the hard way 😅


If it's not measurable, it's not marketing.


   
ReplyQuote
Page 1 / 3