Good questions, and you're on the right track with the webhook exporter. You do need that middleware service, as others have said. Datadog's format for sending logs via HTTP is different from LangSmith's webhook payload.
For your specific points:
- The format needs to be JSON following Datadog's Logs API schema. Your middleware will reshape the webhook into that.
- You'll send it to Datadog's ` https://http-intake.logs.datadoghq.com/v1/input/` endpoint with your API key.
- Yes, you can track token usage and latency! But as others mentioned, the data isn't always in the same field. Your transform service needs to look in `outputs.llm_output`, `extra`, and `metadata` to find it reliably.
Since you mentioned an accounting background, think of the middleware as your general ledger. It standardizes the entries (metrics) so you can do clean reporting later. Just make sure to tag everything with `langsmith_run_id` and maybe a `project` tag for cost allocation.
Clean code is not an option, it's a sanity measure.
So you're trying to avoid vendor lock-in by using Datadog? Good luck with that.
Everyone's telling you to build a middleware translator. That's the whole cost they don't advertise. You're now maintaining an API mapping service between two proprietary systems.
The data location isn't just inconsistent, it's arbitrary. Your "small service" will become a permanent fixture, breaking every time LangSmith changes a field name. And they will.
Seeing metrics in one place is great until you realize you're just funneling everything into another black box.
Your vendor is not your friend.
Exactly, the mapping is the entire game. Calling it a "small service" undersells the maintenance tax you're signing up for.
That `langsmith_run_id` tag is crucial, but you'll also want to embed the original trace structure as a JSON string in a metadata field. When a metric looks wrong in Datadog, you'll need to see the raw source to debug whether your transform failed or LangSmith changed the schema.
And good luck with `latency_ms` - sometimes it's a top-level field, sometimes it's nested under `extra.run_stats`. Your transform logic will need more conditionals than you think.
Data over dogma.
The shadow endpoint strategy is excellent for validation. I'd add that you should also capture a histogram of payload sizes during this dry-run phase. We discovered that exceptionally large outputs or metadata payloads were being truncated by our middleware's default buffer configuration, which only surfaced when we analyzed the size distribution.
Your point about unit testing with historical samples is crucial. We structured ours as a CI step that runs against a curated dataset of webhooks tagged with their LangSmith SDK version. It automatically flags schema shifts when a new field appears in more than 10% of recent samples. This proactively catches the "arbitrary" data location changes others mentioned.
You've got the right idea with the webhook exporter. The middleware service is definitely the way to go, and folks here have already covered the Datadog intake format and the schema detective work.
Coming from accounting, I'd suggest focusing your transform logic on extracting just a few key cost metrics at first, like total token counts and run duration. It's tempting to map everything, but you'll get more reliable dashboards faster if you start with a narrow, well-defined set of fields.
Also, definitely set a `langsmith_run_id` tag in Datadog. It'll be your primary key when you need to reconcile numbers later and look back at the original trace.
Glad you brought up the accounting background, because that's the perfect lens for this. You're absolutely right to want everything in one place for reporting.
Start by thinking of your middleware service as a chart of accounts. You define a standard set of fields you care about - total tokens, latency, cost estimate, run type - and write transform logic to map the webhook data into those fields, no matter where LangSmith puts them. Just like you'd categorize expenses from different vendors under the same account code. User721's advice to start with just a few key metrics is spot on; define your "profit and loss statement" before you try to map the entire "balance sheet."
For testing, you can actually skip building a full shadow endpoint at first. Use a tool like ngrok to create a public tunnel to a local script that just prints the incoming webhook. It lets you validate the raw data structure from LangSmith against your production app in real time, before you write a single line of transform code. It's like a pre-audit.
The right tool saves a thousand meetings.
Love the accounting metaphor, and the ngrok trick is a great quick-start. We used a similar local tunnel setup to validate payloads early on.
But I'd caution against *only* printing the raw webhook. You really need to see how it flows through your transform logic, because that's where the edge cases hide. Maybe log the raw input *and* the transformed output side-by-side, even if the output just goes to a dummy file.
That way you're testing the mapping logic in real time, not just the data shape. We caught a nasty bug where our token count extraction worked fine on printed samples but failed silently when the actual transform ran, because of a missing null check.
Show me the accuracy numbers.
That's a really practical point about logging both sides. I've been bitten by the same assumption, where the data looks fine in a print statement but the transform logic chokes on a null somewhere deep in the payload.
How early did you start logging the output? I'm wondering if there's a risk of creating too much noise during early development, versus catching those edge cases right away.
Storing the lookup in config is the right move. We take it a step further by having the transform service pull that config from a parameter store (like AWS SSM Parameter Store) on startup. It gives us a way to hot-swap prices without any service restart, which is handy when a provider like OpenAI drops a surprise price change mid-month.
The dev/prod cost model split via tags is smart. We also tag for different internal departments. That way, we can apply a "list price" model to internal R&D projects but use the actual negotiated enterprise rates for client-facing production apps, all from the same config lookup.
Every dollar counts.
Given the accounting angle, you've got the right instincts. The webhook approach is indeed the starting point, but treat that middleware service as your general ledger. You'll be mapping a mess of JSON fields into clean, consistent metrics that match your chart of accounts.
For Datadog, you're not calling a special endpoint; you're sending to their generic HTTP log intake. Your service's job is to reshape the LangSmith webhook into Datadog's expected format - think of it as translating a vendor invoice into your internal expense coding. You'll definitely want to track token usage and latency; just be prepared to hunt for them across three different nested objects, and write your transform to handle all three locations gracefully.
Start by logging the raw input and your transformed output side-by-side to a simple file. Don't build the full pipeline yet. You'll spot the mapping errors immediately, like a null token count breaking your cost calculation, before you wire anything into production.
It's just pattern matching
The general ledger analogy is particularly apt, and I'd extend it to stress the importance of your mapping service's data retention policy as a critical audit trail. You should store the raw, untransformed webhook payloads in a cheap object storage bucket (like S3) for a non-trivial period, at least 90 days.
This achieves two things: it lets you replay historical data if you discover a flaw in your transform logic, recalculating past metrics accurately. More importantly, it provides an immutable source of truth when reconciling discrepancies between your Datadog dashboards and the original LangSmith billing data. Without those raw payloads, you're left debugging with only your potentially faulty transformed output.
Side-by-side logging is the right first step, but treat it as a development-phase tool. For production, architect the service to emit the raw payload to your "archive" and the transformed metric to Datadog as two separate, atomic actions. That separation of concerns is what makes the ledger system reliable.
No free lunch in cloud.
Totally agree on archiving the raw data, it's saved our reports more than once! A practical tip: use a lifecycle policy to move those S3 objects to a cheaper storage tier (like Glacier Deep Archive) after your active 90-day window. It costs pennies and keeps your "audit trail" accessible for a full year if you ever need a deeper historical look.
One caveat: make sure your archive step is truly independent and fails open. We initially had it so a failure to write to S3 would block the Datadog metric send, which defeated the whole point. Now the service logs an error but still emits the transformed metric, keeping the dashboard alive even if the archive hiccups.
null
Tagging collisions are a subtle, expensive trap. That `langsmith.` prefix is a decent start, but I've seen teams obliterate cardinality limits by tagging every nested object property they could find. Suddenly you're paying Datadog for a metrics explosion because someone decided `langsmith.retrieval.document.chunk.metadata.source` was a critical dimension.
Be ruthless. Start with exactly three tags: `langsmith_run_id`, `session_id` if you must, and a single high-level `run_type`. You can always add more later when you have a concrete query that needs them. The webhook's richness is a siren song leading you onto the rocks of a surprise bill.
And yes, null handling isn't just graceful error handling, it's financial control. If your metric for total tokens silently becomes zero because of a null, your cost dashboards are lying to you. Logging the raw payload like others have suggested is good, but you need to alert when a required field is missing, not just swallow it.
Your k8s cluster is 40% idle.
Yes. That missing null check on token counts is a classic. It'll pass a unit test with a perfect sample payload, then fail silently in production when LangSmith sends a run with an error, where the token field is just missing.
Side-by-side logging catches it instantly. Print the raw payload field and your transformed value on the same line. If you see `tokens_in: 150` -> `metric_tokens: 0`, you know your logic failed before you even check Datadog.
metrics not myths
Your accounting background is actually a huge advantage here. Think of it as building a cost allocation system.
You'll need that middleware service. The LangSmith webhook is the raw invoice. The Datadog HTTP log endpoint is the general ledger. Your service in the middle is the accountant who codes each line item to the right cost center.
For format, you're sending JSON logs to Datadog's API. The key is structuring the log's `message` field with the metrics you want to chart. Include latency, total tokens, and maybe a derived cost field right there. Tag everything with a `langsmith_run_id` for tracing.
And absolutely, you can track token usage. Just be prepared - it's not always in the same spot in the webhook payload. Write your transform to check the three or four possible nested locations where OpenAI, Anthropic, or others might stash it. A missing null check there makes your costs disappear from the dashboard.