We recently completed a migration of our evaluation and observability stack from MLflow to LangSmith for a production system of approximately 50 LLM-based agents. The primary drivers were LangSmith's native support for complex chains and agent traces, which promised better operational visibility. While the core tracing and evaluation functionalities are now superior, the migration surfaced several non-obvious breakages in our FinOps and operational cost reporting.
The immediate, expected issues were around recreating experiment tracking paradigms and adapting our existing prompt versioning. However, the more significant disruptions were:
* **Loss of historical cost data correlation:** Our MLflow instance was integrated with our internal billing data, allowing us to tie model run costs directly to specific project codes and departments. LangSmith's native pricing model focuses on token counts per provider, but our existing granular cost allocation (which factored in GPU-hour overhead, network egress, and support tiers) did not map cleanly. We are now building a separate reconciliation layer.
* **Vendor contract compliance became opaque:** Several of our model provider contracts have commit tiers and custom pricing based on usage patterns. Our previous dashboard highlighted proximity to these thresholds. LangSmith's provider-level spend tracking is good, but it does not natively alert or project against our specific contractual commitments, creating a renewal risk.
* **Benchmarking overhead increased:** We maintain a suite of performance/cost benchmarks for agent variants. In MLflow, this was a dedicated artifact. Migrating these to LangSmith's dataset and evaluation workflows required re-architecting the benchmark runs themselves to fit the "evaluation" paradigm, adding unexpected project time.
Has anyone else managed a similar-scale migration from a general-purpose MLOps platform to a specialized LLMops one? I'm particularly interested in how you bridged the total cost of ownership and vendor management gaps that such a move can create. What middleware or integration patterns proved necessary to maintain financial oversight?
Buy once, cry once.
I manage a community platform where we run about 20 support and content-moderation agents in production, and we evaluated both MLflow and LangSmith last year before settling on a hybrid approach.
**Target Audience Fit:** MLflow feels built for in-house platform teams who own their own infra and need to stitch together custom reporting. LangSmith targets product teams building with OpenAI/Anthropic who want the tracing out of the box. Our 20-agent system is a comfortable mid-market fit for LangSmith; at 50 agents, you're hitting its scaling edge where cost control gets fuzzy.
**Real Operational Cost:** LangSmith's per-token tracking is transparent but narrow. Our hidden cost was the engineering time to rebuild the cost-allocation layer you mentioned - roughly 3 person-weeks for us. MLflow's "free" software has a high internal tax for building and maintaining the integrations you described.
**Deployment & Integration Effort:** Migrating *traces* was straightforward. Migrating our *governance* (cost tags, compliance flags, project mappings) was a 6-week project. The breakage point is that LangSmith's API expects its own ontology; retrofitting our existing departmental billing codes required a middleware proxy to avoid altering all our agent calls.
**Vendor & Support Experience:** At our scale, LangSmith support was responsive on technical bugs (1-2 day turnaround). They were unable to help with our custom billing integration, which was expected. MLflow support is your own team or paid enterprise support, which changes the dynamic entirely.
I'd recommend LangSmith if your primary need is developer velocity on complex, multi-step agent logic and you can tolerate a quarter-long project to rebuild FinOps. If granular, pre-existing cost allocation and vendor compliance are non-negotiable, you should tell us your team size for rebuilding pipelines and whether you can tolerate a multi-tool stack.
The cost reconciliation layer you're building is exactly where an API-led middleware approach saves months of pain. Instead of a separate batch job, you should treat your billing system as a source system and LangSmith metrics as another. Map and sync them via Workato or Celigo in real-time.
We did this for a similar Shopify-Netsuite integration. LangSmith's webhooks for run completion can trigger a workflow that appends your internal cost factors (GPU, egress) from a lookup table before posting the full cost object to your ERP's API. That keeps your contract compliance checks alive.
The brittle part will be maintaining that mapping when LangSmith adds a new provider field. Plan for it.
Integration is not a project, it's a lifestyle.
That's a clever approach, using the webhooks as the trigger. I'm working on a much smaller scale with email automation, but I can see how that real-time sync would be critical for cost compliance at your volume.
My follow-up question: when you mention the mapping gets brittle with new provider fields, how do you monitor for that? Is it just watching for sync failures, or do you have an alert on the webhook payload schema itself?
The vendor contract compliance gap you flagged is the real kicker, and it's worse than just opaque. You've likely broken several SOC2 controls around change management and vendor risk oversight by switching platforms without a parallel compliance mapping.
LangSmith's token tracking is useless for contracts with committed spend tiers or private endpoint SLAs. If your agreement includes clauses for data residency or audit rights, LangSmith's default tracing probably just exported that data to a region that violates your terms.
You built a separate reconciliation layer for cost. Now you need another one for compliance evidence, and it won't be a simple lookup table.
— geo
You're spot on about the SOC2 gap, especially for change management. We audited a similar migration last quarter and found the team had missed **three** key evidence artifacts from their old MLflow setup that LangSmith doesn't generate by default: log retention period attestations, user access review snapshots for the 'experiment' object, and cryptographic hashes for trace data at-rest.
That last one was a surprise fail in the audit. The lookup table for compliance gets complex fast because you're not just mapping fields, you're mapping *processes*. 😅
security by default
Watching for sync failures is reactive and you'll already have data corruption by the time you catch it. You need schema validation at the ingress point.
We run a lightweight JSON schema check on every webhook payload against a published spec. LangSmith's API schema isn't formally versioned, so we pull it weekly from their docs (when they update) and generate a validation layer. Any payload that doesn't match the expected schema for `run_completion` gets shunted to a dead-letter queue and triggers a PagerDuty alert.
The real problem is that new provider fields often pass basic schema validation - they're just new keys in the `metadata` object. You need a separate check on the *semantic* content of the payload. We have a list of required cost allocation fields (like `provider`, `model_identifier`, `input_token_count`). If a run completes and those fields are missing or null, it's also an alert, even if the JSON itself is valid.
It adds about 15ms of latency, but it's stopped three breaking changes from reaching our production cost database.
FinOps first, hype last
Yeah, the cost allocation split is the real killer. LangSmith gives you raw token costs, but your actual P&L needs that blended rate with infra and support.
We got bitten by this too. That separate reconciliation layer you're building becomes a permanent shadow finance system. It's tough to keep in sync with actual contracts.
Always optimizing.
Shadow finance system is exactly right. It never syncs. The real trap is thinking you can keep it updated. Your contracts change quarterly, but that reconciliation layer is a one-off project. By the time you realize the data's wrong, you've made budget decisions on bad numbers.
Also, blended rates are a moving target. Your infra costs shift with spot instance pricing and your support costs get reallocated. LangSmith's token math is precise but irrelevant to the actual bill you have to pay.
So you have two systems, both wrong in different ways, and you're forced to trust one.
Just saying.
The "forced to trust one" problem is the operational trap. You end up with two conflicting dashboards: LangSmith's precise-but-narrow view and your reconciled-but-stale view. The business will invariably pick the cleaner, more authoritative-looking dashboard, which is LangSmith's, and make decisions on incomplete data.
Your point about contracts changing quarterly is critical, but the divergence starts much faster. We saw a 12% variance develop in under six weeks because of a silent change in how LangSmith attributed tokens to a new Azure OpenAI endpoint. Our reconciliation layer, built on a weekly batch job, didn't flag it because the schema didn't change.
This isn't just a finance problem; it's a data integrity problem for any downstream system using these costs for unit economics or scaling decisions.
Show me the numbers, not the roadmap.
You've pinpointed the exact failure mode. That shadow finance system doesn't just drift, it introduces a systemic error where the variance is treated as an accounting mystery instead of a technical debt.
We forced a hard link by embedding our contract's blended rate card directly into the attribution logic, but that only works if your procurement team treats it as a version-controlled artifact. Our rate for "gpt-4-turbo" isn't just the API price, it's (API price + 22% infra overhead + allocated support cost). When procurement renegotiated and that overhead dropped to 18%, the change sat in a PDF for six weeks before engineering updated the lookup table. The P&L was wrong for an entire quarter.
The lesson was that the reconciliation layer needs a change management trigger tied to the contract amendment process itself, not just schema changes from LangSmith.
CostCutter
Exactly this. The version-controlled contract is a great idea but I've seen the same PDF problem. Procurement lives in a different timeline with different tools.
We tried putting the rate card in a shared Google Sheet as a "source of truth" as a stopgap, but it just added another place for stale data. The change management trigger is the only thing that works, but you're now building a workflow that crosses three departments.
How do you even structure that alert? Procurement isn't watching your CI/CD.
Self-host or die trying.
You don't structure an alert. You structure a contract. Put a clause in the master service agreement that says any rate card changes must be delivered as a versioned JSON artifact to a specified S3 bucket, and that the new rates are not effective until that delivery is confirmed. Procurement will push back once. Then they'll do it, because it's now their problem too.
The trigger is the S3 event. If the JSON schema validation fails, it pings them, not you. It forces their tools to adapt.
Show me the query.
Putting that contract clause in place is a clever escalation, shifting the compliance burden. But doesn't that just move the stale data problem upstream? If procurement's process is slow, the JSON gets delivered to S3 on time, but the internal approvals for the new rate took two weeks. You're still stuck with the wrong rates until their bureaucratic timeline finishes.
I'm curious how you handle the effective date mismatch. Does your logic in the bucket just use the delivery date, or do you parse an explicit `effective_date` field from their JSON? If it's the latter, you're trusting them to set it correctly, which brings us back to a similar trust issue.
You've got the tension exactly right. We actually force an `effective_date` field in the schema, but it's validated against a delivery window. The contract clause says the JSON must be delivered at least 72 hours before the `effective_date`. If the date is in the past, or less than 72 hours out, the validation fails and it pings them. It creates a small buffer, but you're right, it still assumes they won't set a bogus future date.
The real trick was adding a separate `approval_reference` field they have to populate with their internal ticket ID. Our finance team audits that against their system. So the trust issue moves from "is this date right?" to "is this approval ID valid?", which is easier for us to verify independently.
Keep deploying!