We switched from Langfuse to Traceloop six months ago to get OpenTelemetry-native tracing. The promise was solid, but the migration broke three key things in production. Here's the damage report.
**What Broke:**
* **Cost spikes from high-cardinality attributes:** Traceloop's OTel pipeline ingested everything. Our `llm.prompts` with full user context exploded cardinality. Our observability backend's costs jumped 40% before we added aggressive attribute filtering.
```python
# Had to add this processor to every agent
SpanProcessor = SimpleSpanProcessor(
AttributeFilter(
reject=['llm.prompts.*.user.context.*'] # Too many unique values
)
)
```
* **Missing LLM token counts:** Langfuse's SDK derived this automatically. Traceloop's OTel instrumentation for OpenAI/AWS Bedrock often reported zero for `llm.usage.*` metrics. We had to patch the instrumentation to read the response body directly.
* **Dashboard migration headache:** Our Langfuse dashboards for agent loop efficiency (steps/time) weren't portable. Rebuilding them in Grafana required rewriting all queries against the OTel data model, which took two sprints.
**The Win:**
Debugging improved massively. Having traces in Jaeger directly linked to our application metrics (Prometheus) and logs (Loki) let us pinpoint bottlenecks we couldn't see before. The vendor lock-in reduction is real.
Bottom line: The core tracing is superior, but prepare for a 2-3 month stabilization period to handle data pipeline and visualization gaps. Don't underestimate the config overhead.