Skip to content
Notifications
Clear all

Langfuse after 18 months in a fintech startup - the bad parts

2 Posts
2 Users
0 Reactions
41 Views
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
Topic starter   [#17704]

Okay, so we've been running Langfuse for observability on our LLM pipelines for about a year and a half now. Overall, it's a solid product that solved a huge visibility problem for us early on. But as we've scaled and our needs have gotten more specific, some rough edges have really started to show. I'm focusing on the pain points here because the good stuff is covered a lot.

The biggest issue for us is the lock-in with their specific tracing paradigm. Our pipelines are complex—think multi-agent workflows with conditional branching and external tool calls. Langfuse's traces, spans, generations hierarchy feels rigid. Trying to map our execution flow into their model sometimes forces us to either:
* Flatten logic that shouldn't be flattened, losing the real parent/child relationships.
* Create overly nested structures that become a nightmare to query later.
The SDKs feel optimized for simpler, linear chat flows. When you step outside that, you're fighting the abstraction.

Then there's the querying and dashboarding. For deep, investigative debugging, the UI gets sluggish with high-volume traces. We often end up exporting data to run our own analysis, which defeats the purpose. The pricing model also starts to pinch when you're debugging a high-throughput system—you think twice about instrumenting every single helper function because of the volume cost.

Finally, the deployment and integration story has some gaps:
* The Helm chart for self-hosting requires a lot of manual tweaking for a production-grade K8s setup (resource limits, persistent volume configurations, etc.).
* The lack of a native OpenTelemetry collector integration means we have to maintain separate instrumentation paths for our LLM apps vs. everything else.
* Alerting is basic. We ended up building a separate system to monitor trace metrics and latency percentiles from the exported data.

Don't get me wrong—it's still a critical part of our stack. But I'm hoping the team starts addressing some of these scale and flexibility issues soon. Curious if other teams in complex domains have hit similar walls.


Ship fast, measure faster.


   
Quote
(@davidn)
Reputable Member
Joined: 3 months ago
Posts: 305
 

You're right about the mapping struggle. We hit a similar wall with complex order-fulfillment workflows in our ERP. The trace/span model assumes a single, linear "conversation," but real business logic often has parallel branches and independent side effects.

One workaround we used was to abuse the trace_id as a "session" identifier and treat spans as independent action logs, but then you lose the hierarchical view entirely. The query performance you mentioned became our breaking point - trying to filter for a specific error type across thousands of "flattened" spans was unusable.

Did you explore their newer dataset and evaluation features? We found they doubled down on the linear chat use case even more there, which pushed us to build a custom exporter to a simpler log aggregation system.


Measure twice, buy once.


   
ReplyQuote