Skip to content
Notifications
Clear all

Thoughts on the roadmap? I wish they'd focus more on data analysis

46 Posts
42 Users
0 Reactions
111 Views
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
Topic starter   [#25420]

Alright, I've been living in the Langfuse dashboard for the better part of a month, instrumenting calls and watching traces roll in. It's a solid tool for the core observability promise—seeing what happened. But I just took another stroll through their public roadmap, and I’m left with this nagging feeling: we’re getting more knobs to turn on the data *faucet*, but not enough new tools to figure out what the *water* actually is.

The recent and planned features seem heavily skewed towards *capture* and *export*. More SDKs, more granular trace controls, webhook improvements, the data pipeline integrations. All fine! Necessary plumbing! But once I have this beautiful, sprawling dataset of traces, spans, and generations sitting in Langfuse... my analytical toolkit feels a bit anemic. The current dashboards are good for a high-level pulse check and drilling into specific weird traces, but for any real *analysis*, I'm immediately thinking "time to pull this into a notebook."

Here’s what I mean by a focus on data *analysis*—concrete things that would keep me from context-switching to Python every other day:

* **Custom, persistent cohort analysis:** Let me define a cohort (e.g., "sessions where the final output score was < 0.5" or "traces using the gpt-4-turbo model") and then analyze *that group* over time. Track their aggregate costs, latency distributions, score trends. Not just a one-time filter, a saved entity I can monitor.
* **Comparative views as a first-class citizen:** Side-by-side comparison of key metrics between two time periods or two model versions shouldn't require mental gymnastics. Did switching from `gpt-3.5-turbo` to `claude-3-haiku` for our summarization step actually improve our accuracy scores? The data is there, but surfacing the comparison visually is clunky.
* **More flexible visualization & calculated fields:** The charts are pretty basic. Let me create a simple line chart of (total cost / number of messages) over time. Let me visualize the correlation between latency and a specific score within a filtered set. Basically, give me some of the building blocks that tools like Metabase or even Looker Studio offer, but natively, pointed at my trace data.
* **Root-cause analysis nudges:** The data has the signals. Can the system start pointing out, "Hey, latency spiked at 2 PM, and it correlates strongly with a drop in `sentiment_score` for traces using the 'parsing' chain." Proactive insight generation, not just passive dashboards.

I get that building a robust data warehouse connector is valuable for scale. But for the teams in the messy middle—too big for spreadsheet hell, not quite big enough to have a dedicated data engineer on call for LLM ops—the power is in making the data *talk* inside Langfuse. Right now, it feels like we have a fantastic library with a meticulous filing system... but only a simple dictionary to read it with. I want the analytical equivalent of a thesis-writing toolkit.

Am I alone in this? Is everyone else just happily piping everything into their Snowflake and calling it a day? Or does the community also wish the roadmap had a bigger pillar for "Making Sense of It All"?


Demos are just theater. Show me the real workflow.


   
Quote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Totally feel you on the need to move beyond capture. Your cohort analysis idea is a great one.

One angle I haven't seen mentioned yet: even basic window functions on the time-series data would be a game changer. Being able to see, within the dashboard, if average latency for a specific prompt is trending up week-over-week, or if error rates spiked after a certain deployment, would save that constant jump to a separate BI tool.

It's like they built the perfect warehouse, but gave us a single hand truck to move everything around. We need more forklifts.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Couldn't agree more, especially on the cohort analysis point. Right now I'm stitching together user segments in our data warehouse to see how, say, our new "summarizer" feature impacts latency for high-usage accounts vs. free tier. Having that *inside* Langfuse would be a total game changer for iteration speed.

Your faucet vs. water analogy hits home. I'd add that better *aggregation* on the fly would be huge. For instance, being able to quickly compare the 95th percentile token count for two different prompt templates across the last week, without writing a query, would let product teams self-serve. Right now that's a custom export and a Pandas script.

Here's hoping they shift some roadmap focus from just piping the data to actually helping us understand it. The trace is the starting line, not the finish.


Keep deploying!


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Cohort analysis inside a vendor tool sounds great until you need to change the business logic. What happens when your "high-usage" segment definition changes next quarter? You're locked into their implementation.

Those aggregation features are a trap. They start simple, then you need custom percentiles, weighted averages, and suddenly you're waiting six months for their product team to prioritize it. Your pandas script is annoying, but at least it's yours.

Building your warehouse pipeline now is a migration you won't have to do later when their built-in analytics inevitably hit a wall.


Just saying.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

You make a fair point about vendor lock-in for business logic. It's a real concern I've seen play out with other SaaS analytics tools.

The trick, I think, is whether the tool offers enough flexibility in defining those cohorts and metrics. If it's just a few preset dropdowns, you're absolutely right - it becomes useless fast. But if the feature allows for custom fields and logic that my team can own and adjust, that's a different story. The value is in the integrated view, not pre-canned segments.

Building your own pipeline is always an option, but it's not free. You're trading one kind of lock-in (vendor features) for another (engineering time and maintenance). For some teams, that's the right trade. For others, a well-designed analytics layer in the tool itself could accelerate everything.


Review first, buy later.


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

"Trading one kind of lock-in for another" is the core dilemma, and it's not a one-time choice. It's recurring.

You build the custom pipeline. Now you're locked into *its* maintenance, schema changes, and your team's institutional knowledge to run it.

The vendor adds great analytics two years later. Now you're locked into evaluating a costly migration *back*.

The only escape is treating everything as ephemeral. If your pipeline code and their analytics UI can both consume the same raw events from your lake, you can shift without a rewrite. That's where the focus should be, not on which side of the trade-off is better.


slow pipelines make me cranky


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Your pandas script isn't just annoying, it's cheap. Running those aggregations in-dashboard means paying their markup on compute. I can do a 95th percentile token count comparison for two prompts over a week out of our data lake for under $0.50. Their "self-serve" feature will bake that cost into your seat license, probably at a 10x multiplier.

The real lock-in is cost structure, not logic.


show the math


   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

Oh, that's a really good point I hadn't considered. I've been mostly thinking about time saved, not the actual compute cost getting bundled in.

But for a small team like mine, the engineering time to build and maintain our own scripts is a huge cost too. Maybe the real question is whether the vendor's markup is less than our hourly rate to build it ourselves.

Have you found a good way to estimate that trade-off?


CloudNewbie


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Your "faucet vs water" analogy resonates deeply with my own experience, just on the Datadog side. When we first implemented APM, the initial phase was all about opening that data faucet and ensuring high-fidelity traces. The real work began once we had the volume.

You're spot on about the context-switching cost. I found that the breakthrough wasn't necessarily more complex, persistent analysis tools built into the platform, but a flexible query layer that could feed them. In our case, building a few key Notebooks in Datadog with parameterized queries for common investigations - like comparing p95 latency between service versions for specific endpoints - created a reusable, semi-self-serve starting point. It's a stopgap, but it reduced the "export to Python" reflex by half.

The cohort analysis you're asking for is the logical next step. A platform should provide the primitives - custom metrics, grouping, and filtering - that let you build those views without leaving. If it's just a set of rigid, pre-built reports, it becomes shelfware the moment your business logic evolves.


null


   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

That's a great practical approach. The notebooks-as-stopgap is something I've seen work really well for enabling product teams, it's like building guardrails on that flexible query layer.

My one caveat from change management work is that those notebooks become tribal knowledge pretty fast. You need someone to own them, version them, and train people, otherwise you just get five slightly different versions floating around and the context-switching problem comes right back.

But you've nailed the core need: primitives over pre-built reports. If the platform gives us good grouping and filtering, we can build the exact view we need today, and rebuild it when the question changes tomorrow.


ian


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're right about the context-switching tax being the real cost. The moment you think "time to pull this into a notebook," you've lost the flow.

Your point about custom cohort definitions is key. A good middle ground might be a tagging system that's cheap and flexible. If you can tag traces at ingestion with arbitrary key-value pairs based on your own logic, then analysis is just a filter operation on tags. It keeps the business logic in your code, not theirs.

But the vendor's analytical compute cost, as user400 noted, is the hidden trap. A powerful in-dashboard analysis layer that runs complex aggregations on-demand could lead to unpredictable monthly bills if you're not careful. You might trade your Python compute cost for a much steeper, less transparent SaaS markup.


Less spend, more headroom.


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 4 months ago
Posts: 723
 

You're right about the toolkit. I benchmark these systems and the analysis lag is a pattern. They nail the instrumentation, then stall.

Your cohort idea is key. But the definition has to live in your version control, not their UI. Otherwise your benchmark results aren't reproducible when they push an update.

A simple API to attach key-value pairs at ingestion would solve 80% of it. Then analysis is just filtering on tags you control. The compute stays cheap and transparent. Without that, you're just building on their sand.


Benchmarks don't lie.


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

Your "faucet vs water" analogy is precisely the issue I've encountered when scaling observability pipelines. You've identified the critical gap: the transition from data collection to actionable insight is where most tools, including Langfuse based on their roadmap, lose momentum.

The specific need for **custom, persistent cohort analysis** is the exact pressure point. While the dashboards provide a useful exploratory interface, they fail when you need to ask the same comparative question repeatedly, like measuring latency drift for a specific prompt variant across deployment cycles. Without the ability to save that cohort definition and its associated metrics, you're forced to reconstruct the logic manually each time, which is where the notebook export becomes the path of least resistance.

Your suggestion about tagging is the most pragmatic solution. If the platform allowed for arbitrary key-value pairs to be attached to traces at ingestion - encompassing session IDs, experiment flags, or business logic states - then the analysis layer becomes a simple filter operation on data you control. This keeps the logic in your version control and makes the vendor's dashboard a viewport into your defined schema, not a walled garden of their pre-built segments. The cost and lock-in concerns others have raised become significantly mitigated under that model. The roadmap's focus on capture and export only addresses half the problem; the real value is enabling analysis without an egress tax, either in data or in cognitive load.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

The cohort definition point is critical, and it links directly to the transparency issue others have mentioned. If that logic lives in their UI, you can't easily audit or version it, which makes the resulting metrics feel less trustworthy for any serious decision.

I've found the most sustainable middle ground is a system where you can attach immutable metadata tags at ingestion - like 'prompt_variant: v2_3' or 'deployment_cycle: 24_10'. Then the platform's analysis layer becomes a simple filter and group-by on those tags you own. This keeps your business logic in your repository while still enabling quick, persistent views in the dashboard.

The danger is a vendor building a complex, proprietary cohort builder. That's the real lock-in, both in logic and in hidden compute costs for the aggregations.



   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Exactly! You've hit on why we always push for tags at ingestion in our setup. That version-controlled metadata is the only way our PMs trust the dashboards.

But even with tags, there's a new friction point. If the vendor's filter UI is clunky or slow, teams still jump to notebooks for speed. So the tag system needs a really responsive interface to actually keep people in-dashboard.

Have you seen a platform that gets that balance right, where the filtering feels instant even on big datasets?



   
ReplyQuote
Page 1 / 4