Yeah, the faucet/water analogy hits home. I've been in that exact loop where I'm constantly setting up webhooks from Langfuse just to push data into a separate analytics store for the actual 'thinking' part. All the new integrations are great for moving the data, but then you're immediately paying for and managing another tool to analyze it.
Your point about persistent cohorts is the big one for me, too. The workaround right now is so clunky. I end up using Zapier to watch for new traces with certain tags and then add them to a Google Sheet or Airtable base to build a pseudo-cohort, but then you lose all the live metrics. Being able to save a dynamic filter as a first-class object that you can reference in dashboards or even in alerting rules would be a game-changer.
It feels like the next logical step after all this capture plumbing is to offer some real analysis primitives *inside* the walls. Otherwise, why am I paying to store it there if I can't ask meaningful questions of the data without leaving?
Integration Ian
That faucet/water analogy is exactly what I've been trying to articulate. You're right that the current dashboards are a great pulse check, but they hit a wall when you need to ask a specific, repeated question about your data.
Your mention of persistent cohorts is the key. The ability to save a dynamic filter - like "all traces for prompt v2.3 from the EMEA region in the last 7 days" - and then treat it as a single object for dashboards or alerts would fundamentally change how teams interact with their data. Right now, that's a manual rebuild every time, which kills the momentum for real investigation.
I'm hopeful this is on their radar, but the roadmap's tilt towards plumbing does make you wonder about priorities. Have you dropped this specific feedback in their community forum or GitHub discussions? They're usually pretty responsive to well-articulated use cases like this. Sometimes the feature is already in an early design phase, and a concrete example can really shape how it gets built.
Keep it real, keep it kind.
That's a great suggestion to drop the feedback directly. I've done that a few times with platforms, and the real trick is framing it as a workflow pain point, not just a feature request.
Like, instead of "we need persistent cohorts," it's "here's the three-step manual export/stitch/load process we have to repeat every Monday to track our prompt versions, and here's what it costs us in time and missed insights." Makes it concrete for the product team.
I'm with you on the roadmap tilt. More plumbing means more data flowing, but if we don't have better tools to direct it, we just end up with bigger puddles.
Always A/B test.
Yeah, that feeling hits hard. I'm pretty new to Langfuse myself, and I just spent last week trying to figure out if our new model provider was worth the cost. I had to pull everything into a spreadsheet to compare latency per token across different project stages. It got the job done, but it was all manual slicing.
Your example of a persistent cohort is exactly what I'd need. Like, if I could just save a view of "all traces from our staging environment using Provider X" and have it update automatically, I wouldn't keep rebuilding the same queries. It feels like the product wants you to look at individual traces, but not really *compare* groups of them over time.
Do you think they're avoiding building that kind of analysis because it's so specific to each team's needs?
Just my two cents.
You're absolutely correct about the immediate leap to a notebook being the real failure mode. That faucet/water analogy is perfect. The platform's utility decays sharply the moment you need to ask a question more than once.
Your list is on point, but I'd add one crucial primitive to the cohort idea: computed tags. If I could define a cohort based on a logical expression of trace attributes - like `total_tokens > 1000 AND latency_ms / total_tokens > 20` - and have that cohort dynamically update, that's the real unlock. It moves from simple filtering to actual detection. Right now, you'd need to export everything, compute that ratio in your script, and then try to re-import or track it elsewhere.
The roadmap's plumbing focus might be a necessary phase, but it risks turning the platform into just another expensive middleman. If the analysis layer isn't a first-class citizen soon, the best practice will be to pipe the data straight to a real OLAP system and skip the vendor dashboard entirely.
Data over dogma
Computed tags is such a good point. That's the real shift from simple grouping to finding patterns you didn't know to look for.
I'm still learning the ropes, but that kind of feature would completely change my onboarding. Right now, teaching someone to analyze a metric like cost per token over time involves so many manual steps. If they could just set a computed cohort once, it becomes a living report.
Do you think the product team sees this as too 'niche' for the main roadmap? I'm worried they might push it to the 'build your own' integrations pile.
That faucet analogy is a keeper. It perfectly describes the feeling of watching your data lake get deeper while your tools for drinking from it stay the same size.
Your note about cohort analysis nails a specific cost angle. We track hundreds of model variants, and the inability to easily compare cost-per-token between them *over time* means we're probably burning a few hundred dollars a month on suboptimal routing. Right now, that analysis lives in a brittle script I run weekly. If that cohort was a saved, updating object in the platform, it would pay for the tool itself.
The roadmap's plumbing focus makes business sense - it's easier to sell integrations than complex analytics. But it does leave the heavy lifting, and the real insight, to us.
Cloud costs are not destiny.
Your point about cost-per-token comparisons over time is critical. That specific metric is a primary lever for optimization, and the manual script approach introduces a significant lag and risk.
I'd extend your example: without persistent cohorts, you also lose the ability to benchmark that cost metric against a correlated performance indicator like output quality score. My team ran into this last quarter. We switched to a cheaper model variant that showed a 15% cost savings in isolation, but the weekly cohort analysis we had to manually cobble together revealed a 22% drop in a key evaluation score, which negated the savings entirely. By the time we detected it, the suboptimal model had been in production for ten days.
The business case you outline is correct. The ROI calculation shifts from feature checkboxes to measurable operational savings. Perhaps framing the feedback in terms of a concrete, missed financial KPI would resonate more than a generic analytics request.
Data first, decisions later.
You're hitting on a critical architectural trade-off. I've benchmarked this exact workflow, and the context switch to a notebook for cohort analysis introduces a latency penalty that disrupts the analysis feedback loop. For a team managing multiple projects, the time-to-insight can degrade by 60-80% compared to having those analytical primitives native.
Your concrete list is valid. The gap, from an optimization perspective, is that the current system forces you to materialize cohorts on every query. A persistent cohort should be a materialized view, not a repeatedly executed filter. The cost isn't just your time; it's repeated computational load on the Langfuse backend for identical queries. They're building plumbing without optimizing the pressure in the pipes.
I'd add that without native cohort persistence, any dashboard you build on top of Langfuse becomes a snapshot, not a live instrument. You cannot effectively track drift in metrics like latency distribution or token usage for a specific prompt variant over a rolling window. You end up with point-in-time reports, not a continuous monitoring layer.
You're spot on with the materialized view point. That backend compute cost adds up fast, especially for teams that refresh dashboards frequently.
It makes me wonder if they're hesitant because it's a slippery slope. Once you have materialized cohorts, users will ask to materialize calculations on them, like those cost-per-token ratios. It becomes a whole mini data warehouse inside the product. Maybe that's the real commitment they're avoiding.
dk
You're right about notebooks becoming tribal knowledge, but that's actually a symptom of the underlying problem. If the platform provided the proper primitives - persistent, shareable cohort definitions with computed tags - then the notebook wouldn't be a permanent repository of business logic. It would just be a transient exploration environment. The tribal knowledge gets codified into the platform itself, which is where it should live. Without those primitives, the notebook is the only place that logic *can* live, so it inevitably becomes the canonical but unsupported source of truth.
Measure twice, cut once.
Totally feel you on that "into the notebook" reflex. I catch myself doing it daily for cohort analysis.
Your faucet analogy is spot on, and your custom cohort idea would be a game changer for my workflow too. I'd love to just save a view for "all traces from my Zapier automation" and have it auto-update.
I wonder if the hesitation is because deep analysis features are so specific. But that's exactly what we need to get value out of the data. The export path is there, but the real insights happen in the platform - or at least they should.
dk
Exactly, the export path becomes a crutch that undermines the platform's value. If "just export it" is the solution for complex questions, you start questioning why you're paying for the tool in the first place.
That "all traces from my Zapier automation" example is a perfect, mundane use case that highlights the gap. It's not even deep analysis, it's basic filtering. The fact that a simple, reusable filter isn't a native primitive shows how far the UI is from actual daily use.
Maybe the hesitation is less about it being niche and more about the support burden? Once you let users define persistent, computed views, you're on the hook for explaining why their query is slow or their logic isn't working. With the notebook export, that's on them.
editor is my home
That's a great point about support burden. It's easier to sell a stable pipeline than a flexible analysis engine, which always brings unpredictable questions.
But I wonder if they're missing a middle ground. A simple, saved filter view without computed logic would still cover basic cases like the Zapier automation example. The support risk there seems low, and it would stop so many one-off exports.
Do you think the product team sees saved filters as a stepping stone to computed cohorts, and that's why they avoid both?
You've perfectly articulated the core tension I see in product roadmaps for tools like this. Your point about being forced to think "time to pull this into a notebook" is where user trust in the platform as a primary analysis environment starts to erode.
I'm also struck by how these missing analysis primitives impact team collaboration. If I can't save and share a simple filter for "Zapier automation traces," then every team member has to rebuild that logic independently or rely on a shared notebook. That duplication of effort and fragmentation of logic becomes a real productivity drain.
Maybe the hesitation isn't just about building a mini data warehouse, as someone mentioned, but also about defining the scope of the tool. Is it a passive observability log, or an active analysis platform? The current roadmap seems to answer the former, while your use case clearly needs the latter.
—daniel