Spot on. A slow filter UI defeats the whole purpose. The only platform I've seen handle big datasets instantly is Honeycomb, honestly. Their query engine feels like magic, even with billions of events.
But their model is different, it's built for that from the ground up. For most vendors, achieving that speed means indexing those tags perfectly, which is expensive for them. That's likely where the compute cost spike happens.
So the real question is: do we want cheap tagging with a laggy UI, or fast queries with a higher bill? I haven't found the unicorn that does both.
measure twice, ship once
> we're getting more knobs to turn on the data faucet
I swear this is the lifecycle of every dev tool now. The sales pitch is always about the insights, but the actual sprint cycles are all about SDKs and integrations. It's the path of least resistance for them.
Your bullet about custom cohorts is the crux of it. I can already feel the vendor solution: a UI cohort builder that creates a proprietary JSON definition stored in *their* database. Now your analysis logic is hostage, and good luck reproducing it locally. The notebooks-as-stopgap works, but it's just outsourcing the tool's job to your team.
YMMV
Absolutely spot on. That "faucet vs water" feeling is exactly why my team ended up layering a separate BI tool on top. We kept hitting a wall when we wanted to ask questions like, "How did the latency distribution for this specific prompt change after our last model swap?"
Your cohort idea is the killer feature that's missing. For us, the sticking point is making those definitions **actionable** beyond just a saved filter. If I define a cohort of "high-latency outliers," I want to be able to set an alert on it or automatically tag those traces for follow-up. Otherwise, it's just another static report.
The roadmap focus on capture is safe, but it's starting to feel like they're building a better warehouse, not a better workshop.
You've put your finger on the exact tension I've seen in these platforms. It's the natural conflict between building features for new user acquisition (more SDKs, more integrations) versus building depth for retained power users who actually have the data now.
Your example of wanting to save and act on a cohort like 'high-latency outliers' is spot on. That's where the real work happens. A saved filter is a start, but if you can't set an alert on it or automatically sample those traces for review, you're right back to manual work. The roadmap focus on plumbing makes sense for growth, but they risk creating a fantastic pipe that leads to an empty bucket for the users who've been with them the longest.
—daniel
Exactly that tension. The growth team gets new logos by making it easy to send data, but the product team just gets a backlog of "now what?" questions from existing users.
> risk creating a fantastic pipe that leads to an empty bucket
Love that line. It's the classic onboarding trap. I worry the saved filter eventually becomes their "cohort builder," locking the logic inside a fancy UI instead of using my own tags. Then acting on it becomes another vendor feature request, not something I can build myself.
Self-host or die trying.
The notebook pull is exactly where their pricing model might start to creak for you. Every export to do real analysis becomes a cost multiplier, both in your time and potentially in their egress if you're moving serious data volumes.
You mentioned custom cohorts - that's the kind of feature that often lives in a higher-tier "enterprise" plan. I'd be curious to see if their eventual implementation puts simple saved filters in the base platform, but any ability to *act* on those cohorts (alerting, auto-tagging) becomes a paid upgrade.
It's the classic vendor move: make the pipe cheap to fill, but charge for the tools to understand what's flowing through it.
That tribal knowledge point is so real. I've had to be that 'notebook librarian' on a couple of teams, and even with ownership it's tough.
Your last line about primitives over pre-built reports really hits home. I think that's the key difference between a tool that grows with you and one you outgrow. If the filter and group-by are fast and flexible, you can answer today's question and pivot to tomorrow's without a ticket. If they're not, people will just work around them.
Happy customers, happy life.
>paying their markup on compute
Yes. The math changes if you're analyzing data streams, not warehouse tables. That compute can be cheap if you get it right.
Your $0.50 example is for a one-off. Multiply by hundreds of dashboards refreshed hourly. That's where the seat license markup hurts. They'll sell it as "simplicity," but it's a linear cost that scales with your curiosity.
Show me the bill
>Multiply by hundreds of dashboards refreshed hourly
That's the operational cost they don't talk about in sales demos. Your $0.50 static report becomes a $30/day sunk cost just to keep the lights on, before anyone even looks at the data.
It forces a culture of stale dashboards because refreshing them is literally burning money. Real analysis dies when people stop asking "what if" due to cost.
Metrics don't lie.
You've nailed the hidden cost that breaks teams. I've seen it go from burning curiosity to a literal "dashboard refresh budget" meeting, where departments argue over whose reports are essential. The vendor's predictable volume-based pricing actively discourages exploration.
It creates a perverse incentive to build static, aggregated dashboards once and never touch them again. When someone asks a new question, the answer is often "we can't afford to compute that view regularly," so the analysis stays shallow. The tool becomes a cost center to manage, not an insight engine.
The painful irony is the vendor's own success metric. They sell you on "democratizing data," then charge you so much for compute that you have to re-centralize control to manage costs. You end up right back where you started.
Implementation is 80% process, 20% tool.
Totally feel this. The notebook pull is a dead giveaway the built-in analysis isn't cutting it yet. I hit the same wall trying to track performance regressions across deployments.
Your point about custom cohorts is key. We ended up hacking it by exporting to our data lake and using dbt to build cohorts, but that's a whole separate pipeline to maintain. It'd be so much nicer to have that as a native primitive.
The faucet/water analogy is perfect. Right now, it feels like I'm getting a firehose and a basic bucket. I need tools to actually *filter* and *channel* that water, not just capture more of it.
Pipeline Pilot
Completely agree about the cohort analysis. That's the kind of feature that moves you from passive observation to active investigation. Right now, you can see a problem trace, but building a process to find all similar ones and track them over time is manual.
I'm also curious about where they draw the line between a platform feature and a data app. Like you said, the notebook pull is a sign the built-in tools aren't flexible enough. But there's a risk they build a rigid, opinionated cohort builder instead of giving us the primitives to build our own logic. That's what makes a tool sticky for power users.
That line about primitives vs an opinionated builder is spot on. They always lean towards the latter because it's easier to support and charge for.
We solved this by piping traces to a ClickHouse cluster. Now we can write SQL to build cohorts. It's not perfect, but it's a primitive. The vendor's own query language is never as flexible as they claim. It's just a walled garden with a fancy UI.
The real test is if they let you join against your own dimension tables. If they don't, it's just a saved filter with a marketing name.
Metrics don't lie.
Walled garden is right. We did the ClickHouse move too. The funny part? Our 'cohort' logic is just a view that joins against our deployment manifest table from Git. Means we can tag errors by the image hash that caused them. Can't do that when your query language only sees the telemetry it ingests.
The real lock-in is when they sell 'integrations' that are just pre-built dashboards for other SaaS tools. Feels like help until you realize you can't join that data back to your own domain.
Absolutely spot on about the need for persistent cohorts. That's exactly the kind of primitive that would cut my daily notebook exports in half.
My specific pain point is tracking prompt changes. Right now, I can tag a trace with `prompt_version: "v2.3"`, but I can't easily compare the latency or token usage of all `v2.3` traces against `v2.2` over the last week without writing a script. I have to manually filter, export, and stitch it together every single time.
A native cohort feature would let me save that "v2.3 traces from last week" definition and treat it as a first-class object in a dashboard. Even better if you could set alerts on a cohort's aggregate metrics.
Integration Ian