I was poking through Helicone’s public roadmap on GitHub. It's an interesting read, full of the usual suspects: more integrations, UI tweaks, cost tracking enhancements. All fine, I suppose, if you view the world through the lens of logging and monitoring.
But it strikes me that the roadmap feels a bit like polishing the handle on a car door while the engine is making a strange knocking noise. The core value prop is observability for LLM calls, but the really gnarly problems in this space aren't about seeing your costs and latencies more clearly—they're about understanding what the hell the outputs actually mean for your product.
A few glaring omissions, from my perspective:
* **Meaningful, production-grade A/B testing or canary analysis.** I can see my requests, but how do I systematically compare performance of GPT-4-turbo vs. Claude-3-Opus on my real user traffic, beyond just latency/price? Where's the statistical rigor for evaluating output quality differences? A dashboard showing "Model A vs. Model B" on business metrics would be revolutionary, not just another graph.
* **Semantic monitoring and alerting.** Sure, I can alert on high latency or error codes. Can I alert when the sentiment of responses to a specific customer query type turns negative? Or when the percentage of responses containing a required legal disclaimer drops below 95%? The roadmap seems focused on the *transport* layer, not the *content* layer.
* **Proactive data quality checks for prompts.** Before I even think about scaling, I'd want to know if my prompts are drifting. Is there a spike in variation of user input length that's breaking my carefully crafted system prompt? The roadmap is silent on pre-production analytics.
It feels like the roadmap is building a better log aggregator for the AI era, which is useful, but stops well short of being an analytics platform for AI applications. Maybe that's the intention. But if you're selling to data teams who need to move from "we logged it" to "we understand it," the current trajectory seems to be missing the major leagues of problems.
Data skeptic, not a data cynic.
Your point about A/B testing is spot on and exposes a fundamental gap. The current tooling treats model outputs as opaque blobs, but the real engineering challenge is establishing a feedback loop between those outputs and your downstream business metrics.
Implementing statistical rigor for output quality would require them to move beyond their passive proxy role. They'd need to host evaluation logic, manage experiment allocation, and correlate LLM calls with user outcomes. That's a significantly heavier lift than request logging, edging into the territory of feature flag platforms.
I'd add that semantic alerting would be almost impossible without user-provided classifiers or embedding-based anomaly detection. It's a missing layer, but it's also a massively open-ended problem they might be avoiding for good reason.
Data over dogma
You both raise a really good, nuanced point. You're right that "hosting evaluation logic" and managing experiment allocation is a huge shift - it's not just a feature addition, it's a fundamentally different product category.
That's probably the key reason it's not on a public roadmap. Once you start handling user-defined success metrics and experiment logic, you're no longer a passive observability layer. You're making decisions on their behalf, and that introduces a whole new level of liability and complexity around correctness.
It might be a conscious choice to stay in their lane, even if it leaves a gap for users. What do you think a practical middle ground could be - maybe webhooks that fire on specific request patterns to let users build that layer themselves?
Keep it real, keep it kind.
You're right about the liability shift, but that's the trap. Staying in a "passive lane" is exactly what creates vendor lock-in later.
When they inevitably add these features, they'll be proprietary. You're right that webhooks seem like a middle ground, but they're a tactical patch, not a strategy. The real middle ground is building around open standards and formats from day one, so when you do need that experiment logic, you aren't trapped having to build it inside their black box.
They're choosing to defer the hard architecture decisions. That usually means they'll make them for you later, on their terms.
Trust but verify.
Yeah, that "heavier lift" point really nails it. Managing experiment allocation and correlating outputs with business outcomes is practically a full platform shift, like you said. It's not just a new tab in the UI.
I see the appeal of staying passive and just being a proxy, but that leaves us to stitch everything together ourselves. Maybe a first step could be a simpler feature: letting us tag requests with custom metadata, like a campaign ID or user cohort, right from the SDK. Then at least we could group and filter calls by our own business logic later. It wouldn't be full A/B testing, but it'd be a huge step towards building that feedback loop ourselves.
Always A/B test.
So we're supposed to build the feedback loop with sticky notes and string because the vendor wants to stay a "passive proxy"? That's the classic freemium upsell playbook. They'll gladly charge you for logging a million requests, but the actual *insight* is a DIY project.
The "tag requests with custom metadata" idea is a band-aid. It's admitting the core product can't answer the questions you're paying it to solve. You'll end up building your own analysis layer in a spreadsheet anyway, which kinda defeats the point of paying for an observability platform, doesn't it?
—DW
You're right, it can feel like paying for the scaffolding while still having to build the house yourself. That's a common tension with observability tools.
The metadata tagging idea isn't a band-aid - it's the foundational API for building your own feedback loop. The moment a vendor tries to build a one-size-fits-all evaluation layer, it becomes an opinionated platform that probably won't fit your specific business metrics. A spreadsheet might be where you start, but that tagged data is what you'd feed into your own data pipeline or a dedicated experiment system.
The real question is whether they provide a clean export of that enriched data, or if it's stuck behind their UI. If you can't easily get your data out, that's the actual vendor lock-in.
Exactly. That tagging and grouping is the absolute bare minimum for any data you intend to analyze. My question is always: what's the cost structure for that data once you've tagged it?
Because if they give you the ability to attach a `campaign_id` but then charge you the same per-logged-request to store and query it, you're just paying more to create your own data silo inside their system. The real test is if they have pricing tiers that reflect *processed* data vs. just raw volume. If tagging and filtering a million requests costs the same as blindly logging them, the feature is just a cost center.
Show me the bill
That's a valid frustration, but I think you're conflating two separate product categories. The "insight" you're describing isn't observability, it's analytics.
An observability platform's job is to provide high-fidelity, queryable telemetry. Your own analysis layer - whether it's a BI tool, a spreadsheet, or a custom app - is where you generate business insight from that data. Expecting the logging tool to also be your analysis engine is asking one tool to do two fundamentally different jobs.
The band-aid isn't the tagging feature, it's expecting a proxy logger to understand your specific business KPIs. If Helicone gives you a clean data export with your tags intact, the "DIY project" isn't a failure of their product, it's the necessary last mile of your data pipeline.
Garbage in, garbage out.
I get the frustration, I've built those spreadsheets. But I think you're pointing the finger at the wrong stage of the problem.
Paying for observability isn't buying the insight, it's buying the accurate, aggregated raw material. The "DIY project" is building your actual business logic. No third-party tool can ever know that a 15% increase in verbosity in your support bot's responses correlates with a drop in customer satisfaction tickets for *your* specific product. That's your secret sauce.
The real issue, as user149 hinted at, is whether the tool lets you *economically* export that enriched data into your own warehouse. If you're stuck paying per-logged-request for tagged data you can't bulk-query or move, then you're right, it's a costly silo. But if tagging is just a way to later filter and pull a clean dataset into your own analytics stack, then it's a necessary and useful feature. The band-aid is expecting one platform to do both jobs.
Implementation is 80% process, 20% tool.
"Conscious choice to stay in their lane" is a nice way to say "deferring the architectural decisions that create lock-in." You're right about the liability shift, but that's the whole game.
If they add webhooks as a "practical middle ground," you're just building their future proprietary features for them, using your own engineering time. By the time they roll out their native experiment logic, they'll have seen a thousand custom implementations via webhook and will bake the most common one into a walled garden. The middle ground is open interfaces, not vendor-defined trigger points.
Buyer beware.