I'm evaluating a complete overhaul of our observability and analytics stack for a B2B SaaS platform currently migrating from a monolithic EC2 setup to Kubernetes on EKS. The current "stack" is a Frankenstein of CloudWatch logs, a legacy New Relic APM contract that's bleeding money, and a self-hosted Grafana dashboards that nobody trusts because the data sources are inconsistent. It's a mess, and the migration is the forcing function to fix it.
The core requirements are non-negotiable:
* **Unified view for engineering:** Devs and SREs need a single pane for traces, metrics, and logs with solid correlation. The current context-switching between tools kills incident response time.
* **Cost predictability:** The New Relic bill is a variable nightmare. We need a pricing model that scales linearly with our Kubernetes pod count, not with every single span and log line.
* **Strong Kubernetes-native instrumentation:** Automatic discovery of services, pods, and the ability to drill down to container-level metrics without manual tagging hell.
* **Support for business analytics:** Product and sales teams need dashboards on user adoption and feature usage. Currently, this is a separate, duct-taped BigQuery setup. Ideally, the same event pipeline could feed both engineering observability and product analytics, but with different consumption models.
I've done the PoCs. I know the pain is coming.
* **Elastic (ELK) Stack:** The control is alluring, but the operational overhead of managing the stack on k8s (Elasticsearch, Logstash, Kibana, plus APM agents) is a full-time job. I've seen clusters blow up from a bad query. The "managed" cloud offering quickly becomes expensive for high-volume applications.
* **Grafana Stack (LGTM):** Grafana Cloud with Loki for logs, Tempo for traces, and Mimir for metrics. The promise of a unified query language (LogQL, PromQL) is strong, and the vendor lock-in is less severe. My concern is maturity, especially around Tempo's trace analytics and scaling for high cardinality data.
* **Datadog:** The gold standard for "it just works" and the UI is superior. The sales pitch is easy. The bill is not. Every feature is another SKU. Once you're in, the data gravity makes it painful to leave. Their Kubernetes operator is excellent, however.
* **Honeycomb:** Fascinating for event-driven debugging and the high-cardinality analysis is a game-changer for complex microservices. Less traditional for dashboarding and long-term trend analysis, so it might force a "two-tool" reality alongside a metrics-focused system.
What I'm hoping to find here are real-world architectural patterns, not vendor slides. Specifically:
* How are you structuring your ingestion pipelines to keep costs sane? Are you doing pre-aggregation or sampling before data leaves your cluster?
* For those who've unified business and operational analytics, what was the breaking point? Did you use a pipeline like OpenTelemetry Collector to fan out to different backends?
* Concrete Terraform modules or Helm charts for deploying the OpenTelemetry Collector in a production EKS environment, with reliable buffering and failover.
I'll contribute our eventual implementation, including the inevitable terraform state migraines and the custom resource definitions we had to write to make it all work in the GitOps flow. The goal is a stack that engineers actually use, not one that looks good in a architecture diagram.
---
Been there, migrated that
You're hitting on a critical pain point with the unified view requirement. Many teams interpret this as needing a single vendor, but that often leads back to the New Relic cost trap. The architectural goal should be a unified query layer, not a unified ingestion pipeline.
For your migration context, consider separating the pipeline. Use OpenTelemetry for collection, which gives you vendor-agnostic instrumentation for traces, metrics, and logs. You can then route telemetry data to different backends based on use case and cost profile. Engineering gets correlated data in Grafana (with Tempo for traces, Loki for logs, and Mimir for metrics), while business analytics can be served from a cost-optimized columnar database like ClickHouse, queried from the same Grafana frontend.
This approach directly addresses your cost predictability demand. You control the retention periods and sampling rates for each backend independently. Kubernetes-native instrumentation is a strength of the OTel collector, especially with the k8sattributes processor. The risk is operational overhead, but the Helm charts for the Grafana stack are production-ready.
— Harper
You've got the right priorities, but that linear scaling ask with Kubernetes pod count is a fantasy. Every vendor promising that has small print. They'll scale linearly with pods until you need high cardinality tags, then the bill spikes. I've seen it with Datadog's "container-based pricing" and New Relic's "consumption model." The data ingestion fees always creep back in.
Separating the pipelines is smart, but don't just swap one vendor lock-in for another. If you go full Grafana stack (Loki/Tempo/Mimir), you're just self-hosting your variable cost nightmare. Your SRE team now owns the scaling, storage, and uptime of that. The bill might be to AWS instead of New Relic, but it's still a bill.
For business analytics, keep it completely separate from your engineering observability. Pipe the clean, aggregated events to a cloud data warehouse. Trying to make ClickHouse serve real-time container metrics and year-long sales trends is how you end up with another "nobody trusts" dashboard.
-- cost first
This is exactly the kind of situation I'm trying to understand better. The idea of a unified *query* layer versus a unified ingestion pipeline from user1481 is really clicking for me, but your point about scaling with pod count is huge.
When you say you need business analytics dashboards on user adoption, is that data coming from the same app logs and events that your engineering observability uses? I'm trying to wrap my head around if you'd have two separate OpenTelemetry collection flows, or if you'd split a single stream later. It seems like mixing them could get expensive fast, but maybe that's the whole point of separating the pipelines?
Yeah, the separate pipeline strategy is the only way to get both engineering observability and business analytics without a cost explosion.
You asked about collection flows: I'd instrument *everything* with OpenTelemetry from the start, but have it output to two distinct collectors. One collector filters for high-cardinality, high-volume SRE data (traces, error logs, infra metrics) and sends it to your cost-optimized engineering backend. The second collector picks up specific, low-volume business events (feature_used, user_upgraded) and routes them directly to a separate, cheap analytics datastore like Postgres or BigQuery. Trying to split one massive stream later is a filtering nightmare.
This way, your product team's "dashboard on user adoption" doesn't get charged at your observability vendor's premium log-ingestion rates.
Beta tester at heart
Separating pipelines sounds clever until you're debugging a data discrepancy and realize your "cost-optimized" ClickHouse backend for business analytics is missing half the context from the engineering stack. The unified query layer promise falls apart when the underlying data is fundamentally siloed by ingestion path.
And those "production-ready" Helm charts for the Grafana stack? They're a starting point, not a solution. You're still on the hook for scaling Mimir's blocks storage and keeping Loki's queriers from melting down when someone runs a wildcard search over 30 days of logs. That's just swapping a variable vendor bill for a variable cloud infrastructure bill plus your team's unpaid overtime.
null
You're not wrong about the hidden costs in self-managed stacks, but the discrepancy issue hits the nail on the head. A unified query layer over split data is a technical debt time bomb.
I've seen teams burn weeks reconciling metrics because the business event pipeline sampled data or applied different aggregation rules. If your SLO dashboard and your boardroom report use different backends, they will disagree. Period.
The real fix is a contract-first data definition in your OpenTelemetry schema, enforced before the split. It's procurement 101: define the deliverable before you sign with a vendor, or in this case, route to a backend. Without that, you're just building a fancier silo.
—hd
The Kubernetes-native instrumentation requirement is key. I've heard some managed services can handle pod discovery automatically, but what about services that are scaled to zero? Does that break the correlation you need for the unified view?
Still learning.
Scaling to zero is a great point, and it absolutely can break correlation if your tracing and metrics collection aren't designed for it. The pod-level tags just vanish when the pod stops existing.
Most managed services that do auto-instrumentation rely on a sidecar or daemonset agent. When the pod scales to zero, that local agent instance disappears too, so any in-memory buffering or tail-based sampling for traces gets cut off. You can lose the end of a request, which kills your error rates and P99 latency views.
The workaround is pushing telemetry outside the pod immediately, like to an OpenTelemetry Collector sidecar that forwards before shutdown, or better yet, having your app SDK send spans and metrics directly to a central collector service. That way, the data flow is independent of the pod lifecycle. It adds a bit more network configuration, but it keeps your view intact.
catdad
Oh man, I feel your pain on the "Frankenstein" stack, we went through something similar a couple years back. Everyone's making great points about splitting pipelines and query layers, but I need to ask about your last line.
You mentioned product and sales need dashboards on user adoption and feature usage, and that it's currently separate. If you're building a unified engineering view, that's the perfect time to lock down your event taxonomy so the business side isn't flying blind later.
What's your plan for defining those user events? Are you thinking of tagging them at the source with OpenTelemetry, or using a separate tracking library like PostHog or Amplitude? Because if you don't standardize that now, you'll just create a third, even messier silo for product analytics.
Absolutely agree on locking down the event taxonomy upfront. We tried using a separate tracking library initially, and the mismatch in event definitions between that and our internal logs was a nightmare for our product team.
The real trick is to use OpenTelemetry semantic conventions for those business events from day one. Treat "feature_used" or "trial_upgraded" like you would any other telemetry signal. That way, you can enforce the contract at the collector and route it correctly - high-volume app events stay on the engineering pipeline, while the standardized business events get filtered to your analytics datastore.
If you bolt on PostHog later with its own schema, you're just adding another translation layer that will drift. Start with one source of truth.
Scaling linearly with pod count is a pipe dream. Every pod generates variable load. Your bill will still jump when you're debugging an outage and tracing volume spikes 100x.
Forget unified engineering and business analytics in one system. They have opposite requirements. High-cardinality debugging data will wreck your analytics costs, and aggregated business metrics are useless for root cause.
Pick one managed vendor for engineering observability that does correlation well. Use their Kubernetes operator. For business analytics, instrument events with a simple library and shove them into a cheap column store. Trying to architect a perfect hybrid is how you end up with another Frankenstein.
Simplicity is the ultimate sophistication
Your last point about business analytics support is the key unlock here. Everyone focuses on the engineering side during a migration, but if product and sales dashboards stay separate, you're just recreating the silo problem at a higher cost.
I'd push for defining your core business events *as OpenTelemetry signals* right now, during instrumentation. Treat "user_upgraded" with the same semantic rigor as a span. Your collectors can then filter and route these low-volume, high-value events to a cheap columnar store (we use Snowflake for this) while the high-volume engineering data goes to your observability backend.
This way, you get a single source of truth from the start. The sales team's "feature adoption" dashboard and the SRE's "error rate by pod" dashboard are querying data that originated from the same instrumentation contract. It stops the "why are the numbers different?" fights before they start.
Let the machines do the grunt work
You've highlighted the critical tension. A unified engineering pane and business analytics support are often at odds on pricing models. A vendor promising both with linear pod-based scaling is likely using marketing language that falls apart in practice.
Even a pure engineering-focused vendor will struggle with true linear scaling. Your bill is often a function of retention and sampling. If you need 30 days of high-cardinality trace data for debugging, the cost will spike with traffic, not just pod count. The key is finding a vendor where you can control that variable cost knob yourself through aggressive sampling and tiered retention, without breaking correlation.
For the business analytics requirement, I strongly advise against trying to force it into the same pricing bucket. Define the events with OTel, but route them to a separate, fixed-cost system. That keeps your engineering observability costs predictable and prevents a sales team's exploratory query from spiking your SRE bill.
Buy once, cry once.
100% on the pricing models being the hidden trap. That "linear pod-based scaling" promise is the part that always unravels. Your point about retention and sampling being the real cost drivers is spot on.
We got burned because our "unified" vendor's pricing looked great for engineering data at first. But the contract had a huge, opaque multiplier for retaining high-cardinality data beyond 7 days. When we needed to keep 30 days of trace data for a compliance audit, the bill quadrupled. It wasn't about pods at all.
Separating the pipelines isn't just about performance, it's financial sanity. Locking business events to a fixed-cost column store from the start protects the engineering budget. One wild product query shouldn't risk your SLO monitoring.