Skip to content
Notifications
Clear all

Gauge competitors - pros and cons from production users

2 Posts
2 Users
0 Reactions
2 Views
(@devops_shift_lead)
Estimable Member
Joined: 4 months ago
Posts: 136
Topic starter   [#19921]

We've been running a three-month bake-off between Datadog, New Relic, and Grafana Cloud for observability on our 150-node K8s fleet. Budget's tight, but we can't afford blind spots. The marketing pages all look the same. I need the war stories from engineers who've scaled these in anger.

Here's our shortlist and the specific trade-offs we're tracking:

**Grafana Cloud (Loki/Tempo/Mimir)**
* Pro: Unified querying across traces/logs/metrics is real. Cost predictability with BYO cloud object storage for logs is a winner.
* Con: Agent management (Grafana Agent vs. Otel) feels heavier. The "glue it yourself" factor for advanced features burns cycles. Alerting has bitten us with silent failures.
* Verdict so far: Leading for cost control, but operator overhead is non-zero.

**Datadog**
* Pro: Out-of-the-box everything. K8s pod view → logs → traces is seamless. APM instrumentation just works. Their synthetics are robust.
* Con: The bill is a black box that grows with your log volume. You feel locked in. Their agent is a resource hog on our edge nodes.
* Verdict so far: Best "just works" experience, but CFO is already asking for forecasts.

**New Relic**
* Pro: Pricing model (data plus users) is simpler to forecast than pure volume. Their new AIOps stuff actually flagged a weird Cassandra compaction issue for us.
* Con: UI feels slower at scale. The query language isn't as powerful as PromQL for deep dives. Felt like we had to work around it for custom K8s metrics.
* Verdict so far: Strong contender if your primary pain is alert noise reduction.

The question isn't which is "best." It's which trade-offs you can live with at 3 AM. I'm particularly interested in:
* Actual total cost at ~10 TB of log ingest/month.
* Agent stability and resource footprint in production at scale.
* How you handled the transition from one platform to another.

Post your own stack and pain points. Benchmarks and pipeline logs beat opinions.

-shift


shift left or go home


   
Quote
(@elliotn)
Estimable Member
Joined: 1 week ago
Posts: 106
 

You've nailed the Datadog experience. The lock-in is less technical and more operational; their agent's deep integration means you're not just shipping data, you're adopting their entire collection taxonomy. That bill shock is predictable if you instrument everything, but their log management pricing is particularly punitive at scale.

On New Relic's pricing model, it's a double-edged sword. The predictable per-host cost is great for budgeting, but their definition of a "host" gets fuzzy with containers and serverless. We saw a 30% cost overrun because New Relic counted our Kubernetes pods under certain memory thresholds as separate "hosts" for billing purposes. Their newer consumption model might help, but you need to audit your own usage telemetry closely.

Your point about Grafana Cloud's "glue it yourself" factor is critical. The silent alert failures often trace back to the split between Mimir, Loki, and Alertmanager configurations. You think you've set a rule in Grafana, but it's actually dependent on a correctly configured and loaded Alertmanager sidecar with the right template files. The overhead isn't just in cycles, it's in institutional knowledge - you now own the pipeline's resilience, not the vendor.

Have you quantified the operator overhead for Grafana Cloud? We tracked 15-20 hours a month in maintenance for a fleet half your size, mostly around agent updates and storage tier tuning. That labor cost can erase the savings over Datadog's premium.


Data first, decisions later.


   
ReplyQuote