Having just finished another quarter staring at a cloud bill that could fund a small moon landing, I'm forced to revisit the perpetual vendor question. We're currently on Datadog, but the finance team's increasingly pained expressions have me looking at New Relic again.
The sales pitch for both is identical: "One pane of glass for all your observability needs." The reality, as anyone who's implemented either knows, is a labyrinth of SKUs, ingest-based pricing, and features that sound essential until you realize they're a separate contract. My primary metrics are cost-per-APM-host-equivalent and cost-per-GB-of-log-ingest, but the devil is in the details like custom metric volume, synthetics, and whether your "AI-powered insights" are just expensive noise.
I work in fintech, so we have a sprawling Kubernetes deployment with a service mesh, and our compliance requirements mean logs are non-negotiable. I'm looking for concrete examples of how people have structured their contracts. For instance, have you successfully negotiated caps on metric ingest? Do you find New Relic's more granular pricing (e.g., separate pricing for APM, infrastructure, logs) actually lets you optimize better than Datadog's bundled model, or does it just become death by a thousand paper cuts?
I'll contribute our own hard-learned lessons, like this snippet of a Prometheus recording rule we had to implement just to avoid a fortune in custom metrics for a simple SLO:
```yaml
groups:
- name: slo_aggregation
rules:
- record: job:slo_errors:rate5m
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
```
Posts that just say "we switched and saved 20%" are useless without the gritty details of what they stopped monitoring or how their telemetry pipeline changed. I'm hoping for war stories and architectural pivots that moved the needle.
—matt
No SLA, no problem.
I run infra for a 200-person fintech, managing a k8s cluster with 500+ nodes, Istio, and ~300 services. We've been on Datadog for 3 years and just finished a 6-month POC with New Relic.
1. **Cost Predictability**: New Relic wins. Their data cap overages are a fixed % surcharge (usually 10-25%), while Datadog's overages are uncapped and billed at the same high per-GB rate. We got New Relic to agree to hard caps with zero ingestion beyond limit. With Datadog, our log bill spiked 40% twice due to a misconfigured deployment; no recourse.
2. **APM Host Equivalent Cost**: Datadog is ~$31/host/month for full APM. New Relic's "Pro" tier is ~$25, but their "Standard" tier at ~$15 lacks distributed tracing depth. For our service mesh, we needed Pro. Datadog's container-level granularity is more precise but also 2-3x more expensive in dense k8s.
3. **Log Ingestion & Routing**: New Relic is cheaper per GB ($0.25 vs Datadog's $0.50 on our volume), but you pay separately for querying. Datadog's included query scan limit (5GB/day on our plan) was always exceeded. We spent ~$12k/month on logs in Datadog. New Relic estimate was ~$8k, but required us to forward only ERROR+ to them and keep debug locally, adding engineering overhead.
4. **Custom Metrics**: This is the trap. Datadog charges $0.05 per custom metric per host after the first 100. Our Prometheus migration created 10k metrics/host, which would have been a $15k/month adder. New Relic's "Metrics" SKU has a flat 1M DPM (data points per minute) bucket for $25/month. We'd need two buckets, but it's predictable.
My pick is New Relic, but only if you have the team to manage log filtering and can accept their less polished k8s operator. If your primary need is APM for a complex service mesh and you have no bandwidth for log pipeline work, Datadog's integration is turn-key, just wildly expensive.
Tell me your current custom metric volume per host and your approximate log ingest GB/day, I'll tell you which bill will give you heartburn.
latency kills
Ah, the finance team's pained expressions. I know them well.
You're focused on SKUs and ingest caps. Good. That's where the game is played. But you're asking about structuring contracts with these vendors. My experience? They'll give you a cap, sure. Then you'll hit it in week two because a new pod spec goes sideways and starts emitting DEBUG logs to stdout. The "cap" just means they throttle you, and your alerts go dark.
> sprawling Kubernetes deployment with a service mesh
There's your real cost. Both tools will see every sidecar as a separate entity. Your billable host count isn't nodes, it's pods * containers. You think you're buying observability, but you're really just funding their yacht party.
Forget granular pricing. It's a trap. It lets you "optimize" yourself into a corner where you can't actually trace a request without buying three add-ons. Pick the one that doesn't charge you punitive overages and then build your own cardinality limits into the agent config. It's the only control you actually have.
If it ain't broke, don't 'upgrade' it.