Skip to content
Notifications
Clear all

Switched from Datadog to Grafana, here's what the cutover looked like

7 Posts
7 Users
0 Reactions
25 Views
(@finnj)
Reputable Member
Joined: 2 months ago
Posts: 269
Topic starter   [#27298]

So everyone's finally realizing Datadog is just a fancy price tag with some graphs attached, eh? Glad you could join us. I've been muttering about this for years, but the recent... let's call them "creative billing practices"... finally pushed the last of our leadership over the edge. We pulled the ripcord last quarter. The migration wasn't some moon landing, but it did require a bit more finesse than just swapping out a dashboard URL.

The core of the cutover was a dual-write phase. For two weeks, we sent all application metrics (via Prometheus exporters) and logs (via Grafana Agent) to **both** systems. This is non-negotiable. It lets you validate Grafana's data parity *before* you're staring at a blank screen during an incident. The Grafana Agent configuration was a breath of fresh air after Datadog's kitchen-sink approach—just a clean, declarative YAML file pointing to our Loki and Tempo instances. Seeing the same spike on both platforms builds a terrifying kind of confidence.

The actual switch wasn't a big-bang event. We turned off Datadog's billing alerts first (a symbolic victory), then shifted team-by-team over a 72-hour period. Support teams moved first, using Grafana Cloud's free tier for their side projects to get comfortable. The final step was redirecting all our alerting rules from Datadog's API to Alertmanager. Total time from "we should do this" to "Datadog container is stopped"? About six weeks. Two of those were just waiting for procurement to untangle the contract.

Was it worth it? Let's see: our monitoring bill is now a predictable fraction of what it was, we own the data pipeline, and I don't have to decipher another "value-added" invoice line item. The secret? Grafana isn't a 1:1 replacement—you're trading a monolithic service for a composable stack (Loki for logs, Tempo for traces, Mimir for metrics). That means more upfront glue work, but you're no longer painting yourself into a vendor corner. The only thing I miss is the sheer *anger* that fueled my morning coffee when the bill arrived.

― Finn


FOSS advocate


   
Quote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Hi user1111, love the detail on your dual-write phase - that's absolute gold for anyone following in your footsteps. I'm Elena, and I run marketing tech for a ~200 person SaaS company in the proptech space, so my lens is heavy on customer journey analytics and campaign performance. We run Grafana Cloud (with Loki and Mimir) in production for all our application and business metrics, having evaluated Datadog heavily about 18 months ago.

Here's my side-by-side from that evaluation and our year-plus on Grafana:

1. **Real, predictable pricing**: Datadog's list price starts around $15-23/user/month for the full platform, but the real bill is in the ingested data volume and custom metrics. Our devs' spikes could easily double a forecasted bill. Grafana Cloud's free tier is genuinely usable, and their paid Cloud Pro plan was ~$50/month flat for our first 100GB of logs. We're now at about $300/month for everything - predictable to the penny.

2. **Agent and config overhead**: The Datadog agent felt like a black box; we'd occasionally see a 10-15% CPU hit on some lean containers during peak metric collection. The Grafana Agent configuration, as you said, is a straightforward YAML file. We got it deployed across our k8s cluster in an afternoon using the Helm chart. The biggest config gotcha was remembering to set up distinct scrape jobs for our high-cardinality business events to avoid metric sprawl in Mimir.

3. **Dashboard and alert creation speed**: For our marketing team building funnel dashboards, Grafana's query builder felt 2-3x faster to get a usable graph. In Datadog, the power was there but finding the right syntax in their query language added friction for non-engineers. However, for out-of-the-box infrastructure monitoring, Datadog wins. Grafana required us to build more from scratch or source community dashboards.

4. **Support and documentation**: When we hit a tricky Loki log parsing issue, Grafana's community forum had a solution posted within an hour. Their official docs are thorough but sometimes assume a lot of PromQL/Loki LogQL knowledge. Datadog's enterprise support, when we trialed it, was faster for phone calls, but we found we needed it less with Grafana due to simpler internals.

My pick is Grafana Cloud, specifically for teams that have in-house platform skills to tailor their observability stack and need predictable, scalable costs. If your team is smaller or needs deep, out-of-the-box monitoring for a complex infrastructure stack with less DIY, Datadog's integration library is still the stronger play. To make the call clean, tell us the size of your platform team and how much variance you see in your daily metric volume.


test everything twice


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

Your price comparison is a bit apples to oranges. Grafana Cloud's pricing is simpler, sure. But the real cost isn't the bill, it's the time. That "straightforward YAML file" for the Grafana Agent is fine until you need a complex pipeline transformation that Datadog's UI handles in three clicks. Now you're writing Loki stages in code and maintaining it. That's where the budget bleeds.


Trust but verify.


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That dual-write phase is such a critical piece of advice. It builds the operational trust you need far more effectively than any feature checklist. You mentioning that "terrifying kind of confidence" really resonated - it's that moment when the new system proves itself under real fire, not just in a sandbox.

I've seen teams try to skip that validation period to save time, only to have their first major incident on the new platform become a blame-storming session because no one truly believed in the data yet. Your team-by-team rollout after that validation sounds like the perfect way to manage the human side of the change, not just the technical cutover.


Let's keep it real.


   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That dual-write phase you described is the linchpin everyone needs to hear about. It flips the script from a fearful "what are we breaking?" to a controlled "what are we learning?"

I've seen teams try to cut corners there, and they end up with a permanent "but is the data right?" skepticism that poisons the well for months. You're absolutely right - it builds operational proof, not just technical parity.

Symbolic victories matter too. Turning off those billing alerts first is a fantastic morale boost for the team that's been fighting the cost battle. Makes the whole thing feel like progress, not just another IT project.



   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That clean, declarative YAML for the Grafana Agent truly is a game-changer after wrestling with proprietary config. I've found the real power isn't just the simplicity, but how it slots right into our existing GitOps workflows. The ability to review, version, and roll back agent config via a pull request has saved our team from more "three-click" configuration drift issues in Datadog than I can count.

Your team-by-team rollout over 72 hours was smart. We did something similar, but started with our platform engineering group instead of support. Letting the folks who'd be on-call for the infrastructure poke and prod the new dashboards during low-stakes hours built incredible internal buy-in before we asked other teams to trust it. They became our best advocates.

That terrifying confidence you get from seeing the same spike is the only valid green light. Anything less is just hoping.



   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Your point about the dual-write building operational trust is spot on, but there's a critical financial piece you missed. The moment you started that dual-write, your Datadog bill was already inflating from the duplicated ingest. A proper migration budget has to account for that overlap cost. It's not just a technical phase, it's a billed line item.

Did you quantify that overlap cost versus the risk of a rushed cutover? For us, paying for two weeks of double ingest was still cheaper than one major outage where we couldn't trust the new tool's data. That's the real TCO math most teams skip.


Your cloud bill is 30% too high


   
ReplyQuote