Skip to content
Notifications
Clear all

We switched from Datadog to Grafana - here's the cost comparison

1 Posts
1 Users
0 Reactions
2 Views
(@crm_hopper_2026)
Reputable Member
Joined: 3 months ago
Posts: 164
Topic starter   [#5629]

Our organization recently concluded a six-month evaluation and migration project, moving our core infrastructure and application monitoring from Datadog to a self-managed Grafana stack (utilizing Grafana for visualization and alerting, with Prometheus and Loki for metrics and logs). The primary driver was cost predictability at scale, though we also sought greater control over our data pipeline.

I conducted a structured, phased comparison. The methodology was as follows:

* **Baseline Period:** We ran both systems in parallel for 90 days, ingesting identical telemetry data.
* **Cost Tracking:** We itemized all Datadog invoices (hosts, custom metrics, ingested log GB, APM hosts, etc.) and mapped them to equivalent infrastructure costs for the Grafana stack.
* **Infrastructure Allocation:** For Grafana, we calculated the fully loaded cost of the required EC2 instances for Prometheus/Loki/Grafana, associated EBS storage volumes, and the operational overhead for our platform team.

**Quantitative Cost Findings:**

Our annualized run-rate for Datadog was approximately $347,000. The equivalent Grafana stack costs break down as:

* **Compute & Storage (AWS):** $28,500 annually for six i3en.2xlarge instances (for Prometheus/Loki) and two c6a.4xlarge instances (for Grafana).
* **Managed Grafana (AWS):** We opted for AWS Managed Grafana for the visualization layer at $215/user/month, totaling $5,160 annually for our four core users.
* **Labor Overhead:** The largest variable. We allocated 20% of one senior platform engineer's time for ongoing maintenance, patch management, and scaling. At our fully loaded rate, this adds approximately $45,000 annually.

**Total Annual Grafana Stack Cost:** ~$78,660.

This represents a **77% reduction** in direct monetary outlay. However, this is not a pure savings; it is a transfer of cost from a vendor invoice to internal infrastructure and labor. The trade-off is clear: we exchanged capital for control and predictability. Our Datadog costs were highly volatile, spiking with log volumes or new custom metrics. Our Grafana infrastructure costs are now essentially fixed, scaling only with raw infrastructure growth, not usage.

**Migration Walkthrough Summary:**

* **Data-Migration Approach:** We did not perform a historical data migration. We defined a cutover date and began fresh in Grafana. Critical historical data was archived from Datadog prior to contract termination. We focused on migrating the *definition* of our monitoring—dashboards and alerts—not the data itself.
* **Cutover Plan:** A phased, service-by-service cutover over a two-week period. We redirected metrics from our applications to Prometheus first, followed by log streams to Loki. Dashboard URLs in runbooks and chat ops were updated in batches. Datadog remained active as a fallback during this period.
* **Transition Timeline:** The entire process, from initial PoC to full decommissioning of Datadog agents, took 26 weeks. The parallel run phase (weeks 8-18) was critical for building parity in alert coverage and team confidence. The actual technical cutover was the final 2 weeks.

The decision hinges on your organization's tolerance for operational management. For us, the cost predictability and avoidance of vendor lock-in justified the increased internal operational burden. The ROI was clear, but it is not a simple "lift-and-shift"; it requires a committed platform team to own the new stack.



   
Quote