We're a small ops team (10 people) responsible for about 50 microservices across 5 dev teams. We've outgrown our basic logging setup and need real monitoring.
The consensus online seems to be: "Just use Prometheus, it's the standard." But then everyone also says "Use Datadog, it's what the pros use." 🤔
We ran a 3-month trial of Datadog. The out-of-the-box Kubernetes dashboards were amazing at first – everything just worked. But the cost exploded once we enabled tracing and custom metrics for a few key apps. The vendor said we could "manage our data volume," but that just felt like turning off the very features we needed.
So we spent a month setting up Prometheus + Grafana + Alertmanager ourselves. It's powerful and the customizability is real. But the maintenance overhead for our team feels high. Keeping up with storage management, scaling, and making dashboards that everyone can use has been a time sink.
My honest take so far:
* Datadog: Feels like buying a fully-loaded car. Great until you see the monthly bill and realize you're paying for features you barely use.
* Prometheus: Feels like building a car from parts. You get exactly what you want, but you're suddenly a part-time mechanic.
For a team our size, is there a true middle ground? Did anyone else choose one path and regret not taking the other? I'm skeptical of the "easy button" sales pitch now, but also wary of the hidden time cost of self-managed tools.
I'm a senior platform engineer at a fintech company with around 200 microservices; we transitioned from a pure Prometheus stack to a hybrid model about 18 months ago, and I've maintained both.
Here's a breakdown based on what your team is describing.
1. **Team Bandwidth & Skill Drain:** For a 10-person team, the hidden cost of Prometheus is ongoing education and tooling. You'll spend time not just on storage (e.g., Thanos or Cortex), but on making Grafana dashboards idiot-proof for devs and writing maintainable alerting rules. With Datadog, that cost is paid in dollars, not sprint cycles. At my last shop, a team of our size spent roughly 15% of one FTE's time just keeping the monitoring stack itself operational.
2. **Real, Predictable Cost:** Datadog's cost is its custom metrics, spans, and logs. If you instrument liberally, your bill scales directly with your engineering output. I've seen bills jump from $3k to $11k/month by enabling APM for a dozen high-traffic services. Prometheus costs are your infra and labor. The raw software is "free," but you need to provision and manage the storage (cloud object storage is cheap) and the compute for queries.
3. **Integration Friction:** Datadog's out-of-the-box Kubernetes view, service map, and log correlation work after you deploy the agent. For Prometheus, achieving a similar "single pane" requires stitching together kube-state-metrics, node exporters, and often a separate logging pipeline (like Loki) yourself. The initial setup is a multi-week project to get to parity.
4. **Where It Breaks:** Datadog's query language (and its limits on metric cardinality) can feel restrictive when you need complex PromQL queries for derived metrics. Prometheus breaks when you need long-term storage (beyond 15 days is painful native) or truly global query views across multiple clusters without significant added complexity.
Given your description of being overwhelmed by the maintenance, I'd recommend a third path: use the Prometheus Operator for collection and running alerts, but ship your metrics to Grafana Cloud for the managed dashboard and long-term storage layer. You get the standardization of Prometheus without managing Grafana or Alertmanager, and the cost is more predictable than Datadog. If you must pick between the two, stick with your Prometheus stack but immediately invest in Terraform for managing Grafana dashboards and Alertmanager configs to reduce toil.
throughput first
That "15% of one FTE" cost you cite for Prometheus is almost certainly a low estimate once you factor in the context switching and the constant churn in the ecosystem. A new version breaks a plugin, a deprecation forces a rewrite of half your alerts, and suddenly your "free" software burned two weeks.
But calling Datadog's cost "predictable" is a stretch. It's only predictable if your feature usage is frozen. The moment a dev adds a new tag or you decide to debug that one noisy endpoint, you're gambling with the meter. Their pricing model isn't a bill, it's a negotiation where you have to constantly police your own curiosity.
—EB
You've pinpointed the exact operational friction. That hidden tax on developer time and attention, the context switching and ecosystem churn, is often misrepresented as a simple "build vs. buy" calculation.
Where I diverge slightly is on the predictability spectrum. Both models have unpredictable costs, just in different currencies. With the OSS stack, the unpredictability is in time - a critical P1 incident occurring the same week you need to migrate your Prometheus storage layer. With Datadog, the unpredictability is in dollars, but it's also a direct function of engineering activity - a new feature rollout or a forensic investigation automatically increases the bill. That creates a perverse incentive to not investigate issues thoroughly.
The core problem is that both pricing models, time or dollars, scale directly with operational complexity, which is exactly when you need the tooling to be most reliable and cost-effective.
That's a really interesting way to put it, the "different currencies" thing. It makes our experience with Datadog suddenly click.
I've seen that perverse incentive you mentioned firsthand. On our trial, we had a weird latency spike and I was hesitant to even *look* at the distributed trace view because I knew it would spike our bill. It felt wrong, like I was being punished for trying to do my job properly.
So if both models have unpredictable costs tied to complexity, is the real choice just picking which kind of pain your team is better equipped to handle? Like, do you have more spare time or more spare budget? 😅
That car analogy hits a bit too close to home. But you're missing the third option: the used car with a solid warranty.
You don't have to build from parts or lease the luxury model. Look at the managed Prometheus services from most cloud providers. They take the scaling and storage headaches off your plate for a predictable hourly fee.
You'll still need to manage dashboards and alerts, but you're trading pure cash (Datadog) or pure time (self-hosted) for a hybrid cost. The maintenance tax gets cut down, but you're not paying for the gold-plated steering wheel.
always ask for a multi-year discount