You're right about the utility bill analogy, but I'd push on the idea that it's non-negotiable. The predictable bill is only predictable because someone made the engineering investment to build a stable, monitored pipeline for those core metrics.
The false economy isn't in moving the metric itself. It's in assuming that moving it to logs means you can skip that same engineering investment. If you treat a log-based metric with the same rigor as a first-class metric - with its own monitoring, documentation, and change management - the cost profile changes. The problem is most teams don't, which is how you end up with the unmonitored junk you described.
Always check the data transfer costs.
You're right to be worried about shifting the cost. We ran this playbook a year ago, and your tool mix is the key factor. Since you're using Datadog for metrics and Elastic for logs, you're comparing two different pricing models where high-cardinality data can become very expensive, very quickly in Elastic if you're indexing everything.
The rule of thumb we landed on was based on two filters: first, if a metric powered any alert with an SLA under 10 minutes, it stayed a metric. Second, we only moved data where we could apply aggressive log sampling (like 10% sample rate) without losing critical trends. This stopped the log volume from exploding.
Our actual savings were modest, about 15% on the metrics side, but that only held because we spent significant engineering time tuning the log ingestion and aggregation. The real benefit wasn't the direct savings, it was identifying and deleting a whole category of metrics that nobody actually queried anymore.
catdad