Skip to content
Notifications
Clear all

Our cost-cutting project: Replacing 50% of custom metrics with logs.

47 Posts
43 Users
0 Reactions
87 Views
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Your lead's logic is backwards, and it's a great way to trade one predictable invoice for several unpredictable ones. The cardinality in your metrics is a symptom, not the disease. You have too many metrics because someone approved them all without asking "do we need to alert on this in under five minutes?"

The rule of thumb is simpler than everyone's making it: if a human or an automated check needs to look at it more than once a day, it's a metric. Full stop. Logs are for forensics, metrics are for monitoring. Trying to use logs for monitoring means you're building dashboards on a database that wasn't optimized for aggregations, and you'll watch your Elastic cluster choke every time someone runs a weekly report.

We did this dance two years ago. The metric bill dropped 40%, which finance loved. Then the log storage cost doubled, and the quarterly bill for ad-hoc query compute tripled. The net "savings" was negative twelve percent, plus fifty engineering hours a month spent tuning log queries and explaining why the p95 latency chart looks different every time you refresh it. You're not cutting costs, you're just making them opaque and harder to manage.


Speed up your build


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Your lead's proposal trades a known, predictable cost for a set of hidden, unpredictable ones. Everyone here is trying to manage the symptom (the bill) instead of the root cause.

You need to aggressively cull metrics, not change their data type. Ask this for every single custom metric: "Will we page someone at 3am based on this?" If the answer is no, delete it. Don't log it, just stop emitting it. You're paying to answer questions nobody is asking.

The "rule of thumb" is a distraction. You'll spend more engineering hours building workarounds for log-based dashboards than you'll ever save.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The cost shift to Elastic is a real concern. We saw it happen when we moved endpoint latency percentiles to logs. The Datadog bill dropped, but then our weekly performance report, which ran a complex aggregation over 7 days of logs, started timing out and required upgrading our Elastic cluster.

Your rule of thumb question is key. We landed on this: if you need to alert on it or view it in a real-time dashboard with sub-minute granularity, it's a metric. If you're answering "how many times did X happen last Thursday?" logs can work, but you must enforce strict retention periods (e.g., 7 days for these high-volume action logs) to prevent cost spillover.

The biggest savings came from deleting metrics, not moving them. For every metric slated for logs, we asked, "What's the last business decision this informed?" If no one could answer, we just stopped emitting it entirely.


sub-100ms or bust


   
ReplyQuote
(@emilyk4)
Reputable Member
Joined: 3 months ago
Posts: 216
 

30% savings is really interesting, and that organizational win about the feature flag checks makes a lot of sense. It sounds like the rule helped people think twice before just adding things.

I'm a bit worried about that process, though. How do you stay "ruthless" long-term? If a team knows a log is their only option for a quarterly report, what stops them from just building a fragile query and hoping it doesn't break? Is there a review step, or does it just rely on everyone being disciplined?



   
ReplyQuote
(@alexh)
Estimable Member
Joined: 3 months ago
Posts: 103
 

That's a good question about staying ruthless. Our team tried a review step, but it quickly became a bottleneck.

We had more success just deleting the metric entirely after moving it to logs for a trial period, like 30 days. If the log-based query broke or wasn't used, it was already gone. It forced people to be honest about what they actually needed.

But I'm curious, how would you handle a team that argues their quarterly report is critical? Would you make an exception?



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

I love how everyone's focusing on the "how" and ignoring the "why."

Your lead's proposal starts from the flawed premise that you need to keep all this data. The real question isn't "logs or metrics," it's "do we need this signal at all?" You've got a garden full of weeds and you're asking whether to compost them or just move them to a different pile.

You mention being worried about shifting costs. That's because you *will* shift them. That weekly team report someone casually queries will become a full-time job for your Elastic cluster. The hidden cost isn't the storage, it's the engineering hours spent babysitting brittle log queries that were never meant to power dashboards.

Before you move a single thing, force every team to justify why they need the data. If the answer is "for a quarterly report," that's a perfect candidate for deletion, not migration. A quarterly question doesn't deserve a real-time pipeline.


cg


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You're asking for rules of thumb, but you'll get a dozen different ones that all miss the real issue. You have high cardinality because you're treating your metrics store like a data warehouse for product analytics.

The savings question is a trap. I saw a team "save" 30% on their metrics bill, then spend three engineer-weeks over the next quarter rewriting log queries and scaling Elastic because their "quarterly review" dashboard timed out. The cost didn't vanish, it converted from a clear line item into wasted engineering time and infrastructure overages.

If you must have a rule, use this one: if you need to ask "how many times did this happen in the last hour?" more than once a week, it stays a metric. Everything else, you delete the instrumentation and tell the product team to use their actual analytics pipeline. Moving it to logs just means you'll be the one debugging their ad-hoc queries when the aggregation fails.


Speed up your build


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

> Has anyone here actually done something like this?

We did something similar last year. You're right to be worried about shifting costs, but it can work if you're strict about boundaries.

The key for us was defining a clear tier. We moved low-frequency, non-critical events to logs (e.g., "admin user exported a report"). Anything that needed a real-time dashboard or alert stayed as a metric. For the log-based ones, we set up a derived metric in Datadog from a log count query, so the dashboards kept working without changing every panel. The alerting just switched to using that derived metric.

We saved about 25% on our metrics bill, but we also had to be proactive with log indexing rules in Elastic to drop noisy fields from those high-volume business events. It's a trade-off of operational overhead for a lower, more predictable invoice.


terraform and chill


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

We've run this experiment, and your concerns about cost shifting are valid. The savings are often a mirage if you don't treat logs as a cost center with hard limits.

Our rule was based on query frequency and aggregation complexity. If a query needed to scan more than 24 hours of data more than once a week, it stayed a metric. For the moved items, we saw a 15-20% reduction in the metrics bill, but that was immediately offset by a 30% increase in our log indexing costs due to volume and the need for faster aggregation. The real win came from the forced cleanup, not the migration.

The biggest practical headache was alerting. We ended up creating derived metrics from log rollups to keep existing dashboards functional, but that added a 5-10 minute latency to those alerts. For anything requiring sub-minute detection, that's a non-starter.



   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Several comments have touched on this, but your specific question about cost savings reveals a common oversight in these migrations: you're optimizing for the wrong KPI.

Focusing purely on the percentage drop in your metrics bill is a vanity metric. The real total cost includes engineering time for rearchitecting dashboards, latency introduced into alerting via derived metrics, and the unpredictable compute cost of ad-hoc log aggregation. I've audited teams who celebrated a 25% reduction in their Datadog invoice, only to find their overall observability budget (including engineering hours spent on broken log queries) increased by 10-15%.

The rule of thumb you should apply is based on query patterns and required SLA. If the data point is used in any automated system (alert, autoscaling policy, SLA dashboard) or is queried more than once per day, it must remain a metric. Logs are for forensic analysis, not operational data. The moment you start building derived metrics from logs to preserve dashboard functionality, you've admitted the data should have stayed a metric in the first place, and you've just added a fragile, latent aggregation layer.

Your savings will come from deletion, not migration. Every metric you consider moving should first be challenged for deletion. The remaining ones likely belong in your metrics system.


p-value < 0.05 or bust


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your lead's "theory" ignores the real question: why are you collecting 50% useless data to begin with? Moving junk from column A to column B is just an accounting trick.

The rule of thumb is simple: if you're asking about "cost savings after the switch," you've already lost. You'll save a visible line item on your metrics bill and drown in hidden logging compute costs and developer hours debugging aggregate queries.

Seen it three times. The only savings come from deleting things, not playing musical chairs with your telemetry.


-- old school


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

That's a really practical tactic with the review dates. I've found that the review date itself often becomes a formality unless you attach it to a cost center. What worked for us was tying the review to the team's own observability budget. When the calendar reminder pops up, the question isn't just "do you still need this?" It's "are you willing to allocate 5% of your quarterly logging budget to keep this query active?"

It creates a much more honest conversation about value.



   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Great question about the practical side. We actually followed a very similar path last quarter.

For the rule of thumb, we categorized metrics based on their use case in our customer dashboards. If it wasn't used for real-time health checks or immediate alerting, it became a candidate for logs. We saved about 20% on the metrics bill, but like others hinted, we had to be strict with log sampling to avoid just moving the cost.

The biggest gotcha was alerting latency. We used derived metrics from logs for some non-critical dashboards, but the 5-10 minute delay was a dealbreaker for anything customer-facing. Maybe you could start with a small, non-critical subset and measure the latency impact before moving 50%?



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Several replies have the right idea about query frequency and alerting SLAs, but they're missing a critical vendor-specific detail. Your lead's proposal assumes your logging vendor's pricing model is better than your metrics vendor. It often isn't.

With a mix of Datadog and Elastic, you're comparing indexed log volume costs to custom metric costs. If you start generating high-cardinality log events as metric replacements, you'll blow through your indexed log quota. The cost shift isn't just to compute, it's to a different, potentially more expensive, line item on your bill.

For a rule of thumb, add this: never move a metric that has more than 10 unique tag values to logs. The ingestion and indexing cost will erase any savings before you even run a query.


SLA is not a suggestion.


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

That "predictable bill to murky tax" analogy is painfully accurate. I'd add that the tax analogy breaks down because at least with taxes you get a bill. The compute and time costs here are invisible until you get a surprise vendor overage or a critical dashboard fails silently.

Your point about senior engineer time is the core of it. I've seen teams budget for the migration itself but completely ignore the ongoing cognitive load of maintaining a split system. Every time a log-based metric behaves oddly, you're not just debugging data, you're debugging a pipeline abstraction. That cost compounds and never shows up on a vendor invoice.


Test the migration.


   
ReplyQuote
Page 3 / 4