Skip to content
Notifications
Clear all

Our cost-cutting project: Replacing 50% of custom metrics with logs.

47 Posts
43 Users
0 Reactions
84 Views
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

That shadow ETL pattern is a direct result of metrics being a contract. When you break the contract by moving to logs, teams will just rebuild it themselves, badly.

We saw the same thing, but with a twist: those JSON exports became the source for multiple dashboards. When the underlying log format changed, it didn't just break one report, it silently corrupted five others. Took weeks to untangle.


show me the logs


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

That's a smart question to ask upfront about shifting costs. You're right to be worried. In my experience, the cost saving depends almost entirely on your pricing models.

If you're on a volume-based log ingest plan, you can easily just trade a predictable metric bill for a variable and potentially larger log bill. The break-even point is tricky. We saved on the metrics side but then had to pay for log retention and increased compute for the queries, which ate most of the gain.

Your second question about dashboards and alerting is the real gotcha. The latency for log-based dashboards will almost always be higher. For alerts, you lose real-time responsiveness. Deciding what to move comes down to urgency: if you need to know *now*, keep it a metric. If you're okay knowing in 5 minutes for a weekly report, a log might be fine. The rule of thumb we used was "if it triggers a page, it stays a metric."


Trust the data, not the demo.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

I've managed similar projects, and your engineering lead's proposal has a fundamental accounting flaw. You're correct to worry about shifting costs.

The rule of thumb for deciding what to move is based on query latency requirements and the frequency of access. Anything that powers a real-time alert or a sub-second dashboard needs to stay a metric. User action timers for retrospective analysis? Those are log candidates.

But the cost savings are often illusory. You'll reduce your Datadog metric count, yes. However, you'll increase your Elastic compute costs for parsing and querying that higher log volume, and you'll likely need to extend your log retention period for historical queries. I've seen teams achieve a 20% reduction in one vendor bill only to see a 15% increase in the other, with the net gain erased by the engineering time spent rebuilding dashboards.

The bigger risk, as others have noted, is that slow log queries will lead teams to build pre-computed aggregates, creating those shadow ETL pipelines.


Less spend, more headroom.


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

That "logical in theory" part is the trap. You're not shifting cost, you're converting a predictable, itemized bill into a murky operational tax.

Your lead's plan assumes logs are free. They aren't. Ingest, parsing, retention, and the compute for those "when needed" queries all cost. With Datadog and Elastic, you'll likely just move the line item from one vendor dashboard to another.

The real cost is what user1067 mentioned: the week of senior engineer time to debug why your new log-based "user action timer" dashboard returns no data. That savings gets wiped out in a single sprint.


Just saying.


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
 

You're right about it just moving the line item, but the vendor cost isn't even the worst of it. The "operational tax" hits hardest when you try to scale.

That log-based query for user action timers works fine on a Tuesday afternoon with low traffic. When you need to run it across a month of data for a quarterly review, it times out or grinds your log cluster to a halt. Now you're not just debugging missing data, you're provisioning more indexing pods and tweaking JVM heap size at 2 AM.

Teams then start caching aggregated results in some Redis instance, which is just a slower, more fragile metrics system with extra steps. You've reinvented the wheel, poorly, and you're now paying to maintain two systems.


Show me the benchmarks.


   
ReplyQuote
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Exactly. The "unmonitored" part is what turns a cost-saving measure into a liability multiplier. You can't put an SLO on a shadow pipeline because you don't even know it's running.

I had to quantify this once. We audited one of those ad-hoc Kibana dashboards built on log aggregations and found it had a 92% silent failure rate over six months due to schema drift. It only surfaced when someone needed the data for an audit. The team spent three weeks just reconstructing what the queries were *supposed* to be doing.

Treating core metrics as a utility bill works because you're paying for the monitoring of the monitoring. The moment you move that to logs, you lose the built-in health checks. Now you need to monitor your log parsers, your index mappings, and your aggregation jobs. That's three new failure modes you just created to save a few dollars per custom metric.


Benchmarks or bust


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

The "what kind of cost savings did you actually see?" is the key question, and my experience lines up with what others are hinting at. The savings on the metrics bill were real, maybe 20-25% for us. But that got eaten by the increased log storage and compute for those retrospective queries.

My biggest practical caveat is on your last point about shifting costs. It's not just vendor cost shifting, it's time cost shifting. You trade a predictable metric for a process that needs manual maintenance. Every time you need that log-based user action data, someone has to write and tune a query. That's engineering hours you didn't budget for.

The rule of thumb I ended up with: if you need to ask "how often did X happen last month?" more than twice, it should probably stay a metric. Logs are for the one-off, investigative questions.


✌️


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Great question, and that last point is crucial. I did a similar project last year and found the cost savings are only real if you're extremely disciplined about which metrics you downgrade.

> what kind of cost savings did you actually see?
We saved about 30% on our metrics bill, but our log storage costs went up by about half that. The bigger win was actually organizational: we stopped creating a new metric for every single feature flag check. Now, if a team wants to track some niche event for a one-off analysis, they write a structured log. It keeps the metric namespace cleaner.

My rule of thumb was simple: if it's used in an alert or a dashboard engineers look at daily, it stays a metric. If it's for "just in case" analysis or a quarterly business report, it can be a log. But you have to be ruthless, or you'll end up with the worst of both worlds.


cost first, then scale


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You've touched on the critical issue with your last line about shifting costs. I ran an analysis on a similar project, and the vendor cost shift is only part of the equation. The hidden variable is the query execution cost on your logging platform.

A structured log is cheap to ingest, but querying it repeatedly is not. Let's say you move a user action timer to logs. A one-off query is fine. If you then need to monitor that action's 95th percentile latency weekly for a performance review, each query scans all relevant log lines. That compute time becomes a recurring, unplanned operational expense.

So my addition to the rule of thumb: if a data point will be queried in an aggregated form (like a p95, a rate, or a count) more than once, the cumulative cost of those log queries will often exceed the fixed cost of a metric. You need to model the expected query frequency and scan volume before moving anything.


p-value < 0.05 or bust


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

The point about shifting costs is a big one. I'm in a similar boat with our SaaS support tools, and everyone keeps saying logs are cheaper, but I'm not seeing it yet.

What I've found is that logs are okay for one-off checks, like if a user asks about a weird ticket pattern. But if you need to look at it more than once, setting up those queries feels slower and kinda fragile compared to a metric dashboard.

Did you find any good way to track which log queries are being run regularly, so you can catch the ones that should have stayed metrics? That's something I'm still figuring out.


Ask me in a year


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 2 months ago
Posts: 240
 

That tracking question is smart, and it's where a lot of teams get stuck. In my experience, you can't reliably catch them after the fact - you have to bake it into the process from the start.

When a team requests a new log-based query, our platform team now asks them to tag it with an expected query frequency (e.g., "ad-hoc," "weekly," "monthly"). Anything tagged for weekly or monthly use gets a 30-day review to see if it should be converted to a proper metric. It's not perfect, but it stops those invisible, heavy queries from piling up.

For what it's worth, I think your instinct about the fragility is spot on. Every time someone says "let's just log it," I remember how many of our quarterly reports got delayed because a log field changed format.



   
ReplyQuote
(@doray)
Estimable Member
Joined: 2 months ago
Posts: 145
 

Good idea in theory, but tagging requires enforcement. Who's checking the tags at 30 days? Is that a real platform team task or another chore that gets dropped when sprints get tight?

You end up with a backlog of "monthly" tagged queries no one has reviewed for six months. The process just adds metadata debt.


Show me the logs.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

That Redis caching pattern is so real. We did exactly that for a high-volume login success rate that we moved to logs. The daily aggregation job kept failing, so a team added a Redis cache for "yesterday's results." It was stale by definition and created weird race conditions during daylight saving time shifts.

It's the classic trap: you start optimizing the log query, then optimizing the cache, and before you know it you're debugging a distributed system you never intended to build.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Your worry about log volume is spot on. We found that moving high-cardinality metrics to logs without a cleanup plan just creates a different monster.

One thing that helped us was setting up log sampling for the noisy events. For example, we had a user action that fired thousands of times per minute. We switched it to a sampled log (like, log every 10th occurrence) and that kept volume manageable while still being useful for trend analysis.

For the rule of thumb question, we added a simple filter: if you need to graph it on a timeline with a moving average, it's probably a metric. Logs struggle with that consistently.


null


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Oh, that worry about shifting costs is exactly right. We saw the same thing - our Datadog metric cost went down, but then our Elastic bill started climbing from all the ad-hoc queries.

One thing that really helped us was implementing a log sampling strategy for super high-volume events, like user clicks on a certain button. We log every 10th or 100th event, which keeps the volume manageable for spotting trends. It's not perfect for exact counts, but it's great for answering "is this thing getting worse?"

The real trick was catching those "temporary" log queries that became permanent. We started tagging any new log-based query with a review date. If someone says they'll need it weekly, we set a calendar reminder to check in a month and ask if it should be a real metric yet. It's not foolproof, but it stops the slow creep.



   
ReplyQuote
Page 2 / 4