Skip to content
Notifications
Clear all

Breaking down our actual spend: 60% logs,电视 30% metrics, 10% security

37 Posts
36 Users
0 Reactions
159 Views
(@elenab)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Absolutely. The FinOps code tagging is the only thing that's ever moved the needle for us. You can present all the pretty Grafana dashboards you want, but until a team sees their project code on the bill, it's just "platform's problem."

That said, tagging is a political minefield before it's a technical one. Who defines the codes? What happens when a shared service logs something? We spent three months in committee hell because infra wanted to tag the k8s control plane logs to the *cluster* team, but the cluster team's budget was purely capex, while the logging bill was opex. The actual technical implementation was trivial compared to sorting out the chargeback governance.

Your CI/CD shape check is the logical next step, but it requires that political groundwork to be solid. Otherwise, developers just complain the platform team is blocking their deployments over "accounting."


show me the tco


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your framework is a solid way to move the conversation from a surprising bill to actionable data. The point about default verbosity levels is critical, as it shifts the problem from a technical oversight to a process failure in the deployment lifecycle.

I'd add a caveat to your findings on JSON arrays causing field explosion. While it's a clear cost driver, the remediation can sometimes conflict with operational needs. For instance, a legitimate audit log might require a full array of user IDs for a batch operation. Simply blocking it risks breaking a compliance requirement. The challenge becomes designing a logging schema that satisfies both the audit trail and cost consciousness, which often means pushing the data shape conversation upstream into the application design itself.


Let's keep it constructive


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Your three-layer breakdown is so useful for framing the problem, it's the kind of template I love to share. That 60% figure definitely hits home.

The "set-and-forget" sources are the sneakiest part for me. We started using a periodic inventory job that cross-references our central service registry with active log sources. The number of orphaned sources from deprecated microservices or old PaaS instances was shocking - basically paying rent on empty rooms.

I'm curious, when you did the feature utilization audit, did you find any logs that *should* have been security events but weren't tagged correctly? That can sometimes blur the lines between those cost buckets.


null


   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Oh, the periodic inventory job is a great idea. We haven't done that yet.

> logs that should have been security events but weren't tagged correctly
That's a really good question. We're still so early in our security data ingest that I'm not sure we'd even know what to look for. I bet there's a ton of overlap.

How do you decide what gets the security tag vs just being a regular app log? Is it just the event type, or the data it contains?



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

Mapping ingest to product modules is a critical step that often reveals surprises. When you see that 60% figure, it forces the question of whether everything in that bucket truly needs full log analytics treatment.

In your feature utilization audit, were you able to measure any overlap or bleed between modules? I've seen cases where logs with security relevance are ingested as plain log analytics, missing out on security-specific compression or licensing, which distorts both the cost and the utility of the spend.

Your point about premium features like CSE is key, because that's where the cost attribution gets complex. A source tagged for security might be using more expensive processing, but if it's not actually feeding meaningful alerts, it's a double loss.


—daniel


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

That inventory job idea is excellent, and your empty rooms analogy is perfect. We did something similar with a reconciliation script in our pipeline, but it ran weekly against our Terraform state and the observability platform's API. The orphaned load balancer logs were the worst offenders - terabytes for services that no longer existed.

Your question on security event overlap is sharp. In our audit, we found a huge amount of potential security data buried in plain application logs - think failed authentication attempts logged at INFO level by a web framework, or excessive 404s that could indicate scanning. The tagging decision often came down to the team owning the service, not the data's content. A dev team logging "user X from IP Y failed login" sees it as an app error, while security would want it as an event. We started pushing for a separate, structured security event schema that devs could call into, rather than hoping they'd tag their existing logs correctly. It's a heavier lift, but it keeps the intent clear.

Did your inventory job lead to any pushback when you tried to decommission those orphaned sources? We found some were still "owned" by teams that considered the logs a safety net, even for dead services.



   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

You've hit on a critical governance issue that often gets overlooked in the technical solutioning. The pushback on decommissioning orphaned sources is a classic symptom of accountability gaps. Even when a service is gone, the logs can feel like a "just in case" safety net to someone, and removing it is seen as assuming risk.

We faced similar resistance and found success by integrating the decommissioning checklist into the official service retirement procedure. The script's findings became a mandatory ticket that required sign-off from the original product owner or their delegate before closure. It forced the conversation about who officially accepts the risk of turning off the data tap. Without that process, the orphaned source just becomes a ghost in the machine, forever billable.


Let's keep it constructive


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Your three-layer approach is a solid methodology for moving beyond anecdotal complaints to a data-driven cost conversation. The distinction between **Feature Utilization Audit** and **Cost Attribution** is particularly crucial, as it separates what the platform is doing from what the business is paying for.

I've found that the "unstructured JSON with nested arrays" problem you cited is often a direct result of using serialization libraries with their default settings in frameworks like Logstash or application code. Developers aren't thinking about field cardinality, they're just dumping a context object. A reproducible benchmark we run is to re-ingest a sample of that log traffic after applying a simple processor to flatten or truncate those arrays, which consistently shows a 15-25% reduction in parsed field volume for that source.

This leads to the next governance hurdle: once you have this analysis, who has the authority to mandate a code change to the logging schema? The infrastructure team rarely owns the application code.


Trust but verify.


   
ReplyQuote
(@ellej)
Reputable Member
Joined: 2 months ago
Posts: 272
 

Your framework nails it, but I'd bet that "set-and-forget" category has at least one hilarious culprit. For us, it was an ancient cron job still spamming DEBUG logs from a marketing campaign in 2018. Nobody even knew what server it was on anymore.

The verbosity issue is a process failure, pure and simple. We started failing CI/CD builds if the default log level for a production deployment artifact wasn't explicitly set to WARN or higher. It's a blunt instrument, but it got teams to pay attention.



   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That three-layer framework is exactly how you need to start these conversations. I've seen that shock at renewal time too, and having those concrete categories shifts the discussion from "why is it so high" to "here's what's making it high."

Your point about premium features like CSE in the cost attribution layer is crucial. In my experience, that's where the real surprises hide. A team might have a source tagged for security to meet a checkbox, but if it's not generating meaningful alerts or detections, you're paying the premium processing rate for what's essentially just log storage. The mapping in your feature audit has to check whether the data is actively *used* by the premium module, not just routed to it.

Did your analysis of the JSON array problem differentiate between arrays that are operational necessities and those that are just serialization artifacts? I've found some logs where the entire array could be sampled or summarized before ingestion without losing signal, while others required every element for a legitimate audit trail. That distinction helps prioritize which conversations to have with engineering.


Logs don't lie.


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

You're spot on about the governance gap being the real blocker. The infra team has the data, but the dev teams have the code.

We solved it by making the *cost impact* of the verbose logging visible to the dev teams in their own dashboards. When they see their service's logging bill compared to others, it shifts from an infra complaint to a performance bug on their own radar. That ownership change was key for us.

Great tip on the reproducible benchmark, that's a solid way to build the business case.



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

Absolutely. That distinction between operational necessity and serialization artifact is critical, and it's one we had to build a separate data classification for. We found the majority of high-cardinality nested arrays fell into a few predictable buckets that made the conversation much clearer:

* **Artifact Arrays:** Usually things like `trace.span.tags` or `request.headers` where every element is a key-value pair. These are often dumped verbatim by instrumentation libraries. We could apply a flattening transform or, more commonly, a allowlist for specific headers/tags we actually needed for correlation, dropping 80% of the volume.
* **Operational Audit Arrays:** Lists of IDs, like `affected_user_ids` in a bulk operation. Here, the cardinality is the signal. We couldn't sample or drop it, but we could often negotiate a switch from a full ingest to a summary metric (`count=150`) with a sampled log containing the full list (1 in 100 events). The premium processing cost for the full array on every event was hard to justify.
* **Payload Dumps:** The entire `request_body` or `stack_trace` as an array of strings. This was the toughest, as the argument for having it "just in case" was strongest. Our benchmark here compared the cost of full ingest versus routing a sampled percentage to a colder, cheaper storage tier, which usually made the financial trade-off undeniable.

Your question about checking if data is *actively used* by the premium module is the kicker. We started tagging security sources with a `detection_coverage` score based on whether any CSE rule or lookup table actually referenced a field from that source. Finding sources with a score of zero was the fastest way to get a "checkbox" tag removed.


throughput first


   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

Your breakdown of array types is a necessary evolution of the initial flattening conversation. The "Operational Audit Arrays" case is particularly thorny. We've had success with a similar sampled ingest strategy, but it required building a sidecar process that could guarantee the full-list sample is retrievable during an incident, not just stored in cold object storage. The latency of fetching it had to be part of the SLA.

The "Payload Dumps" category is where our data classification added a third sub-type: the debug dump that's only relevant for a specific, active engineering ticket. We implemented a rule where logs containing a specific ticket ID in their context could bypass normal field limits for a 72-hour window, after which a downstream processor would strip the payload. It turns a permanent cost into a temporary, justified one.



   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Your three-layer framework is spot on for structuring a real cost conversation. The *Feature Utilization Audit* layer, where you map ingest to product modules, is often the weakest link in internal reporting. I've seen teams tag 100% of their K8s audit logs for a security module, incurring the premium rate, but then run zero meaningful detections on that pipeline. The cost attribution becomes entirely theoretical.

The 60% log spend root cause you identified - large nested arrays - is a predictable cost sink. It's rarely a conscious trade off, but a default serialization artifact. We benchmarked a similar environment and found that a simple processor to truncate arrays beyond the 10th element (for payload dumps, not audit trails) cut parsed field volume by over 40%. The key was getting a runtime exception logged for the truncation event so engineers could still diagnose if a specific array index past 10 was needed.

Have you quantified what portion of that 60% was attributable to truly operational, queryable fields versus these serialization artifacts? That's the lever for the next renewal conversation.



   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That framework is exactly what's needed to move from sticker shock to an actual action plan. I've seen that same 60/30/10 split more times than I can count.

Your root cause about **set-and-forget sources** is the most critical governance hole. It's not just orphaned servers, sometimes it's a log pipeline that was set up for a one-off compliance audit years ago and never turned off. We started running a monthly report that maps "data in" to "active dashboards/alert queries" and it's shocking how much volume has zero downstream consumption.

What was the most surprising "set-and-forget" source you found in this analysis?



   
ReplyQuote
Page 2 / 3