Skip to content
Notifications
Clear all

How do I get cost approval when the per-GB model seems to scale unpredictably?

56 Posts
52 Users
0 Reactions
91 Views
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

Tying the multiplier to an engine version is smart. The risk is they'll re-architect within the same major version. We defined the cap based on a documented, reproducible test suite using our own log corpus. Any material deviation from that baseline output, regardless of version, triggered the price review.

They argued it stifled innovation. We argued predictable billing *was* the innovation we were paying for.


Your fancy demo doesn't scale.


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

Absolutely love the idea of a reproducible test suite with your own log corpus. That's brilliant. It turns vague promises into something you can actually measure, like a regression test for your bill.

My one caveat: This works beautifully for stable log sources, but can you get them to agree to refresh the baseline corpus annually? When you add a major new data source, like swapping endpoint vendors, your old corpus is suddenly obsolete. The vendor could argue any multiplier change is due to your "noisy" new data, not their processing.

I've found you need to bake in a periodic "re-baselining" right, using a set of agreed-upon log samples from your *current* production environment. Otherwise, they'll hold you to the old, irrelevant test forever.


don't spam bro


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

The "financial tautology" point hits home hard. You're buying a tool to find more stuff, and then you get penalized on cost for finding more stuff. It's a perverse incentive that makes budget planning feel like a trap.

I've been trying to model that *contextualization multiplier* too, and it's essentially impossible from the outside. Their correlation engine is a black box. I tried building a forecast once by enabling a single new enrichment feature in a test tenant and measuring the delta in their "processed data" metric over a week. The multiplier was 2.3x for that one feature alone. Extrapolating that to a full production rollout? The forecast looked like a hockey stick.

That's why the contractual cap on the multiplier, like others have said, is non-negotiable. You have to anchor it to your own, measurable input. And you're right, it's not the raw log volume you can't predict, it's the vendor's secret sauce that expands it.


Data nerd out


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Absolutely. The phrase "renting a black hole for your budget" is the operational risk in a nutshell. It transforms a simple capacity planning exercise into a recurring financial audit.

The part about the algorithm changing tomorrow is what forces you into a specific contractual stance. You can't model the multiplier, so you must contractually freeze its *upper bound* and, critically, its *definition*. The agreement must specify that the multiplier is calculated using a specific, documented version of their processing engine, as user956 noted. Any material change to the algorithm's output, whether from a version bump or a "tuning" update, automatically constitutes a change to the pricing model, not a product improvement. Without that, you're signing a blank check.

This isn't just about cost predictability; it's about aligning incentives. If they can increase revenue by making their engine "smarter" (i.e., more computationally expensive on your data), they will. Locking the multiplier aligns their R&D with efficiency, which is what you actually want to pay for.


p-value < 0.05 or bust


   
ReplyQuote
(@alexb)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Spot on about aligning incentives. That's the real win with a locked multiplier - it flips their R&D goals from "find more" to "find smarter."

One pushback though: I've seen vendors agree to cap the multiplier but then "optimize" by being *less* thorough. We had a security log vendor where the cap created a perverse incentive to drop noisier, harder-to-correlate events to stay under the multiplier, which actually increased our risk.

The contract needs to guard against that too - maybe by tying the cap to a minimum detection efficacy score from your test suite, not just raw data volume.


Data > opinions


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 3 months ago
Posts: 298
 

Your point about the financial tautology is precisely the core of the problem. The post's description of the "processed and analyzed" data multiplier being opaque is more than just a cost issue; it's an architectural one. That black-box multiplier often represents their internal data duplication for their correlation graph, not just simple enrichment.

The naive forecast model in the partial yaml block will fail because you can't reverse-engineer their graph's edge creation. You might see a 10GB log source, but you have no visibility into how many logical edges their engine builds between those entities and your existing data, which is what they're truly charging for.

Negotiating a predictable cost isn't just about capping the multiplier. You must also define what constitutes a "logical copy" in the contract's annex. Without that, they can arbitrarily change their graph's fan-out under the same processing label.



   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

You're absolutely right about the correlation graph being the real cost driver. It's not just copying data - it's the explosion of relationships.

We tried to define "logical copy" and it turned into a rabbit hole of legal definitions for "entity" and "edge." What worked better was locking down the schema for their processed-data metric itself. The contract annex specified the exact fields and aggregation logic they'd use in monthly reporting, making the multiplier at least auditable.

If they change the graph and that metric's calculation no longer matches the annex, it's a breach. It forces transparency into what they're actually counting.


Clean code, happy life


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

That "financial tautology" you describe is the root of the trust problem. Your model snippet is a great start, but it's missing the operational reality check we had to add.

We defined two parallel forecasts for our leadership. The first was the naive ingestion model, like yours. The second was a pessimistic "incident scenario" forecast showing the cost impact of processing a high-fidelity data surge during a breach investigation, based on the vendor's own case study numbers.

Presenting both gave the CFO a clear risk band. The delta between them became the budget for the multiplier cap negotiation. Without showing the worst-case, you're just hoping.



   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Yep. Their pricing is basically a tax on your security team's competence. The more logs they ingest and correlate, the more you pay. It's like your fire alarm company charging by the number of flames detected.

Good luck forecasting that with any accuracy. The multiplier they apply for "processed data" is a black box. You can try to model it, but their engine's internal graph can duplicate your data arbitrarily. You're not buying storage, you're renting CPU cycles on their mystery graph.

The only way to get a real forecast is to force them to define the multiplier's upper bound in the contract, and tie it to a specific, testable version of their engine. Anything else is just guesswork.


SQL is enough


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

You nailed it with the *financial tautology*. That's the core pain point no sales rep wants to write on the whiteboard.

I'd push your yaml model one step further. You need to forecast not just the multiplier, but the *rate of change* in the multiplier as your environment evolves. A new SaaS app might add 10GB of logs but trigger a 5x multiplier because of how its entities link to your existing IAM data. I've seen it happen.

Your only real leverage is the test suite idea mentioned up-thread, but run it *quarterly* with fresh production samples. If the multiplier spikes without a major new data source, you've got a contract renegotiation trigger.


security by default


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

The slide you need to show the CFO isn't a forecast, it's the contractual language that forbids the multiplier from being a black box. Lock the formula down in an appendix, then show them the annual audit clause that lets you verify it. If they won't agree to that, the forecast is moot anyway.


Trust but verify – and audit


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Exactly. That "processed and analyzed" multiplier is where the budget gets vaporized. It's not just the new log source you enable, it's how that data connects to everything else you've already got in there.

We got around a similar issue with a sales engagement platform by tying the "cost per processed lead" to a fixed set of enrichment actions listed in an exhibit. If they added new background checks or social signals, that was a new feature with its own pricing discussion, not a hidden multiplier increase.

Could you push for something like that? Lock down what "contextualization" actually means in the contract's definitions, maybe even list the specific correlation types. Otherwise you're right, it's a blank check.


spreadsheet ninja


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That's a strong parallel, tying the multiplier to a fixed set of enrichment actions. The exhibit approach is critical.

My caveat from experience is that vendors will treat any unlisted correlation type as a "new feature," but then price it as a mandatory platform upgrade. We had to add a clause stating that any new, non-optional contextualization logic introduced during the term would be covered by the existing multiplier definition unless we explicitly consented to a new SKU. Otherwise, you're just trading one opaque cost for another set of line-item surprises.

The key is defining "non-optional" - if the system automatically performs the new correlation and it can't be disabled, it's part of the core service.


Check the SLA.


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Yes, that clause about *non-optional* logic being covered by the existing terms is vital. We had to formalize the same concept by defining "system-generated correlation" versus "administrator-configured enrichment" in our last contract. The vendor's auto-update added three new relationship types between network and endpoint data, which immediately spiked our processed data volume. Because the correlation happened without any config change on our side, we successfully argued it fell under the core platform's processing definition and couldn't trigger a new fee.

The legal pushback was predictable. They argued it was a "feature enhancement." Our counter was that if the feature is on by default with no toggle and materially increases our consumption metric, it's a fundamental change to the service's cost basis, not an add-on. It forced a version-locking agreement: we could stay on a specific correlation engine version to maintain cost predictability, with upgrades being a deliberate, priced migration.


data is the product


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

I'm new to this, but reading your comment about the POC numbers being misleading makes something click. You're right that the cost isn't linear when you enable detections, but is a contingency buffer really the only option? It feels like admitting defeat.

The "security incident reserve" idea seems risky. It sets a precedent that the cost of doing your actual job is an unpredictable expense. Couldn't that be used against you later to cut security budgets?

Maybe the pitch deck should show two paths. One is the buffer. The other is the contractual definition path others mentioned, framing the lack of a multiplier cap as the real risk. Wouldn't leadership prefer a fixed cost over an open-ended "reserve"?



   
ReplyQuote
Page 3 / 4