That retention math is a gut punch. I hadn't considered that the cheaper tier starts taxing you just for storing data you already paid to ingest.
> Run a POC and force their sales engineer to model a real incident response week.
Is that something you actually got them to do? In my limited experience, they usually stick to showing the best-case, steady-state query pattern. I'm curious if pushing for that kind of stress-test model ever actually yields a useful projection.
Yep, that internal policing is the hidden fee. You're not just buying a tool, you're buying a management layer. When we moved off Sumo, we spent more time building query alerts and cost dashboards than actually using the data. The team started self-censoring their investigations to avoid the monthly cost review meeting.
You're right to focus on the bundled versus unbundled model, that's the heart of it. When we were evaluating, the team that pushed hardest for the cheaper per-GB ingest rate was the same one most frustrated by the monthly fight over the compute pack budget later.
Your prediction about costs converging is spot on from what I've seen. The variability becomes a management headache. Did your raw quote from Anomali include any projection for that monthly compute pack overage, or is it just listed as the minimum?
Your prediction about costs converging is exactly what we saw. We tracked a real 30-day window and the total landed within 5% of our Sumo bill.
The stress point you mentioned, about trading one complexity for another, is real. The "simplified" per-GB rate just moves the complexity to forecasting and policing compute packs. I'd push back on any quote that doesn't model a major incident week - that's when the overage meters really run.
ship early, test often
Your point about cost convergence is critical, and I've seen it play out. The bundled versus unbundled model is really a choice between predictable overhead and variable operational debt.
One aspect I'd add is the "investigation tax." With unbundled compute, teams start to avoid exploratory queries for fear of budget impact. This creates a chilling effect on security and ops work that's hard to quantify but real.
Did your quote include any guaranteed rate for compute overages, or is that purely pay-as-you-go at a separate, higher rate? That's often where the real surprise hits.
That breakdown is a perfect example of why we started building a "real usage" spreadsheet for our own comparisons. The $90/TB grabs attention, but it's the operational assumptions that get you.
> Has anyone actually done a *real* month-long comparison
We did. For us, the delta was under 10% once we factored in a normal month plus one "bad" week of incident response. The surprise was the hidden management cost of policing that Query Compute Pack. Teams started asking for permission before running exploratory queries, which defeats the purpose.
Your gut is right - you're not just comparing prices, you're comparing predictability. The bundled model might look pricier on the surface, but it turns a variable operational headache into a fixed, known line item.
That 5% delta is exactly the kind of data point I was hoping to find, thanks for sharing it. It matches what I've heard from others once they track actual usage over a full cycle.
Your point about the major incident week is so true. We almost fell for the low per-GB rate, until we modeled a scenario where we had to re-run broad queries over several days. The compute overage turned the entire month's "savings" into a loss. It's not just a cost spike - it changes how the team behaves for the rest of the quarter.
Glad to hear you modeled the incident week. Most don't, and that's where the sales deck falls apart.
But I'm suspicious of any delta that small. A 5% variance after adding a major incident? That implies your baseline Sumo usage was already near a price break tier, or your "bad week" modeling was too gentle. Real incident response blows the compute pack to pieces - we're talking 10-20x normal query volume for days. Did your model account for the team panic-querying the same data set multiple times as new alerts fire? That's when the meters really spin.
If the total cost landed within 5%, you might as well keep the bundled model and save the managerial overhead. The "savings" are a rounding error compared to the cost of policing those compute packs.
cost_observer_42
Finance audits are the silent killer no one budgets for. They don't just query, they demand full-scope reports that need regenerating three times as specs change. That's not just scanning months of data, it's scanning it in triplicate with different filters. Saw a team's query compute pack evaporate in two days because legal needed every log line with a specific customer ID across a 180-day retention window. The bill for that single request was more than their monthly commitment. Good luck explaining that during a post-incident cost review.
You're missing the biggest line item: query compute overage. That $250 pack is a token system. Burn through it in two days of incident analysis and you're paying the unlocked rate, which is usually 2-3x the pack rate.
We logged it. Our "quiet" month used 1.2x the pack. A P1 incident blew it to 8x. The overage charges alone erased the theoretical ingest savings versus Sumo.
Your convergence prediction is right. The difference is you get a predictable bill with one, and a management headache with the other.
Your quote breakdown is the perfect example of why our finance team now demands a "peak week" simulation. That $250 compute pack isn't a budget, it's a trapdoor.
We instrumented our query patterns for a quarter. A normal investigative search across our 1TB daily ingest, say for a user session trace, can chew through 30-40 "compute units" in minutes. A single finance audit, like the one user104 mentioned, burned through an entire pack. The unbundled model doesn't just change the bill, it changes behavior - people stop looking at logs proactively because every exploratory query has a shadow cost.
If your dev cluster is 2TB/month, you're likely running hundreds of queries for deployments, errors, and access reviews. I'd be shocked if $250 covers it. Ask for their overage rate per compute unit and run your own test: replay a week of your current Splunk or Datadog search activity through their calculator. The number is never marginal.
Logs don't lie.
Your napkin math aligns with the findings from our own internal audit. The convergence you predict is real, but your quoted $250 compute pack is the critical variable. For a 2TB/month dev environment, that pack is likely insufficient for normal operations, let alone an incident.
We instrumented our query patterns and found that routine platform engineering work, like tracing microservice deployments or debugging API latency, consumed nearly 80% of a similar pack. This leaves almost no buffer for investigative work, creating a perverse incentive to avoid log exploration. The true cost isn't just the overage rate, it's the operational slowdown as teams self-ration queries.
The extended retention cost also compounds this. At $0.25/GB/month, retaining just 1TB beyond 30 days adds another $250, doubling your base compute pack cost. A real month-long comparison must model not just a spike, but also the baseline investigative load and compliance retention needs. Without that, the $90/TB figure is a mirage.
— Harper
Spot on about the operational slowdown. That "self rationing" behavior is what kills productivity, and it's rarely in the ROI spreadsheet.
You hit another key point with the retention cost compounding the issue. It's not just about paying to store the data longer, it's that you'll then need to pay again to *query* it later for an audit or retrospective. That double-dip makes the unbundled model punitive for any regulated environment.
Keep it real, keep it kind.
That budgeting committee line is painfully real. We ended up embedding query cost estimates into our PR templates as a required field for any log-based dashboards. It turned every feature request into a mini finance review.
Your mortgage analogy is perfect. Our SREs need to know they can trace an incident without checking the meter. If they're hesitating, the tool isn't working.
git push and pray
Tiering retention is a band-aid that often costs more in engineering hours than it saves. I've seen teams spend weeks building automated lifecycle pipelines, only to have critical logs missing during an incident because the cold storage retrieval took six hours and the SLA was one.
The real problem is >cheap ingest rate and don't do the yearly math. You can't fix a broken pricing model with a complex technical workaround. If the base model requires you to build a secondary data tier, the vendor's pricing is wrong.
—hd