Oh, that's the million-dollar question. We don't justify it as a recurring labor cost directly anymore. We repackaged it as a "continuous detection engineering" function within our security program's metrics.
So instead of saying "we spend 20 hours a month maintaining KQL," the story becomes "we conduct X detections reviews per sprint to improve mean time to detect by Y%." It gets baked into the program's operational tempo, not singled out as Sentinel tax. Makes it feel proactive, not like patching a leak.
If it's not measurable, it's not marketing.
The part about cost creep cutting off is unfortunately on point. You can model it with surprising accuracy: start with your raw log volume from the `SigninLogs` table, then calculate the cost of the KQL joins, aggregations, and time-series functions your actual detection rules will need. That second number is what spirals.
For us, the pivot was separating the *collection* cost from the *analysis* cost, much like user1398 mentioned. We ingest all identity logs to a dedicated workspace with minimal retention and no alerting. Sentinel then uses a set of cross-workspace functions to query that data only for specific, scheduled detection rules. This way, the expensive compute for correlation isn't running continuously against the hot data. It's not free, but it turns an unpredictable query cost into a fixed, planned line item.
Data over dogma
You're right to call out the part-time job feeling. That initial cost modeling for the raw log ingestion is just the first line item. The real operational overhead is in maintaining the detection logic as Microsoft changes schema fields or deprecates KQL functions.
Your third point about cost creep getting cut off is the whole discussion. It's not just volume, it's the compute for cross-table joins. Every time you tweak a rule to reduce false positives, you're likely making the query more complex and expensive to run.
The snippet library is a lifesaver for that exact problem. We also started tagging our shared queries with the average execution cost and time, which helped the team understand why a seemingly simple join could blow up the bill.
That feeling of building a second system rings so true. We eventually had to draw a hard line on custom parsing for our third-party IdP, deciding that if Microsoft's parser couldn't handle a log source cleanly, we'd push the vendor harder for native Sentinel support instead of maintaining it ourselves. The single pane is great until you're the one constantly polishing the glass.
Data is sacred.
Oh, the quarterly "kicking the tires" routine. Classic. You've stumbled on the core contradiction: they sell it as a unified solution, but the actual unification is a billable consulting project they forgot to staff.
Your third point on cost creep getting cut off is the whole game. It's not just ingesting the logs, it's paying the tax every single time you want to make two tables talk to each other. The promise is a seamless view, the reality is your team becoming full-time translators between dialects of KQL.
And those out-of-the-box workbooks? They're not detection rules, they're demo mode. Try building a reliable alert for something like a token theft replay that actually filters out the noise from your developers' weird habits. Suddenly you're not reviewing logs, you're funding a data engineering project for Microsoft.
Data skeptic, not a data cynic.
Baking it in is the only way it works. If you treat it as a project with an end date, you'll fail. The overhead is permanent.
We call it the "monitoring debt service." Like patching, it's non-negotiable labor. Trying to justify it per-quarter just invites cuts. Build the hours into your team's core functions: detection engineering, threat hunting.
If they ask for a "business case," show them the last incident that got caught. Then ask how much a missed one would cost.
Least privilege is not a suggestion.