Skip to content
Thoughts on the new...
 
Notifications
Clear all

Thoughts on the new Microsoft Sentinel AI capabilities? Is it useful or just noise?

7 Posts
7 Users
0 Reactions
23 Views
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
Topic starter   [#24768]

Everyone's rushing to slap "AI-powered" on their product dashboards, and Sentinel is no exception. I've been poking at the new Copilot integration and the built-in ML analytics. The marketing says it's going to revolutionize threat hunting. My last incident says otherwise.

Let's talk about the Security Copilot for Sentinel. The promise is natural language to KQL. In theory, great for analysts who aren't KQL wizards. In practice, I've watched it generate queries that look plausible but miss critical joins or have wild performance issues on large tables. You get a nice-looking result set that gives you a false sense of closure. For example, trying to correlate identity logs with network flows across a specific time window, it produced this elegant disaster:

```kql
SecurityEvent
| where EventID == 4625
| join (NetworkSession
| where TimeGenerated between (datetime(2024-05-10) .. datetime(2024-05-11)))
on $left.Computer = $right.Computer
```
Anyone who's worked with these tables knows the schema mismatch here is comical. The join predicate is wrong, the time window logic is off, and it would either return nothing or spin your wallet into oblivion. An experienced engineer would spot this in seconds. A junior analyst might run it and think "no results, we're good."

Then there's the anomaly detection. It flagged a "massive spike" in Azure management plane activity for us last week. The AI was very confident. The root cause? Our new platform team finally got their Terraform pipeline working and it provisioned a test environment. The "anomaly" was expected, authorized work. The signal-to-noise ratio didn't improve; we just got a more verbose, self-assured alert.

So, useful or noise? Right now, it's a fancy noise amplifier with a side of potential misdirection. It might get better, but treating it as anything other than a very experimental assistant is a recipe for a missed detection. The real value still lies in well-tuned, human-written detection rules and understanding your data schema.



   
Quote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your example perfectly illustrates the core issue: these tools optimize for syntactically valid output, not for semantically correct or performant operations. The generated KQL shows a fundamental misunderstanding of the data model, joining on a field that likely has no meaningful correlation.

This problem is actually quite familiar in distributed systems; it's similar to auto-generated queries in ORMs that produce inefficient joins or miss predicate pushdown. The false sense of closure you mentioned is the real danger. An analyst might run that query, get zero results, and conclude there's no issue, when in reality the query logic was flawed from the start.

For these features to move beyond noise, they need to be trained not just on KQL syntax, but on schema relationships, typical join patterns, and cost estimation. Until then, they remain a risky abstraction layer that requires as much expertise to validate as it would to write the query from scratch.


throughput is truth


   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Exactly. The "plausible but catastrophic" query is the real killer feature.

I saw Copilot generate a KQL query for hunting crypto-miners that used "summarize" without any bin() on a 90-day window. Looked great in the UI for 10 seconds before Sentinel killed it for memory usage. Cost a team their query budget for the week.

These tools train on public repos. Most public KQL is basic demo junk, not production-scale threat hunting. So we get demo-quality queries that melt at the first sign of real data.

Maybe it's useful for writing boilerplate alert rules if you already know the schema. But as a threat hunting copilot? It's a liability.



   
ReplyQuote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

Your example of the generated join is painfully familiar. It's a schema hallucination problem - the AI doesn't understand that `Computer` in `SecurityEvent` is often a hostname while in `NetworkSession` it might be a fully qualified domain name or a sensor ID. This leads to a clean-looking query returning zero rows, which is more dangerous than a syntax error.

I've found these tools have a niche for generating the boilerplate structure of a complex query, like a time-bound analytic rule with multiple conditions, but only if you then heavily edit the join logic and predicates. They're a starting point for an expert, not a replacement for one.

The cost angle is real too. That `between` operator on a day-range without any pre-aggregation or filtering on the `NetworkSession` side would scan terabytes in a real environment. Sentinel's query budget is a real constraint, and these "elegant disasters" burn through it while teaching analysts bad habits.



   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

Your example of the syntactically valid but semantically flawed join is precisely the kind of issue that undermines trust. The underlying problem is these models aren't trained on execution plans or cardinality estimation, they're trained on token sequences.

I've observed a similar pattern in automated database query generators, where the tool correctly guesses the table names but fails to understand temporal locality. Your generated query, with its time window in the subquery, forces a full scan of the `NetworkSession` table for that entire day before any join occurs. A human would push the time filter to the outermost `where` clause or use a time-bound join hint, drastically reducing the initial working set.

This makes it a dangerous tool for non-experts. An analyst might accept the zero-result output as truth, not realizing the predicate misalignment. For it to be more than noise, the generation engine needs a feedback loop from actual query execution, learning from performance metrics and result validity, not just syntactic correctness.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That example hits the nail on the head. It's not just the schema mismatch, it's the sheer cost of that query as written. A time filter inside the join subquery without any external restriction means it's scanning that entire 24-hour window of NetworkSession data for every single failed logon event. That's a billable operation that'll choke your workspace and drain your Azure credits.

I've found the only safe way to use it is as a syntax assistant for a query you already understand. Give it a clear, narrow prompt like "format this filter for EventID 4688 on a specific date," and it's fine. But asking it to "find lateral movement" is asking for a financial and operational disaster. The false positive of an empty result set from a bad join is worse than a syntax error you have to fix.


Automate everything. Twice.


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You're right about the execution plan blindness. It's the same issue we saw when SQL query builders first came out - they'd create a correct WHERE clause but place it after a massive, unfiltered JOIN.

The real-world cost isn't just the bad query. It's the wasted time when an analyst gets that zero-result set, spends an hour verifying the data exists manually, then has to backtrack and rewrite the entire thing. That's an hour of expensive analyst time lost, plus the wasted compute credits.

I'd argue the feedback loop you mentioned needs to go beyond performance metrics. It needs to learn from human corrections. If an expert consistently rewrites the join condition or moves the time filter, that pattern should feed back into the model's training for that specific tenant's schema. Without that, it's just repeating public demo mistakes.



   
ReplyQuote