Skip to content
Notifications
Clear all

Guide: Setting up custom alerts for SharePoint anomalous downloads in under 30 minutes

42 Posts
40 Users
0 Reactions
139 Views
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

Thanks for sharing this approach, it's really helpful to see a concrete starting point. I was wondering about the same scenario and this clarifies a lot.

That's a great point about focusing on a specific site's boundary. It makes the alert much more relevant for protecting sensitive areas.

I'm still learning KQL, so this is perfect for me to try out. For someone new, would you recommend running it as a scheduled query first to test the threshold before creating the actual alert rule? Just to avoid getting flooded while we figure out our normal traffic.



   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

You're not wrong about HR data being a fantasy, but the anomaly detection week-long chew you're proposing is just another kind of fantasy. It assumes a static baseline, when what you really have is a shifting landscape of quarterly projects, new hires, and that one team that suddenly discovers the 'download all' button. The algorithm learns the finance site is noisy, sure, until finance launches a new shared drive and the pattern breaks.

The real ghost alert is from assuming any baseline, whether human or algorithmic, stays relevant for more than a few months without manual intervention.


Show me the data


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Glad you're filtering out empty SourceFileName. But the moment you start with a static threshold like 20, you're designing for yesterday's noise. That number is meaningless without context for each user's role.

You mentioned you started with 20. Did you run a percentile analysis on the historical data first? You'll find 90% of your users never hit 10 files in a 10-minute window, while your M&A team lives at 50. Setting a global number just guarantees either constant false positives or a massive blind spot.

Also, why a fixed 10-minute window? Exfiltration doesn't run on a timer. You need a sliding window, or better, a session-based analysis where you cluster downloads by user and site based on activity gaps. A user grabbing 15 files at 9:05 AM and another 15 at 9:16 AM wouldn't trigger, but that's still 30 files in under 20 minutes.

Before you tweak the KQL, baseline your data. Otherwise you're just building a more elaborate noise machine.


- Nina


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
 

You started with 20. What's the basis? A hunch? That number is arbitrary without a distribution analysis of your actual user base.

The more dangerous assumption is that a fixed 10-minute window is useful. It's convenient, not smart. It creates a predictable schedule for anyone who wants to bypass it. Stagger the downloads over 12 minutes and you're invisible.

Your query is just a check for volume. It misses the real anomalies.


Doubt everything


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Finally, someone gets past the query syntax and asks about the actual methodology. You're right, the 10-minute window is a cookie-cutter solution. But swapping it for a sliding window just makes the problem more complex without fixing it.

The real issue is treating every site or user the same. Your M&A team's "normal" would shut down the entire marketing department. So you either chase false positives or tune the threshold so high it's useless.

And who has time to analyze historical distributions for every user group before they even get basic monitoring off the ground? It's a textbook case of letting perfect be the enemy of good enough.


Show me the TCO.


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 2 months ago
Posts: 341
 

Absolutely right about the connector, seen that exact silent failure before. That checkbox is way too easy to miss.

But I'm curious about the second alert for compliance. You mentioned layering on time series analysis for a user's baseline. Any specific function or approach you'd start with? Building a baseline seems tricky with sparse data for most users.


Automate everything.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Yeah, the arbitrary number is a real problem. But for someone like me just starting out, what's a practical first step if we don't have the data history or time for a full distribution analysis? Maybe start with a deliberately high number, like 50, just to get the alert framework in place and catch the obvious stuff, then tune it down as we see actual data?



   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

That's a very sensible approach, and I'd absolutely recommend starting exactly that way. Getting the alert pipeline working with a safe, high threshold is a perfect phase one.

The key operational step people miss is to log the threshold itself as a configuration item. When you set it to 50, document the assumption ("assume 90% of users never exceed X downloads") and schedule a review for 30 days out. That turns your temporary guess into a managed, improvable control.

Otherwise, you'll get that first alert, respond to it, and then forget to ever refine the logic, leaving you with a "blunt instrument" alert rule that runs for years.


Architect first, buy later


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Totally agree with documenting the assumption as a config item - that's the step from a one-off script to a real operational process. I'd push it one step further and log that 30-day review task directly into your team's project management or ticketing system right when you set the rule. Don't let it live in a random document no one checks.

My team got caught once because we documented the "why" beautifully in a SharePoint wiki, but the calendar reminder to review got lost in an email thread. The alert ran for 18 months on a threshold that was meaningless after our org restructuring.


Happy testing!


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

You're right that sliding windows add complexity, but complexity is already there and hiding it with a fixed window just gives a false sense of control.

The core issue you point out about different user groups is spot on. That's why my team's "phase one" was to just get alerts flowing with a silly-high global threshold (like 100 downloads), but we tagged every alert with the user's department from our HR system. After a week, we could see that legal's normal was 3 and engineering's was 60. It didn't require a full historical analysis, just a week of watching real traffic with a safety net.



   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

So you're using Sentinel, which is effectively Microsoft's way of making sure your entire security alerting budget stays within their ecosystem. I'm sure the licensing for that connector is straightforward and never changes.

Your static threshold of 20 is the perfect example of creating alert fatigue on day one. It's a number that will either scream constantly at the marketing team for doing their job or be utterly silent when finance decides to archive an entire quarter's reports.

And let's not pretend a 10-minute window is anything but a convenience for the query writer. Anyone with half a brain and a stopwatch can bypass it.


Beware of free tiers


   
ReplyQuote
(@brookel)
Estimable Member
Joined: 2 months ago
Posts: 169
 

Good point on the license lock-in. The Sentinel path definitely ties you into their billing cycle pretty tight once you're invested.

But I'm curious about the bypass method you hinted at. Wouldn't a real, targeted exfil just script the downloads to be slow and steady over days? Avoiding any time window altogether?


Self-host or die trying.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Exactly right. A simple script with a random delay between 1 and 10 minutes would be trivial to write and bypass any fixed window you set. It's a low-and-slow attack, not a smash-and-grab.

That's why focusing on a pure volume threshold is fundamentally flawed. You need to layer in at least one other signal, like destination IP or time of day. An accountant downloading 200 files at 2 AM from a new geo-location tells a different story than the same number during business hours.


Your cloud bill is 30% too high


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

So your solution is to just accept that marketing's "normal" is an unmonitored free-for-all? The problem isn't the threshold, it's the assumption that every department's carelessness should be the baseline.

You don't need a historical analysis. You need the policy that says finance's data requires stricter monitoring than the public marketing folder. If you can't define that, no query syntax will save you.


Read the contract


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Grouping by file type sounds logical until you realize most exfil scripts will just rename extensions or zip everything. You're adding complexity for the illusion of control.

And that 30-45 minute ingestion delay is a billing trap. You pay for Sentinel to ingest those logs in real-time, but you're admitting you can't trust the timestamps you're paying for. How do you justify that cost in a FinOps review when your detection is based on a lag you can't fix?


cost_observer_42


   
ReplyQuote
Page 2 / 3