I've been working on tightening up our internal security monitoring, and one area that always felt a bit opaque was user behavior on our SharePoint document libraries. Specifically, spotting anomalous bulk downloads by a single userβpotential data exfiltration or just a user being overzealous.
Microsoft Sentinel is great for this, but setting up a precise alert took a bit of tinkering. Here's a streamlined method I landed on that focuses on high-volume download events within a short time window.
**Core Logic & Prerequisites:**
First, ensure your SharePoint Online logs are flowing into Sentinel via the Office 365 connector. The key table is `OfficeActivity`. We're looking for `Operation` equals `FileDownloaded` and the `SourceFileName` isn't empty.
The KQL query below looks for users who have triggered more than a threshold of download events from a specific site within a 10-minute window. You'll need to adjust two variables:
* `download_threshold`: I started with 20 files as a baseline.
* `workspace_name`: Your Sentinel workspace.
**The KQL Query:**
```kusto
let download_threshold = 20;
let timeframe = 10m;
OfficeActivity
| where TimeGenerated >= ago(timeframe)
| where OfficeWorkload == "SharePoint"
| where Operation == "FileDownloaded"
| where isnotempty(SourceFileName)
| summarize StartTime = min(TimeGenerated), EndTime = max(TimeGenerated), DownloadedFiles = count(), FileList = make_set(SourceFileName) by UserId, SiteUrl, UserAgent
| where DownloadedFiles > download_threshold
| extend CustomEntity = pack("UserId", UserId, "SiteUrl", SiteUrl, "Count", DownloadedFiles)
| project StartTime, EndTime, UserId, SiteUrl, DownloadedFiles, UserAgent, FileList, CustomEntity
```
**Next Steps After the Analytics Rule:**
Once you save this as a scheduled analytics rule (every 5-10 minutes works), you'll want to:
* Create a clear incident classification in the rule settings.
* Consider adding an exclusion for known administrative or backup accounts.
* Integrate the alert with your ticketing system or Teams channel for immediate visibility.
The real power comes from tuning the threshold based on your normal department activity and pairing this alert with user context from your Azure AD logs. Have any of you implemented similar detection for other O365 workloads like OneDrive or Teams file sharing? I'm curious about your thresholds and if you've tied this into a user risk scoring model.
Twenty files in ten minutes isn't anomalous, it's a Tuesday for anyone prepping for a client meeting. Your threshold is going to flood the SOC with false positives from legitimate power users.
The real blind spot is distinguishing between *different* files and someone just hammering F5 on the same document because their sync client is having a fit. You need to add `SourceRelativeUrl` or `SourceFileExtension` to the grouping, otherwise you're just counting retries on a single item. Sentinel's great, but this logic treats all downloads as equal, which they aren't.
Also, you didn't mention the ingestion latency. OfficeActivity logs can take 30 minutes to appear. Your "real-time" alert for a ten-minute window might fire for activity that's already an hour old.
monoliths are not evil
Valid point on the latency. You have to build your detection window around that 30-45 minute delay, or the alert is useless. Basing it on the ingestion time, not the event time, is a more practical fix.
Grouping by SourceRelativeUrl is essential, but don't forget about file type. Someone downloading 50 PDFs from a contracts library is normal. Someone downloading 50 .vhd or .pst files is not. Your query needs both uniqueness and some basic file classification.
They're right about the grouping key, but focusing on URL or extension still treats every library the same. Legal downloading 20 contracts is normal. R&D downloading 20 source files might be a fireable offense.
Your rule needs a baseline per site or library. Build an allowlist of noisy libraries first, then monitor the sensitive ones.
Least privilege is not a suggestion.
Twenty files is a comically low bar. You're going to drown in alerts. Forget a static threshold, you need a dynamic one based on that user's history. Someone who downloads two files a week hitting twenty is a red flag. The legal team doing it is Tuesday.
And you're missing the obvious signal: time of day. A flurry of downloads at 4:55 PM on a Friday is a whole different story than Tuesday at 10 AM. Your query treats them the same.
CRM is a necessary evil
I appreciate the clean KQL start, but that download_threshold variable is a trap. Static numbers ignore departmental norms - what's anomalous for engineering isn't for marketing. You'll drown in false positives unless you build a per-group baseline, which means pulling in HR data or at least site membership. And while you're at it, check the Office 365 connector's latency; your 10m window might be useless if logs arrive in chunks.
APIs are not magic.
HR data for baselines is a pipedream in most shops. You're better off grouping by SiteUrl and letting a simple anomaly detection algorithm chew on the logs for a week. At least it'll learn that the finance sharepoint is always noisy.
And you're right about latency. If you're not using `ingestion_time()` in your window, you're alerting on ghosts.
Prove it.
Great point about starting with a static threshold, but 20 files in 10 minutes is way too low for many teams. I've seen our marketing folks hit that before their first coffee.
You're absolutely on the right track looking at `FileDownloaded` events, but that threshold variable needs context. You'll want to filter out known noisy sites first, maybe by adding a | where SiteUrl !contains "PublicWebsite" clause. Otherwise, you're just building an alert factory.
Also, watch out for that 10-minute window with ingestion latency. Using `ingestion_time()` instead of `TimeGenerated` might save your sanity.
K8s enthusiast
Really solid starting point, thanks for sharing this. Setting that first baseline is the hardest part.
That `download_threshold = 20` is a perfect example of the "tinker trap" though. You'll set it, get flooded with alerts from the marketing or design site, adjust it to 50, and then miss something real happening in the HR site. You almost need to run the query in a "monitor-only" mode for a week across different SiteUrl segments to see the natural variance before you pick a real threshold.
Also, good catch on requiring `SourceFileName` not to be empty. I've seen a bunch of weird ghost events in those logs that can throw off the count.
don't spam bro
Cool trick with the SourceFileName filter, I would've missed that. But when you set `download_threshold = 20`, is that per site or across everything? If it's global, wouldn't a user hitting 20 files across a dozen different sites still trigger it, even if each site is under the limit?
Trying to figure it out.
That's a really sharp observation - you're absolutely right. The way the query's written, `| summarize` groups by `UserPrincipalName`, so it's counting downloads *across all sites* for that user. A user grabbing 3 files from 7 different project sites would hit the 20 threshold and trigger, even though their activity on any single site looks normal.
This gets to the core tension in alert design: do you care about aggregate user behavior, or behavior within a specific security boundary (like a site)? For intellectual property protection, the site boundary is often more meaningful. You'd want to group by both `UserPrincipalName` and `SiteUrl` to catch someone raiding a single sensitive library.
So you could adjust the summarize clause to something like:
```
| summarize FileCount = count() by UserPrincipalName, SiteUrl
| where FileCount > download_threshold
```
That way, you're measuring activity per-user-per-site, which aligns better with how permissions and data sensitivity are usually structured.
Prod is the only environment that matters.
That's a great starting point for the query. Grouping by both `UserPrincipalName` and `SiteUrl` is definitely the way to go for spotting someone targeting a single library.
But since you're already adjusting variables, you could add one more: a list of SiteUrls to exclude. We have a few "all-hands" document libraries where bulk downloads are just part of operations, and filtering those out right at the start kept our alert volume sane while we tuned the real threshold. Just a simple `| where SiteUrl !in~ ()` clause after your time filter.
Also, what are your thoughts on adding a filter for file type? A spike in downloading .zip or .pst files tells a different story than a spike in .jpg files from the marketing site.
spreadsheet ninja
Great starting point! You're spot-on about filtering out empty `SourceFileName` events - those ghost entries can really throw off your counts.
But I'd bump that `download_threshold = 20` way up for an initial test, maybe to 50 or 75. You'll still catch the real outliers while avoiding alert fatigue from those teams who treat SharePoint like a buffet 😅.
Also, consider adding a quick filter for known high-traffic sites at the beginning, like `| where SiteUrl !contains "marketing-share"`. That'll give you a cleaner signal while you tune the threshold.
Keep deploying!
You start with the query? That's putting the cart before the horse. First, check if your connector is even logging FileDownloaded events consistently. Mine wasn't for weeks, and Sentinel just showed empty. KQL is pointless without data.
If it ain't broke, don't 'upgrade' it.
Agreed, but before anyone even touches the KQL, they need to verify the connector's scope. The Office 365 data connector in Sentinel must be configured to collect SharePoint activity specifically. It's not on by default in some tenants; you have to explicitly check the box for SharePoint in the connector's settings. I've walked into three engagements where the team had perfect Exchange logs but zero `FileDownloaded` events because of this oversight.
Your static threshold is a solid operational starting point, but it's phase one. For anything resembling compliance, you'll need to layer on a second alert using time series analysis on the same dataset to spot deviations from a user's own baseline, not just a global number. The static rule catches the obvious smash-and-grab, but the anomaly detection finds the slow bleed.
Mike