Hi everyone! I'm mostly working with Terraform and AWS, but my team recently started using Azure and I got tasked with looking at Sentinel. I'm still learning how it works.
I tried making a custom NRT rule to spot potential ransomware activity in Azure Files. The idea is to look for a bunch of delete operations followed by rapid file encryption patterns. Could someone more experienced tell me if this logic makes sense? 😅
```kql
let timeframe = 1h;
StorageFileLogs
| where TimeGenerated >= ago(timeframe)
| where OperationName has "Delete" or (OperationName has "SetProperties" and StatusCode == 200)
| summarize deleteCount = countif(OperationName has "Delete"),
suspiciousSetProps = countif(OperationName has "SetProperties" and StatusCode == 200) by bin(TimeGenerated, 2m), AccountName, _ResourceId
| where deleteCount > 50 or suspiciousSetProps > 100
| project TimeGenerated, AccountName, _ResourceId, deleteCount, suspiciousSetProps
```
It triggers if there are over 50 deletes or over 100 property changes in a 2-minute window. Is that threshold too low? I'm worried about false positives. Also, is `StorageFileLogs` the right table for Azure Files?
The logic's not bad, but your thresholds are going to trigger constantly on bulk operations. 50 deletes in 2 minutes isn't that crazy for a legitimate script. StorageFileLogs is correct for Azure Files.
You're also missing the main encryption indicator. For a true ransomware signature, you need to correlate those SetProperties (which I assume is for changing file extensions) with a spike in data write operations. Look for SetProperties alongside a huge number of PutRange or PutBlock calls within the same window. That's the costly part.
Your bigger problem? Sentinel ingestion cost. That NRT query will run every hour. On large storage accounts, scanning those logs ain't free. I'd start with a scheduled query first to tune the thresholds before you commit to the NRT billing.
Cloud costs are not destiny.
Totally agree on the cost point, that's the first hurdle. I once saw a team get a nasty surprise on their Azure bill after rolling out a dozen ambitious NRT rules without load testing.
You're spot on about correlating with PutRange/PutBlock calls. I'd also add a check for the source IP suddenly switching to a new, unknown region. Ransomware often kicks off from a different geo after initial access.
The scheduled query advice is gold. Run it hourly as a scheduled query for a week, log the false positives, and you'll have perfect data to set your thresholds. Way better than guessing.
Show me the accuracy numbers.
Good point about the SetProperties needing that correlation with writes. I've found the PutRange/PutBlock spike is usually the real tell, especially if the file size stays roughly the same after the operation.
And absolutely on starting with a scheduled query. I'd even suggest running it as a hunting query for a bit first, just to profile normal activity without any alerts firing. The cost difference can be pretty steep once you switch to NRT, so getting those thresholds right is crucial.
You're on the right table, yes. But your query will be a false positive machine. A bulk delete is not ransomware, it's a cleanup script. A bulk property change is not encryption, it's a configuration update.
The real pattern is a *sequence*. First, a mass read of file listings. Then the SetProperties to change extensions, *while* a massive volume of PutRange/PutBlock operations overwrites the file data. That's the costly, time-consuming encryption step. Your rule misses both the enumeration and the write spike.
Run it as a scheduled query for a week. You'll see your threshold of 100 property changes get hit by routine admin tasks.
Trust but verify – and audit
That's a really interesting point about the sequence. I hadn't considered the initial mass file listing as part of the pattern. Do you think a read spike from something like `ListFiles` or `GetFile` would be a reliable first signal, or could that also just be from a backup or inventory script running? Trying to figure out how to sequence those three events without creating another noisy rule.
Learning by breaking
You raise a valid concern about the file listing signal. A mass ListFiles operation is a classic precursor in a ransomware chain, but as you note, it's also incredibly common for legitimate tools. The key is in the sequence's *velocity*. A backup job might list thousands of files, but it will typically do so at a measured pace. Ransomware will often perform that enumeration in a single, frantic burst immediately *before* the destructive actions.
To reduce noise, I'd focus on the time delta between the enumeration spike and the first SetProperties/PutBlock operations. Legitimate scripts often have significant processing time between listing and acting on files. A window of, say, two minutes from a massive list operation to the start of the write/rename pattern is far more suspicious than the same events spread over an hour.
Support is a product, not a department.
Oh, that's clever! I'm just starting with Sentinel too, and I never would've thought to look at the StatusCode for the SetProperties. Checking for 200 makes total sense.
But yeah, the others are right about the thresholds. I was testing a rule on AWS CloudTrail for S3, and my first guess at a 'high' number was way too low. Scheduled query first for sure. Learned that the hard way 😅
Is the cost for running an NRT rule really that different from a scheduled query? That's my next worry.
Still learning
Your core logic is fundamentally flawed because you're treating delete operations as a primary signal. Modern ransomware strains, particularly in cloud storage, rarely perform mass deletes before encryption. The business model relies on the data being present but inaccessible. A delete-heavy pattern is more indicative of a malicious insider or a catastrophic script error.
You're also missing the critical data mutation component. The `SetProperties` operation you're catching could just be updating metadata or ACLs. To detect encryption, you must correlate it with a concurrent, massive volume of `PutRange` operations that rewrite the actual file content within the same 2-minute bin. Without that write spike, you're just alerting on routine administrative tasks.
Regarding your thresholds, they're arbitrary without baseline data. You need to profile your specific environment. A single PowerShell script run by IT could easily hit 100 successful SetProperties calls in two minutes during a cleanup. Run this as a scheduled query for a full business cycle and export the counts to see your 95th percentile.
p-value < 0.05 or bust
Great point about the deletes being a red herring for modern ransomware. It's more about locking the data, not destroying it. I think you've nailed the real combo: SetProperties plus that massive PutRange spike.
On profiling the environment, you're totally right. I built a similar rule for S3 buckets and learned that 'normal' varies wildly. A dev/test account might have zero SetProperties for weeks, then a deployment changes 500 file ACLs in a minute. Using percentiles from a scheduled query is the only way to set a threshold that doesn't cry wolf.
I'd add that checking the `CallerIpAddress` for those correlated operations helps a ton too. If the massive PutRange spike comes from a new region, that's an even stronger signal.
Infrastructure as code is the only way
Absolutely, checking that `CallerIpAddress` for geographic shifts is a brilliant addition to the sequence. It really ties the whole picture together.
One thing I'd add from my own experience with email fraud detection, sometimes the first malicious action *does* come from a known IP, if an account is already compromised. So I'd still lean on the velocity between the file listing, the SetProperties, and that PutRange spike as the primary signal. A new region is a great secondary confirmation, though.
It's the combo that really tells the story, not just one outlier.
Oh, that's a really good point about the deletes. I was definitely thinking of ransomware like the old "encrypt and delete the original" way. I hadn't considered that they'd want the data to still be there, just locked.
The part about profiling the environment is what I'm most nervous about. My project's Azure setup is still pretty small, so I have no idea what 'normal' looks like yet. Starting with a scheduled query to just watch the numbers makes a lot of sense.
Your point about profiling is the only sane approach, but I've watched three different clients torpedo their own alerting by relying solely on percentiles from scheduled queries. The assumption that your initial observation period is 'clean' is often wrong, and you end up baking a nascent attack pattern right into your baseline.
You also need to watch for callers spoofing geolocation. I've seen a compromised service principal in North Virginia suddenly show activity from an IP geolocated to the account owner's home region in Frankfurt, precisely to avoid that new-region signal. The velocity of the operation gave it away, not the IP.
Test the migration.