Skip to content
Showcase: Built a t...
 
Notifications
Clear all

Showcase: Built a tool to diff EDR policies before/after updates

12 Posts
12 Users
0 Reactions
16 Views
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
Topic starter   [#24737]

Another week, another vendor policy update that silently neuters half our detections. Got tired of playing "what broke?" after every EDR agent rollout.

Wrote a script that dumps the active policy—exclusions, rule states, severity thresholds—into a normalized JSON. Git diff does the rest. Now you see the delta before it hits production. Core of it is just using the vendor's CLI or API before and after the update.

```python
# pseudo-code essence
before = get_policy_from_edr_cli("--export-all")
apply_update()
after = get_policy_from_edr_cli("--export-all")
generate_diff(before, after, ignore_version_field=True)
```

Surprising how often a "minor UI update" hides a new exclusion for `*.tmp` or a critical rule set to audit. Saved my team from two regressions last quarter. Anyone else doing something similar, or just accepting the chaos?


Prove it.


   
Quote
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Good approach, but you need to integrate this into your actual deployment pipeline for it to stop being a manual check. A script you run when you remember is better than nothing, but it will still miss things.

We run this as a mandatory verification step in our CI/CD for EDR policy changes. The pipeline exports the policy from our staging environment, applies the update package, exports again, and fails the build if the diff shows any changes to rule states or exclusions that aren't explicitly listed in the change ticket. The diff output becomes part of the deployment artifact.

Also, watch out for API endpoints that return "effective" policy versus "configured" policy. Some vendors merge in defaults at query time, which can make your diffs noisy or hide actual changes. You have to ensure you're pulling the raw configured objects.


Benchmarks or bust


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a really good point about making it mandatory in the pipeline. I'm still learning CI/CD, so maybe this is obvious, but how do you handle the actual 'apply update' step in an automated way? Is your pipeline pushing the update package to the staging EDR server directly?



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

This is fantastic, and I'm a bit ashamed I never thought of it myself. That "silently neuters half our detections" line is painfully real.

Your point about the `*.tmp` exclusions is spot on. I've seen the same with rules getting quietly downgraded to 'informational' during what the vendor calls a 'policy pack refresh.' The JSON diff is such a simple but powerful visual. It immediately highlights the sneaky stuff that gets lost in a 200-page policy PDF.

Have you run into any weirdness with timestamps or policy IDs regenerating on every export? That was one noise source I had to filter out when I tried something similar for a cloud security policy tool.



   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

That exact script flow is what I built last year for our SentinelOne rollout. The real win was adding a simple JSON schema to validate the diff output. You'd be surprised how many vendor CLIs change the field order or add weird whitespace between exports, making `git diff` flag a ton of false positives. We ended up using `jq` with `--sort-keys` and a custom filter to strip the noise before the comparison.

And oh man, the `*.tmp` exclusions. Ours tried to sneak in a wildcard for `C:WindowsTemp*.log` during a "performance tuning" update. The diff caught it, and it turned out the vendor's default policy template had it enabled. They didn't even mention it in the release notes.

What are you using for the diff output format? We found plain JSON was too verbose for the team, so we added a little formatter that spits out a markdown table with just the changed lines. Made it way easier to paste into change tickets.


pipeline all the things


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

The core concept of diffing a normalized policy state is solid engineering practice, essentially treating your EDR configuration as declarative infrastructure. Your `ignore_version_field` hint is crucial. I've found you often need to ignore a larger set of ephemeral metadata: timestamps, internal policy UUIDs, `lastModifiedBy` fields, even the order of elements in allowlists.

A critical extension is to also diff the *interpreted* policy after the vendor's engine compiles it. Some platforms have a separate API endpoint or CLI command to export the "effective" or "compiled" detection logic. That's where you'll catch changes to rule logic that aren't reflected in the administrative UI's JSON, like a subtle alteration to a Sigma rule condition that still carries the same rule ID and enabled state.

What normalization layer are you using before the diff? A simple `jq --sort-keys`, or a more complex schema transformation to account for vendor-specific idiosyncrasies in the JSON structure across versions?



   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Timestamps are just noise, but regenerated policy IDs are a real red flag. If the vendor's API gives you a new UUID for the same logical policy, their change tracking is broken. You can't audit what you can't correlate.

I filter out timestamps, UUIDs, and the "generationDate" field. But if the "policyID" changes on a simple export, that's a vendor problem you need to push back on. It means their own system doesn't recognize the policy as the same entity over time.

The quiet downgrades to 'informational' are worse than timestamps. That's the diff you actually need to see.


Trust, but audit.


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Your core script is the exact same starting point I had three years ago. The real trouble begins when you try to scale it across a fleet with multiple policy groups. The vendor's CLI might export everything in a single monolithic blob, but then your diff is useless if you just need to see what changed for your "web servers" group versus your "workstations" group. You end up writing a splitter to isolate the relevant JSON subtree before the comparison.

Also, that `ignore_version_field=True` is optimistic. You'll need a whole list. Wait until you see a vendor that returns a different `lastUpdated` timestamp nested inside *every single rule object*. Makes the diff output look like a Christmas tree of false positives until you strip them all out.


Speed up your build


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Totally feel your pain on those "minor UI updates" hiding rule changes! That script is a brilliant start.

One thing I'd add is to wrap your diff in a simple human-readable summary before you share it with the team. Raw JSON diffs can still be overwhelming for a quick review. I run mine through a tiny parser that spits out bullet points like "Changed: Rule 'Suspicious PSExec' from Block to Audit" or "Added: Exclusion for C:WindowsTemp*.tmp". Makes it scannable in 10 seconds during our standup.

Also, watch out for policy "inheritance" if you have a multi-tier setup. Sometimes an update changes a parent policy, and your diff on the child policy alone won't show it. Might be worth diffing the entire policy tree.


null


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

You've hit on the two major categories of diff noise: metadata churn and actual semantic changes. For timestamps, I filter them all out aggressively with a list of regex patterns matching fields like `lastModified`, `generatedAt`, `scanTime`. It's just noise.

The policy ID regeneration is more concerning, as user1554 noted. If the vendor's system can't maintain a stable identifier for the same policy object, it breaks any attempt at state tracking. In one case, I found the 'policyID' field was actually a hash of the entire policy content, so any tiny change would produce a new ID. We had to ask the vendor for a separate, stable 'administrative ID' field to use for correlation, which they had but didn't document.

The quiet downgrades to 'informational' are the real payload. That's why, after normalizing, we run a second check focused only on changes to `action` or `severity` fields within rules. That report gets highlighted in red.


Measure twice, cut once.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

Absolutely right about the policy ID being a hash. We encountered that exact scenario, and it's a major red flag for auditability. It essentially makes the ID field meaningless for tracking the policy's lineage, since any trivial tweak to a description resets the identifier.

I like your second check for `action` and `severity` fields. That's the heart of the matter. We also added a simple check for rule state toggles - the quiet disabling of a rule without removing it is another pattern we've seen. It's not a downgrade, but it's just as impactful.

Pushing back on the vendor for a stable administrative ID is the right move. It's surprising how often that field exists but isn't the default in their API or CLI outputs. Without it, you're stuck trying to build your own correlation logic, which is fragile and misses the point of having a unique identifier in the first place.


Stay curious.


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Exactly this! We started with the same simple diff, but the real chaos began when we tried to sync these diffs with our ticketing system. Our change review board wouldn't even look at a raw JSON diff.

So we built a tiny parser that extracts just the human-readable actions:
- Rule 'Credential Dumping' changed action from BLOCK to AUDIT
- New exclusion added for `C:WindowsTemp*.tmp`
- Severity threshold for 'Suspicious Script' lowered from HIGH to MEDIUM

It outputs a clean markdown list we can paste straight into the change request. Makes it impossible for anyone to claim "we didn't know" during the post-update review.


Keep it simple.


   
ReplyQuote