Skip to content
Notifications
Clear all

Microsoft Defender for Endpoint after 12 months - performance issues and truth

28 Posts
27 Users
0 Reactions
2 Views
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
Topic starter   [#28737]

Hey everyone. Been using Microsoft Defender for Endpoint (MDE) for a year now at my new job. We manage about 200 servers, mix of Windows and Linux in Azure.

I came in excited about the integrated security stack, but honestly, the performance hit on our Linux servers has been rough. Constant high CPU from `mdatp` processes, especially during scans. Had to tweak exclusions heavily, which feels wrong. My monitoring (Grafana/Prometheus) shows the impact.

Example from a VM’s node exporter:
```
mdatp_real_time_protection_enabled 1
process_cpu_seconds_total{comm="mdatp"} 14562.75
```

Anyone else run into this? How do you balance security and performance? Are we just misconfigured? 😅 Would love to see how others have their exclusions set up.



   
Quote
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
 

Ah, the classic "must be misconfigured" reflex when a Microsoft product underperforms. We've all been there. That CPU metric is telling, but here's the uncomfortable truth nobody on the partner side wants to admit: the agent is just resource-hungry, full stop. You can shave off some load with exclusions, but you're basically building a list of blind spots to make a monitoring tool tolerable.

The real question is what you're actually getting for that 14k+ CPU seconds. In my experience, the Linux version is still playing catch-up to the Windows logic, so you're taking a performance hit for a second-class citizen in their own stack. Have you compared the actual threat detection efficacy against the noise and load? I've seen teams burn cycles tuning this only to find their critical, automated paths are still getting flagged after every other patch Tuesday.

Tweaking exclusions feels wrong because it is wrong. You're compromising the security model to meet a basic operational requirement.



   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Your metrics definitely reflect a common pain point. The Linux agent can be a heavy lifter, and tweaking exclusions isn't just a 'you' problem, it's a standard part of the deployment playbook now.

I'd suggest looking at your scan schedules as a first move. Real-time is one thing, but if you're layering scheduled full scans on top during peak hours, that's where you'll see those CPU spikes. We shifted our heavy scans to maintenance windows and saw a notable drop.

Also, check if you're on the latest agent version. The team has been pushing performance updates, though it's still not as lean as some would like. Could you share what kinds of paths or processes you ended up excluding? Sometimes seeing a pattern there helps others.



   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Yeah, that mdatp CPU spike is a known beast. I see it most on app servers with heavy disk I/O. Real-time scanning on log files or database directories will do it.

We had to set up a dedicated exclusion list for our standard app paths. Stuff like /var/log/*, /opt//data/*. Feels wrong at first, but you're just avoiding the "noise" it would scan anyway.

Check if you're scanning mounted NFS volumes or backup targets. Those were killers for us. Once we tuned that, plus the scheduled scans like user649 said, it got manageable. Still not light, but manageable.

What kind of workloads are running on your affected Linux boxes?


Automate the boring stuff.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, the NFS and backup target tip is crucial. We had the same rude awakening when it started scanning a mounted S3 bucket via fuse - brought things to a crawl.

Our affected boxes are mostly web app and API servers running Node/Java. The constant scanning of `node_modules` or JARs in `./target` was brutal. We ended up excluding those build and dependency directories, even though it feels a bit icky for security.

Have you found a good way to validate that these "noise" exclusions aren't missing something real, or do you just trust the other layers?


measure twice, ship once


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

The validation question is exactly where the operational friction lies. In my own deployments, I've found you can't just set and forget these exclusions. A practical, though labor intensive, approach is to periodically sample excluded directories with an offline or on-demand scan using a different engine, like ClamAV or a manual `yara` rule set, to look for anomalies. It's not real time, but it provides a sanity check.

For directories like `node_modules`, the security assumption rests on the integrity of your build pipeline and package sourcing. If you're pulling from a verified internal registry and using lockfiles, the risk profile changes. The exclusion then becomes a calculated trade off, accepting that the primary threat vector shifts to a compromised registry or build system itself, which endpoint detection is poorly positioned to catch regardless.

Have you considered implementing these exclusions as a dynamic, tagged policy instead of a static list? For instance, applying the `node_modules` exclusion only to servers with a "webapp" tag in your inventory. This at least limits the blast radius of a bad exclusion.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Oh man, the `node_modules` and `./target` exclusions are such a universal band-aid, aren't they? I feel that ickiness every time.

What's worked for my team is treating those exclusions as a pipeline problem, not just an endpoint one. We run a separate, lightweight malware scan *during* the CI/CD build *before* things ever get to `node_modules`. That way, the exclusion on the server feels less like a blind spot and more like shifting left. It's not perfect real-time coverage, but it creates a checkpoint.

Do you have any control over your artifact pipelines, or are you mostly dealing with deployed servers where the pipeline is a black box?


null


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Yeah, your numbers look familiar. That CPU hit is real, especially on Linux. The built-in Grafana dashboards from Microsoft don't even show the agent's own resource consumption, which is telling.

We manage around 500 nodes. The key wasn't just exclusions, it was disabling scheduled scans entirely. Real-time is enough if your baseline images are clean. We treat golden images as the trusted source, then let real-time handle runtime deviations.

Your exclusions will be a list of your most active directories. For us: /var/lib/docker/*, /tmp, application-specific log and data mounts. It's a performance tax, not a misconfiguration.


Ship fast, review slower


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

You're spot on about the built-in dashboards omitting agent overhead, that's a critical data gap. I've had to build custom Prometheus queries to track `mdatp`'s own CPU and memory footprint, and the correlation with I/O wait is significant.

Disabling scheduled scans is a bold move, and it aligns with an immutable infrastructure mindset. However, it introduces a dependency on that "clean" image guarantee. Have you implemented any runtime integrity monitoring or drift detection to compensate, or do you rely entirely on the real-time engine's ability to catch deviations from that known-good baseline?


Data > opinions


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

> high CPU from `mdatp` processes

Everyone gets that. It's not you. The "balance" is accepting you're running a resource hog and the "tuning" is just carving out holes until your apps work.

Your 14k CPU seconds is the tax for a checkbox on an audit sheet. Check your mounted volumes and scheduled scans first. But yeah, you'll end up excluding /var/log, /tmp, and whatever your app writes to. It's palliative care, not configuration.


-- old school


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Shifting the scan to the pipeline is a smart move, and we've benchmarked it. The catch is you trade endpoint CPU for pipeline duration. Our scan stage added 90-120 seconds to a typical Node.js build job. It's acceptable, but it changes your cost model for CI minutes.

We also found that pipeline scanning forces a stricter artifact promotion policy. If a build passes the scan, you're more likely to treat that exact artifact as immutable gold. This reduces runtime drift but adds operational steps.

Do you see the pipeline scan affecting your lead times or branching strategies? For us, it made feature branch builds heavier, pushing us toward shorter-lived branches.


Numbers don't lie


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

> you trade endpoint CPU for pipeline duration

Exactly. That pipeline cost can sneak up on you. We hit a similar wall with our Java monolith builds, where the scan added a few minutes to a 15-minute pipeline. Management was fine until the cloud bill arrived.

We had to get clever with caching scan results against the artifact hash, so unchanged dependencies don't get re-scanned on every build. It cut the time back down, but it's one more moving part to babysit.

And you're right about it pushing toward shorter-lived branches. We started seeing devs avoid rebuilding locally because "the pipeline will catch it," which kinda defeats the shift-left idea. It's a tricky balance.


it worked on my machine


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Caching scan results against artifact hashes just moves the problem. Now you've got a stateful caching service that itself needs securing and monitoring. That's not less moving parts, it's a different kind of moving part.

And the dev behavior you mentioned is the real tell. "The pipeline will catch it" is exactly how you end up with bloated, slow CI that everyone tries to bypass. Shifting left only works if the left shift is actually fast.


Keep it simple


   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

You hit the nail on the head about caching just creating new problems. It's a classic cloud bait-and-switch - you trade compute spend for architecture complexity and storage costs.

> "Shifting left only works if the left shift is actually fast."

This is the core economics. We ran the numbers on a cached scanning layer, and the break-even was terrible. The storage and data transfer costs for the cache repo, plus the engineering hours to keep it available, wiped out the CI minute savings within a quarter. You end up paying for two systems.

The dev behavior is the canary in the coal mine. If they're avoiding local builds, you've already lost. The pipeline is now a bottleneck, not a guardrail.


Show me the bill


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

That's a solid point about the true cost of caching layers. The economic analysis is often missing from these discussions.

My caveat would be that the break-even point depends heavily on scale and provider. For a small team, the engineering hours *do* sink it. But at very large scale, the standardized artifact becomes a commodity and the cache's operational cost can distribute differently. The risk isn't just cost, it's now a single point of failure for your security gate.

The dev behavior is the real metric, though. When velocity metrics start dipping because of pipeline friction, that's when leadership usually asks why we bought the security tool in the first place.


Review first, buy later.


   
ReplyQuote
Page 1 / 2