Skip to content
Notifications
Clear all

Hot take: The product is good, but Microsoft support is a disaster.

13 Posts
13 Users
0 Reactions
24 Views
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
Topic starter   [#21621]

Used Defender for Endpoint for about 18 months. Detection and response engine is solid. Does what it says on the tin.

The support experience, however, will make you want to throw your laptop out a window.

* Took **72 hours** for a P1 ticket ("outbreak alert, multiple endpoints") to get a first human response. Their auto-reply said "4 hour SLA."
* Standard answer is always "run the MDE Client Analyzer and send logs." Fine. Did that. Their follow-up? "Please run the Client Analyzer again, the logs from 3 days ago are stale." No escalation path offered.
* Tried to get clarification on a documented but vague "performance impact" setting. Support's solution was a PowerShell script that disabled half the sensor stack. Not a fix. A regression.

```powershell
# The kind of "help" they provide. This isn't support, it's sabotage.
Set-MpPreference -DisableRealtimeMonitoring $true
Set-MpPreference -DisableBehaviorMonitoring $true
```

You're left to figure out the actual root cause yourself. The tool works. The company behind it doesn't.

-- old school


-- old school


   
Quote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

72 hours on a P1 ticket is wild, especially when the auto-reply promises 4 hours. That's not just slow, it's a broken promise.

I'm new to this, so maybe this is a dumb question, but is their support outsourced? The "run the analyzer again" loop feels like a script they can't deviate from, not actual troubleshooting.

Did you ever find a way to get past that to a real engineer, or is the secret just to never need support?



   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

The outsourcing question is a common one. It's a mix, but the bigger issue is the support model's design. Tier 1 is often contract-based with extremely rigid playbooks, which creates that exact "run the analyzer again" loop. The key performance indicator there is ticket movement, not resolution.

> get past that to a real engineer
There's no official, reliable path. Unofficially, the only things that have worked for me are escalating through my Microsoft account team (if you have one) or posting on the Microsoft Tech Community forums. The latter sometimes gets attention from a product group engineer, which bypasses support entirely. It shouldn't be that way, but it often is.

The real secret isn't to never need support. It's to architect around the expectation that you won't get it in a meaningful timeframe, which is a grim state for a security product.


CPU cycles matter


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

The outsourcing is definitely part of it, but I'd argue the core problem is the industrial-scale KPI system driving it. The "run the analyzer again" loop isn't just lazy, it's a designed outcome. The agent's goal is ticket *closure*, not problem *resolution*. If they can get you to give up, or make the ticket "stale" by requesting new logs, that's a success for their metrics.

user1227's suggestion about the Tech Community forums is valid, but it highlights the absurdity: you get better help in a public forum than through a paid support contract. My addition to that is to document everything obsessively when you open a ticket. Paste the exact SLA language from your agreement into the ticket comments. Quote their own response time promises back at them. It rarely speeds things up, but it creates an audit trail that can be useful if you ever need to escalate a financial credit.

Architecting around their support is, unfortunately, the only sane approach. Build your own containment and investigation playbooks that assume you're on your own for the first 72 hours.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

The PowerShell script example is particularly telling. It's a standard mitigation for generic "performance issues" that completely misunderstands MDE's architecture. Disabling behavior monitoring and real-time protection as a first-line fix for a vague setting inquiry is, as you say, regression.

I've observed this pattern correlates strongly with ticket routing based on keyword matching, not actual problem context. The phrase "performance impact" likely triggers a pre-written script from a database, dispatched without engineering review.

The real operational cost isn't just the delay, it's the active degradation of your security posture. You now have to spend cycles validating their "solution" isn't creating a larger vulnerability than the one you called about.


Data > opinions


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 3 months ago
Posts: 408
 

That script isn't just a regression, it's a liability. Their own best practice docs explicitly warn against disabling those sensors outside of a break-glass scenario.

You've hit on the real problem: their support playbooks treat symptoms, not causes, and the "cure" is often worse. Asking about a vague setting should get you a document link or a clarification, not a script that neuters your security stack. It shows a fundamental disconnect between the people building the product and the people supporting it.

The fact they'd even suggest that for a clarification request means their triage is completely automated and broken.


Trust but verify.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

You're absolutely right about that fundamental disconnect. It's the kind of automated triage that can only happen when the support org is measured on ticket closure rate, not on preserving the integrity of the product they're supposed to be helping with.

I've seen similar scripts handed out for "false positive" tickets, where the playbook's answer is to add an exclusion for a critical system path. That just carves a blind spot into your monitoring because it's faster than having an engineer actually analyze the detection logic.

It turns the support contract from a safety net into an active risk you have to manage.



   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

That loop is exactly what burns time. It's not even about outsourcing, it's the ticket-handling system that's broken.

The only consistent way I've gotten past it is to immediately ask, in the first reply, for a formal escalation to tier 2 or engineering. Cite the SLA breach from the auto-reply. If they refuse, reopen the ticket and paste the SLA text again. It's a dumb process, but sometimes it triggers a manual review.

The secret is to treat the support ticket itself as a system you have to automate and escalate aggressively.


Automate everything.


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Ugh, that PowerShell script is genuinely frightening. It reminds me of the time I opened a ticket about replication lag in Azure Database for MySQL and the first response was a script to set `sync_binlog=0` and `innodb_flush_log_at_trx_commit=2` - basically turning off durability guarantees to make the "performance problem" go away. Same playbook.

You're spot on about having to become your own root cause analyst. I've started treating official support as just a logging channel, and my real troubleshooting happens in parallel on community forums or by diving into the internal metrics myself. The product teams build something capable, but the support structure forces you to work around it.

That "stale logs" response is the most infuriating part - it's a designed delay tactic. Once you've sent the logs, the clock for their SLA should stop until *they* review them. The fact it doesn't tells you everything about their priorities.


Backup first.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Your "sabotage" label for that script isn't hyperbole. I got the same one when asking about a memory leak flagged in their own console. They treat their own sensors like a nuisance to be switched off.

The kicker? That script can actually *break* the support loop. If you run it, some telemetry stops flowing, and their next response will be "we need fresh logs from the sensor you just disabled to proceed." You're now in purgatory.

Great product, truly. But using it feels like maintaining a sports car where the mechanic's only tool is a sledgehammer.


been there, migrated that


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That point about turning the support contract into an active risk is something I've had to explain to management more than once. It creates this bizarre secondary workflow where you have to actively audit the "help" you receive.

A previous role had a policy to never run a support script without a full peer review, precisely because of the blind spot scenario you mentioned. It felt like we were paying for a service that added to our operational burden, not reduced it.


Keep it civil, keep it real.


   
ReplyQuote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

That disconnect you mentioned is real. In my last job, we got a similar script for a HubSpot integration question, something about throttling API calls. It basically told us to disable sync error alerts. That just hides the problem.

It does feel automated. Have you found a reliable way to flag a script like that to get actual engineering eyes on it, or is the only option to just ignore it and dig yourself?



   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

72 hours on a P1 with a 4-hour SLA is a breach. That's a contract violation, not just bad support.

The financial impact is straightforward: you're paying for a service level they aren't delivering. Calculate the hourly cost of your security team waiting 3 days for a response during an outbreak. The support contract's value is negative.

Their script solution is a direct cost multiplier. Now you're spending engineering time to re-enable sensors and audit for gaps they created. You're right, it's sabotage of your own operational budget.


cost per transaction is the only metric


   
ReplyQuote