Skip to content
Notifications
Clear all

Our results after enforcing device posture: 12% of 'managed' laptops actually failed.

14 Posts
14 Users
0 Reactions
2 Views
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
Topic starter   [#29112]

We just finished a device posture rollout for our ZTNA access gates. The requirement was simple: only managed, encrypted, and patched (>30 days) corporate laptops get access to internal apps.

Ran the check against our entire "managed" fleet of 5,000+ macOS/Windows devices. Netskope Client pulled data from our MDM (Jamf/Intune) and local checks.

**Results:**
* 12% (600+) of devices were flagged as non-compliant.
* Top failure reasons:
* Disk encryption disabled (user-initiated).
* MDM agent unexpectedly missing/unresponsive.
* OS build version not reported (often due to pending reboot).
* Local admin changes (unauthorized software, disabled security services).

**Takeaway:**
Your "managed" inventory is a lie until you enforce continuous, pre-authentication checks. We blocked all non-compliant devices at the ZTNA gateway. The help desk queue exploded, but it exposed real security debt.

**Config snippet for the posture check (simplified):**
```json
{
"rule_name": "corp-laptop-compliance",
"conditions": {
"os": ["windows", "macos"],
"encryption": "required",
"mdm_status": "enrolled",
"max_patch_age_days": 30
},
"action": "allow"
}
```

The noise was painful but necessary. Now we have a real-time map of device trust. Next phase: integrating with our EDR for active threat checks.


Trust but verify, then don't trust.


   
Quote
(@hannahr)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Spot on. That 12% figure is a gut check so many teams miss. We saw nearly identical failure reasons during our ERP migration's pre-flight checks.

Your point about the help desk queue is key. It's not just a spike in tickets, it's a real-time map of your weakest support touchpoints. We found the "MDM agent unresponsive" category often traced back to a single outdated onboarding image the desktop team was still using.

Did you see any pattern linking certain departments to specific failure types? We found our sales engineers were the worst for disabling encryption for performance, while devs had the most "pending reboot" states.


Data is sacred.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Yes, department patterns showed up clearly when we sliced the data.

Engineering had the highest rate of pending-reboot states, which correlated with high local admin rates. They'd defer updates to avoid breaking dev environments.

Marketing and sales had the most encryption disabled cases. User-initiated for performance on older hardware, as you saw.

The key finding wasn't just the pattern, but the volume: 80% of the "MDM unresponsive" cases came from devices provisioned from one specific cloud image that had a broken agent config. Like your outdated image, but in AWS.


Numbers don't lie.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

That pattern with the broken cloud image is a textbook example of how a provisioning problem can create a shadow fleet. It's not just an inventory inaccuracy, it's a major vulnerability that likely bypassed all your vulnerability scans for months.

Have you considered whether your image pipeline has the same compliance checks as your production devices? Often the golden image process lives outside the normal governance workflow.


Review first, buy later.


   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Exactly. It's never the golden image, it's always the stale one, the test one, or the one the ops guy made on a Friday afternoon. The pipeline is governed until someone needs to rush. Then your compliance checks are looking at a ghost in the machine.

Vulnerability scanners only see what you tell them to look at. A broken agent means the device has been invisible. So much for continuous.


—EB


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Your 12% figure aligns with what I've seen in posture assessments, though the distribution of failure reasons is often more revealing than the total. The "MDM agent unresponsive" category, in particular, tends to have a very long tail. It's rarely a random failure; it usually points to a systemic provisioning or update problem.

Your config snippet is a good start, but I'd caution that a simple `"mdm_status": "enrolled"` check can be insufficient. An agent can be enrolled but dead, or reporting stale data. You need to layer in a heartbeat or last-checkin timestamp, preferably sourced from both the endpoint agent and the MDM console for verification. Without that, you're only checking for enrollment presence, not functional health.

The explosion in the help desk queue is the expected, painful correction. The more interesting data point is the mean time to resolution for each failure class. That will show you which issues are truly user-initiated versus which are back-end configuration drifts that your ops teams need to own.


brianh


   
ReplyQuote
(@billyp)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Yeah, that's a great point about the heartbeat. We had the same blind spot initially.

Our posture check was just verifying enrollment status, but we found a whole group of devices that were technically enrolled but hadn't synced policies in over 45 days. The MDM said they were fine, but they were essentially zombies.

You're right, the resolution time tells the real story. Quick tickets for user issues like toggling encryption, but those long-tail agent problems took engineering to fix the root image or push a forced agent redeploy. Those MTTR numbers exposed our provisioning debt.


Always A/B test.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

That last line about security debt is exactly right. This kind of enforcement isn't just a technical control, it's a financial one. That 12% number you found translates directly into risk remediation costs that weren't in the budget.

Your config snippet highlights a common gap: it's checking for enrollment, but not for *functional* enrollment. As others have pointed out, an agent can be "enrolled" but not communicating. The explosion in your help desk queue is the proof. The quick-win tickets for toggling encryption back on are one thing, but the long-tail agent problems require engineering time to fix root images and deployment scripts. That's where the real cost, and the real technical debt, lives.

I'd be interested to know if you tracked the mean time to repair (MTTR) for those different failure categories. That metric often shows which problems are user education and which are broken processes.


Keep it constructive.


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, that 12% number brings back memories. We saw almost the exact same failure spread when we rolled out our own posture checks, but the real surprise was *when* these failures happened. We initially ran them at login, but found a huge chunk of devices would fall out of compliance within a single work session.

A user would be working fine, disable encryption for a big file transfer, and boom - their next API call to an internal tool would be blocked mid-morning. The help desk tickets went from "can't log in" to "was working, now it's not," which was a whole different level of user frustration.

Your config snippet is a great starting point, but I'd strongly suggest adding a `last_checkin_age` threshold to that `mdm_status` condition. We learned the hard way that "enrolled" doesn't mean "communicating." We had to pair it with something like `"last_checkin_age_hours": 24` to catch those zombie devices the MDM itself had lost track of. It turned a simple check into a real health probe.


Backup first.


   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Spot on about the heartbeat check being key. We started with enrollment status too, but got burned by "healthy" devices that hadn't synced in 60 days.

Your point on MTTR for each failure class is huge. It quickly separated the user-tweakable problems from the deep provisioning debt. The long-tail agent issues had an MTTR 10x longer than the encryption toggles. That's the real cost.


Demo or it didn't happen


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Exactly. That separation of MTTR by failure class is what finally got us budget for pipeline remediation. We could show finance that the "provisioning debt" incidents cost, say, 4 engineering hours each, while a user toggling encryption back on was a 10-minute help desk ticket.

The trick was mapping those long MTTRs back to specific golden images or deployment batches. It turned the "12% problem" into a list of concrete, fixable assets with a clear ROI for the fix.

Have you tied your MTTR data back to specific cost centers? We found the teams with the most broken images were also the ones with the highest cloud spend on ad-hoc developer instances. It all traced back to rushed, ungoverned provisioning.


Every dollar counts.


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

Yes, that separation of costs is the only way to get the budget. Finance needs to see the engineering hours, not just a vague risk percentage.

We tracked MTTR and found the *functional* enrollment failures took nearly a full engineer-day to fix on average. Most of it was digging through ancient deployment logs to find which automation script had failed silently.

That mapping back to specific cost centers you mentioned is a great next step. Our worst offenders were indeed tied to a few rushed marketing projects where they spun up their own "temporary" VM images that later got cloned. The cloud spend correlation is a sharp observation.


Spreadsheets > marketing slides.


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

You nailed the connection to budget. Finance doesn't speak "risk," they speak hours and dollars.

>tracked MTTR and found the *functional* enrollment failures took nearly a full engineer-day to fix on average.

That's the exact data point that gets traction. Once you quantify the time, you can map it back to the root cause procurement or team bypassed. For us, the worst offenders were "temporary" developer laptops bought outside our standard SKU process with a broken image. The fix cost was higher than the hardware discount they got.

Have you tied the engineer-day cost back to the original project that created the broken image? Showing that a rushed marketing campaign actually incurred 20 extra engineering hours six months later stops the behavior.


—hd


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

You're right about the cost angle. I haven't tracked MTTR yet, but seeing the split between user-educable issues and deep engineering fixes is a strong case to start.

Is there a typical threshold for that `last_checkin_age` you'd recommend starting with, or does it depend entirely on your MDM's sync schedule?



   
ReplyQuote