The stale `managedBy` attribute is one of those problems that's blindingly obvious once you've seen it a few times, but it's completely invisible from the Intune console. You're right that cleaning the object is the fix, but I've found you often need to go a step further than a simple delete.
If you just delete the device record in Entra ID, the local dsregcmd state on the machine can still have the old registration details cached. The machine will often just re-register with the same stale attributes unless you also run `dsregcmd /debug /leave` and clear the local registration before rejoining. It's a two-part cleanup.
Also, watch for machines that flip back to non-compliant a few days after a "successful" manual fix. That usually means you have a GPO or a leftover SCCM client action somewhere that's periodically re-writing the old hybrid join info, triggering the whole cycle again.
Show me the benchmarks
That 23H2 callout is interesting - it's usually a good instinct to check the OS build. I've seen similar things happen after feature updates because a specific Windows component version can change how the compliance check interacts with the TPM or Secure Boot.
But your point about the machines being identical is key. In my experience, when a few machines in a seemingly identical set start flipping, it often comes down to a subtle difference in their workload or connection pattern. The machines that stay green might just be getting their compliance checks at a slightly different time, avoiding a resource conflict with another process like a scheduled antivirus scan or a VPN tunnel reconnection. Have you checked if the flipping coincides with any other scheduled task logs on those specific boxes?
ship early, test often
You're right to focus on the workload angle, but I'd argue against assuming it's purely about timing collisions with scheduled tasks. In our environment, the correlation wasn't with the tasks themselves, but with the *completion state* of certain tasks.
We traced an identical flipping pattern to a specific version of our disk encryption agent. When its daily verification task finished with a particular warning code (which didn't cause a failure in the task scheduler log), it left a temporary file lock on a system resource. That lock didn't block the next scheduled task, but it *did* block Intune's compliance provider when it ran its check within the next 90-second window. Machines with even a one-minute difference in their agent check schedule would avoid the lock entirely, appearing identical in configuration but divergent in compliance results.
The logs to scrutinize aren't just Task Scheduler; look at the sequential timestamps in the DeviceManagement-Enterprise-Diagnostic-Provider log for the failing device and compare them to the Application log for your security/encryption/VPN agents. The overlap is often milliseconds.
That's a really sharp observation about the completion state, not just the schedule. I've seen the same pattern with a network inspection driver that would briefly hold a registry key handle open after its scan finished - a clean exit, but it left a lock for maybe 500ms. The compliance check hitting in that exact window would get an access denied on a simple registry query, flagging the device.
Your point about checking the sequential timestamps is spot on. I'd add that you can sometimes catch this in a ProcMon capture filtered for the MDM components, looking for `SHARING_VIOLATION` errors. The trick is getting the capture timed right, which is nearly impossible unless you're proactively monitoring a known flipper.
Latency is the enemy, but consistency is the goal.
The `SHARING_VIOLATION` lead from a ProcMon trace is a solid diagnostic path, but its utility is limited by the temporal resolution of the logs you're correlating with. The MDM-Diagnostic-Provider operational log (Event ID 2150) captures the result code, but the 0x80070020 error for a file lock is often logged as a generic failure without the specific offending handle.
A more deterministic approach is to run the compliance evaluation on demand while a controlled lock is in place. You can use the `Sync-MgDeviceManagementManagedDevice` cmdlet via the Graph PowerShell SDK to trigger a fresh compliance check while simulating the resource conflict, say, by using `Handle.exe` to lock the suspected registry key. This bypasses the need to catch the random 500ms window.
Nullius in verba
The 23H2 angle is definitely a valid starting point, but I'd lean toward the workload timing theories posted after you. We've seen similar flipping where the OS version was a constant, but the intermittent failures tracked back to resource locks from non-Intune services during the compliance check window.
If you're seeing the Intune console stay green while Entra goes red, you could test the sync timing artificially. Force a compliance check using Graph PowerShell while you temporarily lock a registry key or WMI namespace - sometimes you can reproduce the failure on demand and confirm it's a resource collision, not a sync bug.
Any chance those "identical" machines have a slight variation in their third-party security client or VPN version? Even a small patch difference can shift their internal task schedules enough to cause these random collisions.
Every dollar counts.
Absolutely on point about testing with artificial locks. That's the kind of controlled experiment that moves you from guesswork to a root cause.
Your mention of third-party client version differences is huge. I've chased this exact ghost in a fleet where the "identical" machines had a one-week stagger in a CrowdStrike sensor update rollout. The newer sensor version changed its scan completion routine, which inadvertently held a WMI namespace open a fraction of a second longer. That tiny shift in schedule was enough to make a subset of machines collide with the compliance check like clockwork.
It makes you wonder how many of these "random" flips are actually just deterministic chaos from micro-variations in the environment.
That "deterministic chaos" line really nails it. We've been pulling our hair out over a similar pattern where the compliance flips looked random, but when we finally mapped them against patch Tuesday schedules across different departments, the "chaos" suddenly had a very clear timeline. The micro-variation wasn't even in the security agent version - it was in the reboot lag after updates. One department reboots same-day, another waits a week. That slight shift in uptime cycles was enough to desync all the other scheduled tasks.
So, would you say the fix is more about hardening the compliance check itself against these tiny locks, or is it about enforcing stricter synchronization across those third-party service schedules?
The ProcMon trace idea sounds useful, but you're right about the timing. How do you even set that up without already knowing which machine will fail next? It seems like you'd need to catch it in the act.
I've seen similar lock issues in AWS with CloudWatch agent configs, where a script holds a log file open just a bit too long. Feels like the same kind of race condition, just in a different place.
The whole "catch it in the act" problem is exactly why ProcMon is often a dead end for this. You're chasing a ghost. The better trick is to *create* the condition artificially on a test box. Use a script to lock the resource you suspect, then manually trigger a compliance check. If it fails, you've got your smoking gun without needing perfect timing on a production flipper.
And yeah, the AWS comparison is spot on. It's the same core race condition, just dressed up in different vendor jargon. Makes you wonder why we keep buying systems that are this fragile to millisecond locks.
FOSS advocate
Sync delay can explain it, but more often it's a false negative from the compliance check itself. The fact that Intune shows green while Entra shows red tells me the compliance evaluation succeeded locally, but the result that got synced up failed on a transient error.
We chased a similar ghost last quarter. It turned out the compliance provider was hitting a WMI query for disk encryption status right as our third-party AV did its daily scan. The AV locked the namespace for a split second, causing the WMI call to fail. That failure got logged and synced as "not compliant," even though the device was technically fine.
Check the MDM-Diagnostic-Provider logs on a flipper for event 2150 right around the time it flipped. Look for `0x80041001` or similar access-denied codes. If you see them, you've got a resource lock, not a sync problem.
Automate everything. Twice.
That's a perfect example of the false negative pattern. It's exactly what I was getting at with the "deterministic chaos" idea - a tiny, predictable lock from another service creates a real but meaningless compliance failure.
Your point about Event 2150 is the right diagnostic path. We found that `0x80041001` (WBEM_E_ACCESS_DENIED) was the most common, but we also saw a fair number of `0x8004106C` (WBEM_E_PROVIDER_LOAD_FAILURE) codes. Those were even trickier, because they often pointed to a WMI provider that was temporarily unavailable during a scan, not just locked.
One caveat I'd add: sometimes the failure in the logs is for a *different* compliance rule than the one that shows as non-compliant in the portal. We'd see an access-denied error for a BitLocker WMI call, but the portal would flag the device for "Microsoft Defender status not up to date." It's like the sync process gets confused by any failure in the evaluation batch. Makes the logs feel misleading until you piece it together.
Oh, the *different compliance rule* mismatch you mentioned is a massive headache. We logged a support ticket for that exact behavior last month. The Intune engineer confirmed the portal can show the *last* successful sync state for a specific setting if a subsequent rule fails in the batch, making the error look like it's in the wrong place.
It really undermines trust in the dashboard. You start chasing Defender updates when the real culprit is a BitLocker WMI glitch.
Your WBEM_E_PROVIDER_LOAD_FAILURE note is spot on, too. We traced one of those back to a Dell Command Update process holding WMI hostage during its driver check.
Cheers, Henry
That's a solid starting point, and it matches the pattern. We ran into a similar cert-related drop, but the root cause was slightly different: a required intermediate certificate wasn't present in the local machine's trusted intermediate store on a subset of our cloned workstation images. The device's own cert was valid, but the chain broke during the policy check, causing a transient failure.
It often appears random because the check might succeed if the chain can be fetched from a CA in time, and fail if there's any network latency during validation.
Measure twice, buy once.
Oh yeah, that "Intune green, Entra red" mismatch is a classic. We had the same heartburn last year. Everyone jumps to sync delay, but more often it's a race condition on the local compliance check itself.
Like others said, check the MDM logs for WMI failures (event 2150). But also, don't assume the non-compliant setting shown in the Entra portal is the actual culprit. The dashboard can be misleading. We spent weeks chasing a firewall rule when it was actually a BitLocker provider timing out.
Keep it simple.