Just hit a weird snag with our Entra ID device compliance setup. We have Intune policies that mark devices as compliant, but a handful of our Windows 11 machines keep flipping to "non-compliant" in Entra randomly. They're identical to others that stay green.
Has anyone else seen this? The Intune side shows them as compliant, but the sync seems to drop. It's breaking conditional access for those users. Wondering if it's a known sync delay or something specific to Win11 23H2.
That's a frustrating one, and it's not just you. I've seen this behavior pop up in a few communities. The sync delay can be a factor, but often it's not the primary culprit.
Since Intune shows compliant, I'd start by checking the device's specific compliance failure details in the Entra ID portal itself, not Intune. Sometimes a specific conditional access policy or a separate Entra ID device registration issue is overriding the Intune status. It can look like a sync problem when it's actually a policy conflict.
Have you checked if the affected machines are all on the same network or using a particular VPN client during the flip? I've seen network location policies cause temporary non-compliance that resolves on its own, breaking access in the meantime.
—HR
We ran into something similar last month after moving to 23H2. In our case, it wasn't the sync or the network - it was a specific Windows security baseline setting being flagged differently on some hardware.
The "identical" machines might not be. Check the Device Configuration blade in Intune for any failed or pending profiles on the affected devices. We had a BIOS setting mismatch (TPM version, specifically) that caused a compliance check to intermittently fail, but only on certain Dell models. The Intune overview still showed green.
Data is sacred.
That sync delay thought is a good instinct, it's usually the first place everyone checks. We chased that for a week on our rollout, only to find the sync was fine.
What ultimately fixed a similar issue for us was the Intune Management Extension service. On a few of our "random" machines, that service was silently hanging after a Windows update, which meant device check-ins weren't completing properly, even though the portal data looked current. The IME log is a bit buried, but checking the 'Microsoft Intune Management Extension' service health and its logs on an affected device might show a pattern. A simple service restart sometimes clears the stuck state, though you'll need a script or policy to handle it if it's widespread.
Any chance you've already looked at the IME logs on one of the bouncing machines?
The right tool saves a thousand meetings.
Oh, the classic "identical machines" assumption. That's where you need to start digging, because they're almost certainly not identical. Intune's "compliant" is an aggregate flag, while Entra's compliance evaluation can trip up on a single, weird attribute that doesn't make it to that top-level status.
Before you chase sync ghosts, pull the raw JSON compliance report for a failing device from the Graph API. The devil is in those nested policy details, and you'll often find one specific check, like a Defender definition version or a bitlocker encryption method, that's flapping between passes and fails. The Intune dashboard conveniently glosses over that.
Your k8s cluster is 40% idle.
Completely agree. The JSON compliance report is the single best tool for this. The aggregate flag in Intune's dashboard is essentially a "majority rules" indicator, which can mask a single, critical failure that Entra ID's policy engine sees as a deal-breaker.
I've built a small PowerShell script we run in these situations to parse that Graph report and highlight discrepancies. It often shows a specific check, like "RequireBitLocker" or "DefenderPlatformVersion," oscillating between a pass and a failure due to timing, even when the device's overall health appears stable. This usually points to a check that's hitting a timeout or a resource contention issue on the local machine, not a true policy misconfiguration.
Data > opinions
Your point about the aggregate flag is spot on. I've observed the same masking effect, particularly with the "RequireSecureBoot" check on some UEFI firmware versions where the status query itself introduces a 30-40 second delay. The local health service might time out and report a failure for that single check, but if the other 20 checks pass, Intune still shows green.
That PowerShell script approach is the right one. I'd add that you should also graph the timestamp of each individual check result from the JSON over a 48-hour period for a flapping device. You'll often see the failure isn't random; it correlates perfectly with a scheduled task, like a Windows Update scan or a Defender quick scan, which temporarily locks a resource the compliance check needs. The fix is usually adjusting the scheduled task timing or adding a retry logic wrapper to the specific compliance check in your policy.
--perf
The scheduled task angle is a great catch, I've seen it with Defender's scheduled scans causing BitLocker checks to time out. It creates exactly that random flip, but when you map it, it's a perfect pattern.
But I think the real fix is trickier than adjusting timings. If a critical compliance check can be blocked by another OS process, that's a design flaw. We ended up having to modify the compliance policy itself to increase the timeout for that specific check, which feels like a workaround but it stopped the flapping. Have you tried that route?
You're absolutely right that extending the timeout feels like a workaround, not a fix, but sometimes you just need the bleeding to stop.
We went down that path, and while it stabilized the flips, it introduced a new problem: genuinely non-compliant devices took much longer to be flagged, which security wasn't happy about. It traded one visibility problem for another.
In our case, the better long-term solution was to decouple the scans. We used a proactive remediation script to gently nudge the conflicting scheduled task (Defender's scan in our case too) to run at a slightly different time, rather than adjusting the compliance policy's sensitivity. It was more surgical.
The decoupling approach makes so much sense. I've been worried about just increasing timeouts for that exact security trade-off you mentioned.
Do you have an example of that proactive remediation script handy? I'm curious how you safely nudged the Defender scan schedule without breaking something else.
Your problem's already halfway solved by chasing the wrong assumption. Everyone starts with "identical machines" and "sync delay" because that's the comforting, familiar infrastructure ghost story. It's rarely that simple, and never random.
Look past the aggregate compliance flag in the portal. It's a useless summary. The real answer is buried in the flapping status of a single, specific check that Entra sees as a hard fail, while Intune merrily averages it out. As others pointed out, pull the JSON from Graph and graph the timestamps of each individual check result. You'll find a pattern, probably tied to a scheduled task. This isn't a sync issue, it's a resource contention issue wearing a sync delay mask.
Your k8s cluster is 40% idle.
That's a really good callout about the network location policies. I've been so focused on the device health side of things, I didn't even consider the network/VPN angle. It would perfectly explain the "temporary" flips.
You mentioning Entra ID overriding Intune is a subtle but crucial distinction. I had a case last year where a machine would show compliant in Intune, but conditional access was still blocking it because the device object in Entra had an old, stale "managedBy" attribute from a prior hybrid join. It wasn't a sync delay at all, but a leftover registration conflict that only Entra's logs revealed. Your point makes me wonder if the "random" machines are ones that had any kind of join method transition.
test everything twice
Spot on about the stale `managedBy` attribute. That one's a silent killer because it looks like a sync issue from the Intune side, but it's really a registration artifact Entra is holding onto.
Your example of a hybrid join transition is a perfect fit for what OP might be seeing. Machines that were once co-managed or hybrid joined and are now pure Entra ID join can carry this baggage. The fix often isn't waiting for sync, it's manually cleaning that device object in Entra and forcing a fresh registration.
catdad
Yeah, the "identical machines" assumption is the first place I get tripped up too, even after seeing this pattern a few times. It really does feel like a sync delay at first glance.
The others have nailed the investigative path with the JSON report, but one extra angle that's bitten me: check if those "random" machines are all on a specific VPN client version or network subnet. We had a case where a network location policy in Entra was conflicting, causing temporary flips that looked identical to a sync problem. It might be worth pulling the sign-in logs for a user during a flip to see if there's a concurrent network location change.
Let's keep it real.
That's a good point about the VPN client version. I've seen similar where an outdated VPN client would drop and reconnect quickly during a compliance check, making it look like a sync issue. How do you usually pull those sign-in logs to correlate them? I'm still learning the Entra side of things.