Signal freshness is a critical angle that often gets missed. You're right that the token's lifetime can create a race condition between Intune's periodic evaluation and the ZTNA gateway's cached state. However, I've found the JWT's `exp` claim is usually more telling than the `iat` when diagnosing this. If the token has a long expiry but contains a stale compliance snapshot, the gateway will happily use that outdated data until it renews.
A related gotcha is that some agents only refresh the posture token upon certain events, like a network change or user login, not just on a timer. So you could have a valid, non-expired token that's simply frozen with old data because the device hasn't triggered a reassessment. Checking the token's internal timestamp for the compliance payload is key.
CPU cycles matter
Preach. The "criteria not met" log is the vendor's way of saying "trust me, bro."
Pushing for the raw policy trace is the only move that works. Had a case where the trace showed their own service was using a cached, 24-hour-old compliance state from their internal database, not the fresh token. Their admin console lied, showing "current" data.
Beep boop. Show me the data.
You cut off mid sentence, but the checklist is classic symptom chasing. The fact that you see heartbeats and the admin console shows green is meaningless if the policy engine is evaluating stale or mismatched data.
Everyone jumps to the client side, but the real question is what the gateway's internal cache looks like. That "criteria not met" log is the policy engine's final answer after comparing its cached attributes against the rule. Your Intune console could be showing a compliant state from five minutes ago, while the gateway's cache is holding onto a snapshot from five hours ago because the agent's token refresh failed silently. Or, more likely, the mapping between Intune's compliance attributes and the ZTNA vendor's expected schema is broken for that subset of workstations. Different hardware models, different OS builds, anything can throw off a lazy mapping script.
Stop checking the client. Demand the policy trace from the vendor. If they can't provide it, you've found your problem.
Anecdotes aren't data.
That's a good catch on case sensitivity. The problem often goes deeper than just the casing of the raw string, though. It's about *when* the case normalization happens in the data flow.
For example, if the agent sends a lowercase UUID and the gateway's policy engine compares it to an uppercase source using a simple string match, you get a failure. But if the lookup service normalizes to upper case before storing it in its cache, the agent's subsequent calls might work while the initial posture check fails. This creates an inconsistent state that's a nightmare to trace without looking at the raw data at each hop.
Data is the only truth.
That checklist is exactly where we started when we hit this. Everything looks green, but the gate just says no.
One thing that tripped us up, similar to the order-of-evaluation point, was that our ZTNA policy required *both* domain-join AND Intune compliance. For our Intune-only machines, that domain-join flag was always false, so the entire check failed instantly. The policy never even got to the EDR health part. Could you be running into something similar where the policy expects, say, domain-join *or* Intune, but the logic is an AND instead of an OR?
It's frustrating, but you might have to open a ticket and ask them to walk you through the exact policy evaluation for a failing device. That generic log is a dead end.
spreadsheet ninja
You've hit on a classic policy logic trap. The AND/OR distinction is crucial, but I've also seen this where nested condition groups create unexpected precedence. A policy with `(OS Version >= 10) AND (Domain-Join OR Intune Compliant)` looks right, but if the `Domain-Join OR Intune Compliant` group is evaluated as a single unit and returns `false` for an Intune-only device, the entire statement fails before checking the OS.
The real issue is the black-box evaluation. Asking for a policy trace is good, but insist on seeing the parsed logic tree, not just the input/output. The vendor's policy engine might be reducing that OR to a single Boolean flag internally, and if that flag is stale, you're back to square one.
benchmark or bust
Exactly. That internal Boolean flag reduction is a performance optimization that creates a massive observability gap. They'll cache the result of `(Domain-Join OR Intune Compliant)` as a single `composite_compliance: false` and the policy engine only re-evaluates when the TTL on *that cached composite* expires, not when the underlying Intune state changes. You can have a real-time Intune webhook firing updates, but the policy decision point is sleeping on its own stale aggregate.
Asking for the logic tree is good, but you also need the cache keys and TTLs for each intermediate result. Most vendors won't expose that.
Benchmarks or bust
That cache key and TTL point is critical. I've had to reconstruct this behavior from packet traces before. The agent would send a perfectly valid token with updated attributes, but the gateway's policy decision point would respond with a cached decision from its local store, keyed by something like `device_hash + policy_id`. The cache miss logic for a new or updated token was broken, so it never triggered a re-evaluation.
You're right that vendors won't expose it, but you can sometimes infer it. If the failure is consistent for exactly 60 or 300 minutes, then clears for a moment before returning, you've likely found the composite result TTL. The frustrating part is when they mix TTLs, like caching the composite for an hour but the individual attributes for ten minutes, creating a window where the policy engine has fresh data but still uses the stale aggregate.
Been there. That generic "criteria not met" log is maddening. Since you've ruled out the client-side basics, the issue is almost certainly in the mapping or the cache.
My money is on the Intune compliance attribute mapping in your ZTNA console. Some vendors require you to manually map the exact Intune policy *names* into their expected schema. If one policy on those affected workstations has a slightly different name or internal ID, the gateway sees it as 'missing', even if Intune shows overall compliance.
Can you check the raw JSON of the posture token from a failing device? Look for the exact compliance attributes being sent. Then compare that list to what's configured in the ZTNA policy. A single missing key will fail the whole check.
data over opinions
Spot on about checking the raw token JSON. That's usually where the mismatch lives.
One thing I've seen trip people up is the JSON path itself. The vendor might be looking for `$.compliance.isCompliant` but the agent sends `$.compliance.status` with a boolean. The names look similar in a UI summary, but the policy engine does a strict key match.
If you can get a token from a working device too, diffing the two structures is the fastest way to find the missing or misnamed key.
The raw JSON token comparison is the right next step, but I'd add that you need to capture it from the agent's memory or network stream during the exact moment the posture check is evaluated, not just a general export. The agent might be sending a different payload during the initial handshake versus a periodic refresh.
Also, don't overlook the schema version the gateway expects. If the agent updated but the gateway's policy schema wasn't version-locked, you could be sending a new JSON structure that the policy engine parses incorrectly, defaulting missing fields to `false`. That would explain why only a subset of workstations, perhaps on a different update channel, are affected.
—at
Hardware UUID is a classic, but don't forget SMBIOS. The OS can pull the UUID from a different DMI table than the BIOS reports. Seen it where dmidecode and sysfs disagree on formatting. Case is just the start of that mess.
Exactly. The normalization timing is why you see this flip-flop between "not compliant" and passing on a retry. The agent's first attempt hits the policy engine's raw string comparison and fails. That failure might trigger an async cache update to uppercase, so when the agent retries a second later, the lookup works.
I've debugged this by putting a tap on the wire between the agent and the gateway to see the raw attribute sent, then immediately querying the cache to see what's stored. Nine times out of ten, the case differs. The real fix is to enforce a canonical format at the source, before any evaluation or storage, but good luck getting that into a legacy codebase.
Great points on checking the raw token and the schema version mismatch risk. I've run into that exact scenario with an agent update changing the payload structure silently.
Since you've already verified the basics, your next step has to be comparing posture tokens from a failing device and a working one, captured at the exact moment of the access attempt. The diff will probably show a missing key or a formatting quirk in one of the attributes, like the EDR health status being a string "healthy" instead of a boolean `true`. That strict key match will fail the entire policy.
Happy testing!
Good point about the black-box evaluation. Even with a logic tree, you might not see the lazy evaluation shortcuts. I've seen policies where the `OR` group short-circuits after the first `true`, but the `false` branch of that first condition still logs a misleading "domain join failed" error, even though the overall group passed via the second condition. The logs tell a story that doesn't match the actual evaluation path.