Graph's latency variance is the main reason we stopped using it for real-time decisions in our IR runbooks. You can't have a playbook step that fails 20% of the time because an endpoint decided to take a coffee break.
We switched to polling with exponential backoff and a hard fail after 30 seconds, then falling back to a cached state from our own sync service. It's more infrastructure to manage, but at least it's predictable.
The real joke is that the portal's "loading..." spinner is often more accurate than the API. If the spinner is still up, your script will almost certainly timeout or get a 5xx.
Your fancy demo doesn't scale.
The point about portal load times being a proxy for backend health is critical. I've observed the same pattern with Graph API latency variance. That `$filter` operation on a conditional access policy often bogs down because it's not a simple key-value fetch; it triggers a real-time evaluation against a distributed policy engine, and the portal's UI is likely making several of those calls sequentially to render a single view.
The propagation delay for service principals is a separate but related architectural issue. It's not just eventual consistency; it's the lack of a clear, queryable state. The system doesn't expose whether the principal is in a provisioning queue, replicating, or ready. This forces the arbitrary sleeps you mentioned in scripts, turning declarative operations into a stateful guessing game.
brianh