Our security team recently presented an incident post-mortem where an attacker used stolen OAuth tokens to move laterally. The critical finding was that our endpoint detection and response (EDR) solution registered no alert. The session appeared legitimate, as the token itself was valid.
This highlights a known gap: EDR and network monitoring are often blind to the misuse of a legitimately issued token after the initial authentication event. The attack chain bypassed MFA and conditional access policies because those controls were satisfied at token issuance.
I am analyzing our detection capabilities and find them insufficient. I am interested in the community's operational experience.
* What telemetry sources are you using specifically for token theft anomalies? We are evaluating Azure AD sign-in logs (specifically `IdentityLogEvents`) for impossible travel, token issuance anomalies, and unfamiliar sign-in properties.
* Are you correlating this with workload-specific logs (e.g., Microsoft 365 audit logs, Azure activity logs) to detect anomalous resource access patterns post-authentication?
* What is your threshold for response? Given a medium-confidence alert of token theft, do you revoke all sessions for the user, require re-authentication on all devices, or take a more targeted action?
Our current hypothesis is that a statistical model analyzing the sequence of accessed resources post-login may be required, as traditional signature-based detection fails. I am reviewing academic literature on lateral movement detection using authentication graphs, but practical implementation guidance is scarce.
prove it with data
You're on the right track with the Azure AD sign-in logs. That's your primary signal. But you can't just look at the identity layer.
Correlate it with the resource access logs immediately. A token issued from a user's usual IP but used minutes later to hit the Azure management API from a new ASN? That's your anomaly. Our threshold is low - if we get a match on impossible travel and an unfamiliar resource pattern, we force a token refresh and isolate the session. It's noisy, but the false positive cost is lower than a breach.
Don't expect EDR to see this. It's an identity problem.
slow pipelines make me cranky
Absolutely, correlating identity and resource logs is key. We built a similar detection using the Microsoft Graph API to pull sign-in and audit logs into a SIEM, then set up alerts for mismatched client apps or sudden spikes in token usage to unfamiliar services.
One caveat we found: the "noise" you mentioned can spike with legitimate VPN usage or developers in CI/CD pipelines. We had to add a short allow-list for known service principal access patterns to keep it manageable.
Have you run into issues with alert fatigue from automated tools acting on those same tokens?
Alert fatigue from automated tools is a real problem, and it's a strong sign your initial logic needs refinement. When we implemented similar token theft detection, we found that automatically forcing a token refresh on every alert caused major disruptions for mobile users with dynamic IPs.
Instead, we created a two-stage response. The first alert is for investigation only and goes to the security queue. Only if the correlated resource access log shows an action like role assignment or secret creation do we then trigger an automated session revocation. This added a crucial minute of human review but stopped the "crying wolf" effect that makes teams start ignoring alerts.
—HR
Your point about adding an allow-list for service principals is crucial. We followed a similar path, but found the maintenance overhead for that list became its own problem as our CI/CD and integration sprawl grew. We had to automate its curation using service principal metadata, like tagging those created by Terraform or a specific pipeline orchestrator, to keep it dynamic.
This ties directly to the alert fatigue question. Our automated tools weren't the main culprit, the sheer volume of low-fidelity alerts was. We had to move beyond simple allow-listing and implement a scoring model. A mismatch on its own gets a low score; a mismatch combined with a new geographic region and an attempted sensitive API call crosses the threshold. This dramatically cut down the investigation queue volume. Have you considered weighting different anomaly signals to reduce noise?
null
Scoring model is the way. We started similar but found that weighting wasn't enough. We had to make the score time-bound, too. A mismatch from a new region gets a medium score, but if that same session makes a low-risk API call five minutes later, the score decays. It only *accumulates* for high-risk actions. That stopped us from chasing every blip.
Totally feel you on the service principal sprawl. Automated tagging was our fix as well, using the app registration owner as a tag source. Saved so many hours.
dk
>force a token refresh and isolate the session.
That's too broad for our setup. We tier our response based on the resource. Impossible travel to a user's OneDrive? Log and investigate. Same signal targeting a key vault or management group? That's an automatic kill and immediate page. Treating every resource the same created operational noise we couldn't handle.