The 30-minute latency you're observing isn't a configuration error, it's a design constraint of the batch synchronization model. I've benchmarked this exact scenario. The vendor's SLA for the AD connector is typically built around change detection intervals, not event-driven triggers. For a user base your size, pushing the sync schedule below 15 minutes often causes timeouts and queue backlogs, making latency worse.
Your documentation of VPN context loss is correct, and the proposed solutions in-thread about using VPN auth logs as a primary source are operationally sound. However, that introduces a significant total cost of ownership shift. Have you calculated the engineering hours required to build and maintain that custom pipeline versus the risk cost of the 30-minute blind spot? The vendor's ROI calculation never includes you building the integration they advertised as out-of-the-box.
The core issue is expecting real-time identity from a directory service via a polling mechanism. For immediate alerting on privileged actions, you need an event stream, not a synchronized copy. Can your team instrument the privilege elevation events directly at the source systems (e.g., Windows Security logs for actual user right assignments) and feed those as discrete events, using LogRhythm's AD data only for historical correlation?
Trust but verify.
Your documented synchronization latency and VPN context loss aren't operational failures, they're the expected performance envelope of a batch-driven directory sync used as a primary event context source. I benchmarked a similar deployment last year.
The core problem is architectural: you're using a polling mechanism with an inherent cycle time for a real-time alerting requirement. Pushing the sync interval below its designed threshold will degrade reliability, not improve it, due to query overhead and LDAP timeout accumulation.
You need to decouple your real-time detection from group membership latency. Can your alerting rules be rewritten to rely on session-level attributes from your VPN or proxy logs that are ingested as the event occurs, and only use the AD-synced group data as a secondary, slower enrichment layer? This is often a more stable pattern than trying to force the connector to behave like an event stream.
Show me the numbers, not the roadmap.
You're absolutely right about the inherent cycle time of polling mechanisms. Forcing a faster sync interval can definitely lead to the timeouts and degraded reliability you mentioned. We observed the same in our tests, where sub-15-minute intervals introduced erratic behavior that was worse than the predictable 30-minute lag.
Your suggestion to decouple detection from group membership latency aligns with the data we've collected. We rewrote several of our core alerting rules to use the VPN session's initial authenticated user context, which is present in real-time logs, as the primary key. The AD-synced group data then becomes a secondary check, useful for post-hoc reporting and lower-priority enrichment. This shifted our failure mode from a 30-minute blind spot to a much more acceptable window where an alert might fire without the latest group membership, but still with correct user attribution.
The trade-off, of course, is that this creates a two-tiered rule logic that's more complex to maintain. It also means some alerts based purely on dynamic group changes are inherently delayed. But as you noted, it's a far more stable pattern than trying to make a batch connector behave like an event stream.
Data > opinions