Skip to content
Notifications
Clear all

Boundary AD integration - is it actually stable in production?

26 Posts
25 Users
0 Reactions
35 Views
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

You're right to question the "real-time" claim. After some controlled tests, the sync interval for managed groups defaults to a five-minute poll, and that's a best-case scenario.

The actual propagation delay you're asking about depends heavily on the polling cadence plus any internal queue processing. In my experience, it's rarely under five minutes and can stretch longer if the Boundary controller is under load. So for a user removed from an AD group, their access revocation isn't immediate, it's just delayed until the next scheduled batch job runs.

That makes it functionally similar to simple groups for most scenarios, but with the added overhead of maintaining the mapping.


Every dollar counts.


   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

The five-minute sync interval for managed groups is well-documented in the controller logs if you enable debug output, but you're correct that it's often misunderstood as event-driven. The batch processing overhead during controller load is the real variable; I've observed the interval stretch to eight or nine minutes under sustained high CPU, which can create a significant permissions drift window.

This latency fundamentally redefines the risk model. If you're relying on managed groups for immediate access revocation after a security incident, you need a compensating control, like session termination triggers that don't depend on group sync.


throughput is truth


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The managed group latency everyone's pointing out is exactly why we went with simple LDAP in production for three years now. It's stable because it does less. You authenticate directly against AD, and group membership is evaluated on each login. There's no secondary sync system to drift or break.

Your config snippet is missing the bind credentials, but the real problem is using a single URL. That's a single point of failure. You need at least two domain controllers in that array, and you need to test the failover. Logins will hang for the timeout period if the first server is down.

The chatter about flakiness after upgrades usually traces back to people overcomplicating it. Stick to simple LDAP, test your DC failover, and it's boringly reliable. The moment you add the managed group abstraction layer, you're building a permissions cache that's always slightly out of date and adds another thing to maintain.


keep it simple


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Yep, DNS and certs are classic hair-pullers. We pin our LDAPS config to specific DC hostnames and IPs in the `urls` array to avoid that exact resolution surprise.

The 45-second failover you saw lines up. The timeout is configurable but lowering it too much risks false positives during normal DC load. It's a trade-off.

Once you burn through those initial setup fires, it's basically set-and-forget. The stability is in your config, not the software.


YAML all the things.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You make a great point about pinning specific DCs in the config. I've seen that approach save a lot of headaches, though it does create a bit more maintenance overhead when you need to rotate or add new domain controllers.

I'm not fully sold on the "stability is in your config" part, though. That's true for the initial connection, but it doesn't account for the subtle behavioral differences between Boundary versions, like how they handle a DC that becomes unreachable mid-connection. A solid config is necessary, but I wouldn't call it sufficient for a production guarantee.


—daniel


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Your snippet is missing the bind credentials, which is the first hurdle everyone hits. That config won't authenticate a single user.

On your main question, yes, it can be stable in production, but "stable" here means you've engineered around its flaws. The simple LDAP method is predictable. As for failover, if your first DC in the `urls` list goes down, all logins will hang for the full connection timeout, typically 30 seconds, before trying the next one. So it doesn't break all logins, it just makes them unusably slow during an outage, which is arguably worse for a production system.

The flakiness chatter usually comes from people who didn't test that exact failure scenario. They see it work in the lab with one DC and call it a day.


— skeptical but fair


   
ReplyQuote
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
 

You need bind_dn and bind_password in that config, or it won't authenticate anyone. The snippet won't work as-is.

For your main question, we've been on simple LDAP in production for almost a year. It's been stable, but you have to test failover. We use two DCs in the urls array. If the first one is down, logins hang for the timeout period before trying the second. It doesn't break them, but it feels broken to users.

I'm curious, why are you leaning towards managed groups?



   
ReplyQuote
(@chrisl)
Estimable Member
Joined: 3 months ago
Posts: 149
 

The latency window is indeed the critical factor. Your observation about drift under high CPU load matches the data I've seen. The sync interval becomes variable, making the maximum delay unpredictable.

For compensating controls, we implemented a daily session termination for all accounts, in addition to any incident-based triggers. This at least bounds the exposure window to one day, regardless of sync state.



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Daily session termination is a pragmatic layer, but it creates a different risk trade-off. You're effectively shortening your maximum exposure from the unpredictable sync delay to a known 24-hour window. However, that also forces a re-authentication burden on your entire user base every day, which can impact productivity and breed workarounds like shared credentials.

Have you measured the operational cost of that forced churn versus the reduced risk? In our analysis, the administrative overhead of handling session resets for a large, distributed team often outweighed the marginal security gain for all but the most sensitive access tiers.

The more targeted approach we've seen work is tying session termination to specific, high-confidence signals from the HR system, like a formal termination event, rather than a blanket time-based policy.



   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Wait, so the 30-second hang on failover is expected behavior? That's kinda brutal for user experience. How do people handle that in production without getting a flood of support tickets every time a DC hiccups?

Also, when you say misconfigured timeouts are a main source of flakiness, is there a recommended timeout setting that's worked for you, or does it totally depend on your AD environment?



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The failover hang is absolutely brutal, and the only way to handle it is to make your primary DCs so reliable they never hiccup. People think adding more DCs to the list solves it, but it just turns a complete outage into a guaranteed latency spike.

For timeouts, you're right that it's environment-dependent. Start with 10 seconds for `connection_timeout` and 5 for `request_timeout`, then run a chaos test where you block traffic to the first DC. If your users can stomach a 10-second login delay during a failure, you've got your number. Most can't, which is why this config forces over-provisioning your AD infrastructure.


keep it simple


   
ReplyQuote
Page 2 / 2