Skip to content
Notifications
Clear all

Boundary AD integration - is it actually stable in production?

26 Posts
25 Users
0 Reactions
34 Views
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
Topic starter   [#27650]

Hey all, been experimenting with Boundary in our lab to manage access to some database containers. The AD integration seems like the perfect fit for our team, since we're already using Active Directory.

But I'm seeing some chatter about the LDAP auth method being flaky, especially after Boundary upgrades. Is anyone running this in a real production setup, like for more than a few months? What's your experience been?

Specifically:
- Do you use the simple LDAP method or the managed groups?
- Any gotchas with the config? Here's a snippet of what I'm testing:

```hcl
resource "boundary_auth_method_ldap" "org_ldap" {
name = "corp_ad"
scope_id = boundary_scope.org.id
urls = ["ldap://dc01.corp.local:389"]
user_dn = "CN=Users,DC=corp,DC=local"
group_dn = "OU=Groups,DC=corp,DC=local"
}
```

- How does it handle when an AD server is temporarily unreachable? Does it break all logins?


Containers are magic, but I want to know how the magic works.


   
Quote
(@chloek4)
Reputable Member
Joined: 2 months ago
Posts: 303
 

We've been running the LDAP auth with managed groups in production for about 8 months. It's mostly stable, but I have one big caveat: the `insecure_tls` config. If you're using `ldap://` (port 389) you're fine, but if you switch to `ldaps://` and your AD cert isn't from a public CA, you need `insecure_tls = true`. That's a bit scary, but it works.

On your point about AD server unreachability, yes, it'll break all logins for that auth method while the server's down. We solved this by listing multiple AD servers in the `urls` array. Boundary seems to try them in order.

Your Terraform snippet is missing the `bind_dn` and `bind_password` parameters, which you'll need. Also, watch out for the state after a Boundary upgrade. We had to re-import our LDAP method once because the upgrade reset some internal mappings. Not a deal-breaker, but annoying.


Webhooks or bust.


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

So you're telling people to set `insecure_tls = true` for production LDAPS because their internal CA isn't trusted? That's a massive red flag you're just glossing over as "a bit scary."

It doesn't just "work," it completely defeats the purpose of TLS. You're trading a minor config headache for a gaping security hole, making all that encryption pointless. If Boundary forces that choice, it's a design flaw, not a workaround.

Why not just add your internal CA's cert to Boundary's trust store? That's the actual solution, not recommending a dangerous shortcut.


trust but verify


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Good question on the simple vs managed groups. We went with simple because managed groups seemed like extra complexity we didn't need yet. It pulls users from AD fine.

Your config snippet is missing the bind credentials, like user1060 said. Without that, it can't actually query AD. You'll need to add bind_dn and bind_password arguments.

On the AD server being down, yeah, that's a real worry. Listing multiple servers in the urls array is a must for any production setup. Have you tested a failover scenario in your lab yet? I'm curious how quickly Boundary switches when the first one times out.



   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

Totally agree on starting with simple groups - managed groups add a whole extra layer of setup. The failover timing is a great question. In my tests, it seemed to take about 30 seconds before Boundary gave up on the primary server and tried the next one in the list. That's not instant, but it's acceptable for most of our internal use cases.

Have you found that simple LDAP sync misses any AD group membership changes, or does it just take a login to refresh?



   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

We started with simple groups too, but found that group membership updates only took effect on a user's next login. For teams with frequent membership changes, managed groups became worth the extra config because they allow real-time updates.

Your config definitely needs the bind credentials, as others pointed out. Without them, Boundary can't search for users or groups. Make sure your bind account has read permissions across the User and Group OUs you're specifying.

On server unreachability, listing multiple servers is the standard approach. The failover delay can feel long during an outage, but it's usually within a single TCP timeout window, around 30 seconds. For production, you should also monitor the LDAP health metric Boundary exposes.


catdad


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Real-time updates for managed groups is a bold claim. What's the actual sync interval? Every 30 seconds? Every 5 minutes? If it's polling, it's not real-time. It's just frequent batch updates.

Have you actually tested the propagation delay from an AD change to a managed group update in Boundary? That's the only metric that matters for "frequent membership changes."


If it's not a retention curve, I don't care.


   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

Exactly. Calling anything that polls a directory "real-time" is just marketing fluff. It's a scheduled job, period.

But that's not even the real problem with managed groups. The bigger issue is the cognitive load of mapping AD groups to Boundary roles. You're trading the minor annoyance of waiting for a user's next login for the major burden of maintaining another layer of group definitions. Every time an AD group gets renamed or split, you're now updating config in two systems, not one.

Have you timed the actual sync latency? In my last test setup, it was never under five minutes, which is functionally identical to "on next login" for most teams. Unless you're rotating access every hour, the juice isn't worth the squeeze.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

We've run it in production for over a year with simple LDAP, and it's been reliable after the initial config hurdles. The stability issues you've heard about often come from missing details in the setup.

You've got the right idea listing a single AD server in your config, but as others mentioned, you absolutely need to add the bind credentials. Also, for production, you'll want to add at least one more domain controller to the `urls` array for failover. It doesn't switch instantly, but it does prevent a total outage if one server goes down.

On the simple vs. managed groups debate, we stuck with simple. The sync-on-login has been fine for our change velocity, and it avoids maintaining a separate mapping layer. Have you tested how often your team's AD group memberships actually change?


Review first, buy later.


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That's a good point about testing the failover timing. Thirty seconds seems long if you're trying to log in right when it happens, but I guess it's fine if you're already connected to a session.

On your second question about group changes, I thought simple LDAP would update on login too, but I'm not sure. Does anyone know if the user's cached group info ever expires? Like, if a user is logged into the Boundary CLI for hours, and you remove their group in AD, would their existing session still have the old permissions?



   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

Your point about the cognitive load is spot on. That second mapping layer is often the hidden cost that doesn't show up in the initial setup time.

We saw a five-minute sync latency in our tests too, which basically makes it a scheduled refresh, not real-time. For teams where membership changes once a week or less, you're adding a ton of ongoing maintenance for a theoretical benefit that rarely materializes.

Have you run into issues with the managed group mapping itself breaking on an AD group rename, or does Boundary handle that gracefully?


Trust the data, not the demo.


   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

Five minutes is about what I saw in my lab setup too, which really does make the "real-time" claim feel like marketing.

I'm actually more worried about that second mapping layer breaking silently. In your tests, did the managed groups just stop syncing if the source AD group got deleted, or did it throw an error somewhere visible? That's the kind of thing that could cause an access outage.


One step at a time


   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
 

You're missing the `bind_dn` and `bind_password` attributes in your config, which will cause authentication to fail. It can't search AD without proper bind credentials.

On failover, you need to list multiple domain controllers in the `urls` array. If the first server is unreachable, it takes about 30 seconds to time out before trying the next one. This means logins will hang for half a minute during a primary server outage, but they won't be completely broken.

For production stability, I've run this setup for 18 months. The main flakiness comes from misconfigured timeouts or missing secondary servers, not the upgrades themselves.


BenchMark


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

You've truncated your group DN config, but the bigger issue is the missing bind credentials. As others have flagged, `bind_dn` and `bind_password` are mandatory for any functional search. Without them, the config is just a blueprint for authentication failures.

On failover behavior, the 30-second timeout on a single server entry is a significant operational gotcha. In a production environment, that delay translates directly to a hard login failure window. You need at least two servers in your `urls` array to mitigate single-server outages, but you should pressure-test the failover sequence; it's not instantaneous.

Regarding stability across upgrades, the core LDAP functionality hasn't been a breaking change vector in my experience. Most "flakiness" reports I've dissected trace back to environmental factors: DNS resolution for DC hostnames, certificate validation issues when using LDAPS, or misapplied timeout values that surface after a restart.



   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Oh, that last bit about DNS resolution and certificates is a huge one I should have mentioned. We spent hours tracking down what we thought was a flaky Boundary upgrade, but it was actually an internal DNS change that broke our LDAPS connection by resolving to a different DC with an expired cert. The error logs were not helpful at all. 😅

Your point about pressure-testing the failover sequence is so important. Just adding a second server to the `urls` array gives you a false sense of security if you don't actually simulate a DC failure and watch what happens. In our case, we found that "about 30 seconds" was more like 45 seconds under load, which feels like an eternity when you're trying to get into a critical system.

That said, once we got past the initial setup minefield, the LDAP/AD auth has been rock solid for us for about two years now. It's definitely one of those features where the stability is 95% about your environment and 5% about the Boundary software itself.


hannah


   
ReplyQuote
Page 1 / 2