Skip to content
Notifications
Clear all

Guide: Automating Boundary host set updates from your CMDB

30 Posts
29 Users
0 Reactions
101 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
Topic starter   [#23298]

A common operational friction point when deploying Boundary is maintaining accurate host sets for dynamic infrastructure. Manually updating IPs and tags defeats the purpose of a zero-trust network. This guide outlines a practical pattern for synchronizing Boundary host catalogs with an external CMDB or inventory source, using Boundary's Go SDK and a simple reconciler pattern.

The core concept is a control loop that fetches the current desired state from your CMDB (e.g., via its API) and reconciles it with the existing host sets in a designated Boundary scope. We'll use a service account with appropriate permissions in Boundary, stored via a `boundary` auth method. Here is the essential reconciliation logic in Go:

```go
package main

import (
"context"
"github.com/hashicorp/boundary/api"
"github.com/hashicorp/boundary/api/hosts"
"github.com/hashicorp/boundary/api/hostsets"
)

func syncHostSet(cmdbHosts []CMDBHost, boundaryClient *api.Client, hostCatalogId, hostSetId string) error {
ctx := context.Background()

// 1. Read existing hosts in the catalog
hClient := hosts.NewClient(boundaryClient)
hostList, err := hClient.List(ctx, hostCatalogId)
if err != nil { return err }

// 2. Map existing hosts by external ID (from CMDB)
existingHosts := make(map[string]*hosts.Host)
for _, host := range hostList.Items {
if host.ExternalId != "" {
existingHosts[host.ExternalId] = host
}
}

// 3. Determine creates, updates, deletes
for _, cmdbHost := range cmdbHosts {
if _, exists := existingHosts[cmdbHost.ID]; !exists {
// Create new host in Boundary catalog
_, err := hClient.Create(ctx, hostCatalogId,
hosts.WithName(cmdbHost.Name),
hosts.WithHostAddresses(cmdbHost.IP),
hosts.WithExternalId(cmdbHost.ID))
if err != nil { /* handle */ }
}
delete(existingHosts, cmdbHost.ID)
}

// 4. Delete hosts no longer in CMDB
for _, toDelete := range existingHosts {
_, err := hClient.Delete(ctx, toDelete.Id)
if err != nil { /* handle */ }
}

// 5. Re-fetch all host IDs and update the host set membership
updatedHostList, _ := hClient.List(ctx, hostCatalogId)
var hostIds []string
for _, host := range updatedHostList.Items {
hostIds = append(hostIds, host.Id)
}

hsClient := hostsets.NewClient(boundaryClient)
_, err = hsClient.SetHosts(ctx, hostSetId, 0, hostIds)
return err
}
```

Key implementation notes:
* The `ExternalId` attribute is crucial for idempotent mapping between CMDB entities and Boundary hosts.
* Always use the version field (`0` in the example) for the `SetHosts` call to manage concurrency; fetch the current version from the host set object in a real implementation.
* Run this reconciler as a periodic job within your CI/CD pipeline or as a dedicated microservice. The Boundary service account should have `ids=*;actions=*` permissions on the host catalog and host set.
* For large inventories, implement batch operations and consider rate limits.

This approach reduces drift and ensures that Boundary access policies are consistently enforced against your current infrastructure baseline, a significant improvement over static configuration.

benchmark or bust


benchmark or bust


   
Quote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

This is exactly the kind of pattern we've been looking to implement, thank you for putting it together. I'm curious about one practical aspect, though. In a manufacturing context, our CMDB often has hosts that are temporarily offline for maintenance or in a decommissioning queue. Does your reconciliation logic account for a soft delete or a status flag, or would it simply remove those hosts from the Boundary host set entirely? I'm thinking we'd need to preserve them in Boundary but perhaps adjust their attributes or move them to a separate "quarantine" host set based on the CMDB state. How would you extend the example to handle that transition gracefully without losing the host object?



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

That's a great question, and it's a scenario we've had to handle too. The pattern I sketched just shows the core loop; you'd need to incorporate the status from your CMDB into your comparison logic.

You could add a field like `DesiredState` to your internal `CMDBHost` model - values like "active", "maintenance", "decommissioning". Then, in your reconciler, instead of just adding/removing, you'd have a third path: updating host attributes or moving hosts between sets.

For a quarantine set, you'd need the IDs of both the primary and quarantine host sets. Your sync function would iterate through the CMDB list and decide for each host: if status is "active", ensure it's in the primary set; if "maintenance", ensure it's in the quarantine set (and maybe update a `reason` attribute). You only delete from Boundary when the CMDB record is gone or marked "decommissioned" *and* a grace period has passed.

The key is making the host's external ID (from your CMDB) a stable, stored attribute in Boundary, so you can track it across sets.


Sleep is for the weak


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Adding a 'DesiredState' field is the obvious move, but I'm always suspicious of how cleanly that maps from a real CMDB. What's your source of truth's actual schema? Most I've seen have three different status fields across different tables, not a single, reliable column.

That external ID point is critical, though. If you don't get that right from day one, your reconciliation turns into a mess of duplicate hosts. But you're also assuming Boundary's host attributes can handle the load. Have you stress-tested updating attributes on thousands of hosts every sync cycle? The API can get chatty.



   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Good catch on the attribute update overhead. The API calls can become a bottleneck at scale. For bulk operations, you're better off using the host set's `set_host_sources` method to apply changes in batches rather than updating each host individually. A delta comparison in your reconciler is key to avoid unnecessary API chatter.

Regarding the CMDB schema mismatch, that's why I'd advocate for a translation layer in the sync service. Let it normalize those three status fields into a single operational state your Boundary logic can consume. It adds complexity but prevents your sync code from being tightly coupled to the CMDB's quirks.


Numbers don't lie


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

You've identified the central challenge in any external sync pattern: aligning two systems with fundamentally different data models. Your suggestion for a dedicated translation layer is the correct architectural answer, even if it feels like extra work initially.

That layer doesn't just map fields; it should encapsulate all the business logic for what constitutes an "active" host from your company's perspective. It can consume those three disparate CMDB status fields, apply rules about maintenance windows, and output a normalized state the reconciler understands. This keeps your Boundary-specific code clean and testable.

The batch operation point raised by user458 is crucial for performance, but it also interacts with this translation layer. Your delta logic should operate on the normalized output, not the raw CMDB data, to decide which batch operations are required. This decoupling means a change in the CMDB schema only affects one component.


Let's keep it constructive


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

You're right about the core reconciliation pattern being the key. I've seen teams get tangled up trying to handle every edge case in the first version of the sync loop. Starting with a simple "add missing, remove stale" logic often reveals the actual business rules you need, like those status flags, without overcomplicating things from day one.

One thing I'd add is to make that first `List` call for existing hosts more resilient. If your host catalog grows large, you need to handle pagination in that initial fetch, or your comparison will be working off incomplete data. The SDK handles it, but it's easy to miss in a first draft.


~Harry


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

> If your host catalog grows large, you need to handle pagination in that initial fetch

Yes, absolutely. It's a classic footgun in the early versions of a sync service, and you'll only spot it when your host list gets long enough that the first page is incomplete. The SDK's `List` methods return a `*api.ListResult`, which has a `PaginationToken` for getting the next page. It's a straightforward loop, but forgetting it means you're silently missing hosts, and your reconciliation will start deleting things it shouldn't.

Here's the quick fix you'd add around that `hClient.List` call to be safe:

```go
var allExistingHosts []*hosts.Host
for paginationToken := ""; ; {
hostList, err := hClient.List(ctx, hostCatalogId, hosts.WithPaginationToken(paginationToken))
if err != nil {
return err
}
allExistingHosts = append(allExistingHosts, hostList.Items...)
if hostList.PaginationToken == "" {
break
}
paginationToken = hostList.PaginationToken
}
```

Doing this from the start saves a nasty debugging session later. It makes that initial state capture a bit heavier, but it's the only way to be sure your comparison logic has the full picture before deciding what to add or remove.


— francesc


   
ReplyQuote
(@annie82)
Reputable Member
Joined: 3 months ago
Posts: 232
 

Oh, that pagination trap is a really good shout, I'd have totally missed that. It's exactly the kind of thing you don't think about until everything breaks after your first big success with the script. Thanks for the code snippet too, it makes it super clear.

I'm curious though, when you're doing this loop to fetch everything, does it make the initial comparison a lot slower if you have thousands of hosts? And if so, is that just the price you pay for correctness?



   
ReplyQuote
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

It's the foundational pattern, sure. But glossing over pagination from the start is a classic rookie mistake that'll bite you later. Skipping it doesn't just affect performance. It completely undermines reconciliation correctness once your host count exceeds the default page size. The first "big success" quickly becomes a production incident where hosts start disappearing.


Prove it


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Exactly. It's not a performance tax for correctness, it's table stakes for the script to function at all after you cross a trivial threshold. The default page size is what, 100? It doesn't take much.

The real cost isn't the loop, it's the memory footprint. Pulling thousands of hosts into a slice for comparison can bloat your service if you're not careful. You can mitigate that by streaming the comparison or working with smaller batches, but you can't skip the pagination fetch.



   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

You're right to be suspicious of a single `DesiredState`. Most CMDBs don't work that cleanly. The value of declaring it isn't the field itself, it's forcing you to define the business logic for what that state *means* before you write a line of sync code.

On the external ID, I've seen duplicates happen not from a missing ID, but from one that can change. If your CMDB's unique key for a server can be reassigned during a reprovisioning cycle, you'll get the same mess. The ID needs to be immutable for the host's lifecycle.



   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

That's the real world. We built a mapping where "DesiredState" boiled down to a ruleset: CMDB status=Active AND maintenance flag=false AND decommission date null. The logic sat in the translation layer, exactly as described.

On the immutable external ID, I've had to push back on teams using CMDB-generated IDs that got recycled. The fix was demanding a truly persistent, unchanging asset tag or UUID from the source system's data owners. If it can change, it's not an ID.


Where is your SOC 2?


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Missing the pagination loop there. That code will break silently once you pass the default page size. Add the loop before you do any state comparison.



   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Yeah, they mention it just breaks silently. That's the scariest part, because your script could seem fine for months until it suddenly isn't. Makes me wonder, are there any easy ways to test for this kind of pagination failure early on? Like, mocking a large dataset somehow?


Still learning.


   
ReplyQuote
Page 1 / 2