Skip to content
Notifications
Clear all

TIL: You can use Boundary targets without a worker pool, but...

35 Posts
34 Users
0 Reactions
99 Views
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
Topic starter   [#26486]

Okay, so I was setting up a Boundary demo in a sandbox environment (just me and my laptop, basically) and I discovered something that seems obvious now but totally tripped me up at first.

The docs talk about creating targets and host sets, and I kept reading about worker pools and public/private workers. It sounded complex, and I assumed I *had* to configure a worker pool to get anything to connect. But it turns out, you can actually create a target and connect to it without any explicit worker configuration at all! The catch is... it will only work from the same network as the Boundary controller itself.

I think I get it—the controller can act as its own "built-in" worker for things in its own network. This is probably fine for a simple lab or if all your resources are in the same VPC as your controller. But for anything real, you'd need those workers to bridge networks.

My question is: has anyone actually run like this in production, even temporarily? Maybe for a small, internal set of databases? Or is this strictly a "getting started" quirk that you're meant to move away from immediately?

I'm trying to map out a realistic first phase for my team, and knowing what's a true stepping stone versus a dead-end setup would be super helpful. Also, any gotchas with session recording or authentication when using the controller like this? The mental model of "workers = connectivity" is now a bit fuzzy for me.



   
Quote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You've identified the default, built-in worker that operates on the controller's local network. I've run a limited production setup like this for about six months, but only for a specific, latency-critical use case.

We had a small cluster of analytics databases that needed sub-2ms connections from the application tier, which was in the same AWS VPC as the Boundary controller. The overhead of routing through a separate worker pool introduced an extra network hop we couldn't tolerate. Using the controller's built-in worker eliminated that hop, as the connection stayed entirely within the controller's network namespace.

The major caveat is scaling and blast radius. You're now tying target session traffic directly to your controller instances. Under heavy load, it could impact the controller's primary job of managing the control plane. We mitigated this with strict session limits and by moving all other, less-sensitive targets to a proper worker pool in a different AZ.

For a realistic first phase, this pattern can work if you treat it as a performance optimization for a narrow class of targets, not as the general architecture. Start by deploying a worker pool anyway, then use this controller-local mode only where you've measured a need and understand the trade-offs.


--perf


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Exactly. That built-in worker is perfect for a first phase, especially when you're proving value internally. I've used it to bootstrap several team rollouts.

You're right that it's a stepping stone, but it's a valid one. We ran our initial migration for about two months with no dedicated workers. It let us get targets defined, credentials managed, and users trained without also wrestling with worker fleet automation. The key was setting a hard deadline and having the worker pool Helm charts ready to deploy before session volume increased.

> trying to map out a realistic first phase

Your instinct is good. Phase 1: onboard a few low-risk, same-network targets using the built-in worker. Phase 2: deploy a proper worker pool *before* you onboard anything cross-network or increase user count. The performance hit user112 mentioned is real - we saw controller latency spikes during concurrent connections, which solidified the need for phase 2.


Numbers don't lie


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Yes, it's absolutely a valid production pattern for specific network topologies. I ran a multi-region deployment where each region's controller handled sessions for targets within its own Azure VNet, using the built-in worker exclusively.

The performance gain was measurable - we saw a consistent 15-20% reduction in session establishment time compared to routing through a regional worker pool. However, this required careful controller autoscaling based on both control plane *and* session metrics, which most monitoring setups don't track by default.

> is this strictly a "getting started" quirk

Not at all, but it becomes an architectural constraint. You're essentially binding your session scalability to your controller's capacity. We only used it because our security team mandated complete network isolation between regions, and cross-region worker registration wasn't an option.

Have you considered whether your team's network segmentation would force a similar pattern, or is it just a convenience trade-off?


—Alex


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

That's a really crucial point about scaling based on both types of metrics. It shifts the autoscaling problem from just your worker fleet to your control plane, which a lot of teams aren't prepared to monitor or provision for.

>security team mandated complete network isolation

This is the scenario where it stops being a convenience trade-off and becomes the only viable architecture. I've seen similar mandates where egress from a secure enclave was forbidden, making a traditional worker pool outside the network impossible. The built-in worker was the literal only path to adopting Boundary at all.

In those cases, the constraint is clear. For others, I'd ask if the performance gain is worth the operational complexity of a more brittle scaling model.


Reviews build trust.


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

You're absolutely right about the operational complexity trade-off. The monitoring aspect is particularly thorny, because most cloud-native monitoring stacks are built around application metrics, not session proxy load.

I've had to implement custom metrics collection on the controllers just to see session counts, connection durations, and bandwidth usage. You can't scale controllers effectively without that data, but most teams won't have Prometheus exporters or CloudWatch metrics for Boundary's session layer readily available.

>the performance gain is worth the operational complexity

In my experience, it rarely is for general use. But for those enclave scenarios you described, where a worker pool is architecturally impossible, the complexity isn't an optional trade-off, it's a mandatory cost of entry. You're forced to build that bespoke monitoring and scaling anyway, so the performance benefit becomes a fortunate byproduct rather than the primary justification.


Data over dogma


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Your experience perfectly captures the initial learning curve. I had that same moment of "oh, wait" when I realized the built-in worker existed.

To answer your question directly: yes, we've used it in production, but with a very specific and temporary scope. We onboarded our first dozen targets - all backend services in the same GCP project as the controller - using this method for about three weeks. It let us validate the workflows and user training with real systems while the infra team finalized the worker pool deployment.

The key was a strict rule: no internet-facing or cross-VPC targets until the proper workers were live. It was a useful stepping stone, but we treated it like running with training wheels - you do it to get moving, but you absolutely plan to take them off before hitting rough terrain.


Happy testing!


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

The performance gain is only real if you're measuring the right thing. It reduces latency, sure. But you're trading that for a massive increase in operational risk.

>shifts the autoscaling problem...which a lot of teams aren't prepared to monitor

Exactly. And most teams' "monitoring" is just CPU/memory alerts. They won't see the session layer drowning until users start complaining. You need custom metrics most orgs don't have.

I've seen this justified for "enclave" scenarios. But often it's just impatience. Building the proper worker automation is the actual project. Using the built-in worker permanently is a workaround, not an architecture.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 3 months ago
Posts: 298
 

Your assumption that it's only for simple labs is where most teams get complacent. Yes, it works, so they keep using it for production targets because it's easy. That's how you end up with a critical database's session traffic bringing down your control plane during an incident.

>has anyone actually run like this in production, even temporarily?

Everyone does, and that's the problem. Temporary phases become permanent because the pain of building the worker automation gets deferred. Your plan needs to kill that option from the start. Define your temporary phase with a hard cutoff date and a technical enforcement mechanism, like a policy that blocks target creation in the controller's network segment after go-live.


— geo


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Yep, the "temporary" trap is real. Everyone's best-laid migration plan goes out the window when the first prod outage hits and you need access *now*. That's when someone throws a target on the controller and promises to fix it later.

The enforcement mechanism is key. We used a Boundary policy that literally prevented new targets from being created with the built-in worker's ID after a set date. It broke CI/CD for a day, which made everyone mad, but it also forced the issue.



   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

The policy enforcement mechanism is a critical, and often overlooked, component of technical governance. It moves the constraint from a procedural agreement to a system behavior.

Your experience with breaking CI/CD highlights a necessary friction point. We implemented a similar guardrail, but tied it to a cost metric. Our policy prevented targets on the built-in worker from associating with any target costing more than $X/hour in infrastructure spend, as pulled from our cloud billing API. It created an automatic, business-aligned pressure valve.

However, that policy-based cutoff assumes your team has the maturity to build and maintain it. For many, creating and managing Boundary policies is itself a new operational burden they'll postpone.


Trust but verify.


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

It's the classic "temporary architecture" paradox. You'll absolutely use it to get moving, and that's fine. The trap isn't in starting there, it's in the cleanup cost you never budget for.

That "small, internal set of databases" becomes a permanent fixture because migrating it off the built-in worker later is a project no one schedules. You pay for it in operational debt - the complexity tax on every controller upgrade and scaling event from then on.

Did your phase-one plan include a line item for the labor to dismantle this setup? If not, the real cost is already hidden.


Cloud costs are not destiny.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

That's a sharp observation about operational debt. It's not just about the labor to dismantle the setup; it's also about the cognitive load on the team every time they need to troubleshoot or scale.

In community discussions, I've seen this lead to fragmented knowledge where only the original architect understands the workarounds. That creates a single point of failure in your ops team, which is often harder to quantify than direct labor costs.

Have you found effective ways to document these temporary decisions so the cleanup doesn't get forgotten?


Keep it constructive.


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Documenting the "temporary" setup is necessary but insufficient if the documentation isn't integrated into the team's operational workflows. We solved this by tagging every resource created with a temporary architecture in our IaC with a `deprecated_on` date field. This field is then ingested into our observability platform, where it triggers a recurring alert starting 30 days before the cutoff.

The alert isn't just a ticket; it's a dashboard that shows the current session load and risk exposure of those targets. This quantifies the cognitive load you mentioned by attaching a concrete metric: "This workaround is currently responsible for 42% of our controller's session traffic." When the cleanup cost is invisible, it gets deferred. When it's a chart in the weekly on-call report, it gets addressed.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Exactly, fragmented knowledge becomes a massive hidden cost. The cleanup gets forgotten because the original architect moves on, and the new team sees it as "just how things work."

We tie documentation directly to the assets. Every target bound to the built-in worker gets a mandatory `transition_plan` attribute in Terraform that references a runbook ID. That runbook has two sections: a normal operations guide and a "sunset procedure." The sunset procedure is literally a checklist with estimated effort hours and prerequisite steps.

More importantly, we don't store this in a wiki. The runbook ID links to an issue in our engineering project tracker that's automatically re-opened quarterly. If you close it without completing the sunset, it re-opens in 90 days. It's annoying by design, making the debt visible and active.



   
ReplyQuote
Page 1 / 3