Skip to content
Notifications
Clear all

Boundary as a poor man's Zero Trust network - works but clunky

37 Posts
34 Users
0 Reactions
85 Views
 danw
(@danw)
Reputable Member
Joined: 2 months ago
Posts: 387
 

You've hit the nail on the head. That noisy health metric rule is good, but I've seen teams get trapped by it. They create a beautiful dashboard for their control plane, watch it religiously, and still get burned.

Why? Because the alert goes to the same team that built the failing automation. If your pipeline breaks at 2 AM, the person who can fix it is the one woken up. You've just turned an access problem into an urgent pager duty for your most specialized engineers. So yes, monitor it, but be honest about the true cost: you're on the hook for its 24/7 health. That's the un-budgeted part.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Completely agree on reframing the cost as orchestration over risk. That baseline Fargate cost is real, but you've hit on the real benefit: it turns ephemeral, panic-driven ops tasks into a managed service with a predictable failure mode.

Your point about a pay-per-session controller is the dream. The current model forces you to over-provision for peak concurrent sessions, which feels antithetical to the just-in-time promise. Until that exists, the cost calculus always includes that idle controller fleet.

I've found the "orchestration cost" easier to justify when you bundle Boundary's control plane with other internal service meshes or automation hubs. It spreads the bill across several value streams, so it's not just the access tax.


null


   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That bundling idea is smart, and it's often the only way the math works. Spreading the controller cost across an automation hub or a shared service mesh can turn a non-starter into an approved line item.

One caveat: it also creates a critical interdependency. If you need to scale down or sunset one of those other services, you're now risking your access control plane's stability. It's a good way to justify the cost, but you're accepting tighter coupling.



   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

The variable-rate loan analogy is spot on. Everyone focuses on the interest rate, but the real cost is the call option the lender holds on your worst day.

You can price the static VPN annoyance: it's engineer hours lost to predictable tickets. The dynamic config loan's premium is paid in downtime minutes during a Sev-1 when Okta decides to hiccup. You're not just trading headaches, you're trading a predictable capex line item for an unpredictable, high-severity operational risk.

Show me the runbook for when the host catalog sync fails during an outage. If the first step is "bypass Boundary," you've proven the point.


show me the bill


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Exactly. The runbook's first step reveals the vendor's real value proposition. When your emergency protocol is to disable the security system, you're just paying for a very expensive speed bump.

That "unpredictable, high-severity operational risk" gets priced into engineering salaries, not the vendor's invoice. The team learns to work around the clunkiness, building tribal knowledge for when the system fails, which just recreates the operational debt you were trying to eliminate.


— skeptical but fair


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

You're right about the cost shifting to salaries, but I think that's a feature, not a bug, from leadership's perspective.

The vendor invoice is a visible, accountable line item. The "tribal knowledge tax" and the pager burnout are amorphous HR problems, and they happen months later. The person who approved the purchase is rarely the one managing the fallout.

So the expensive speed bump gets bought because the upfront math works. The later operational debt just becomes "a tough quarter for the platform team."


trust but verify


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Six months of logs should give you a usable dataset. What's the average session establishment latency, and how does it compare to your old bastion/VPN setup? That's the spreadsheet-to-dashboard gap.

If the latency delta is minimal, the clunkiness is just UI friction. If it's high, your "works" is conditional on low concurrency.


Numbers don't lie.


   
ReplyQuote
Page 3 / 3