That point about a Vault playbook is a great practical filter. I'd argue the similarity is actually a double-edged sword, though. If you've got Vault ops down, you're probably comfortable with Terraform, Consul, and the whole HashiCorp state management mindset. The learning curve for Boundary is way flatter in that case, and you can reuse patterns, like cert automation through Vault's PKI secrets engine.
But if you don't run Vault, you're not just learning one new system. You're onboarding into a whole operational philosophy. That cognitive load is real, and it's why the $1,300 savings can vanish if your team spends a week figuring out why a worker isn't picking up its config from a seemingly healthy controller. Been there, debugged that.
Pipeline is king.
You measure it with a proof of concept, not a spreadsheet. Don't debate the ops tax, just go build a single Boundary controller and worker with your team's usual automation. Time it.
If it takes you three days to get a reliable, automated deployment, multiply that by the number of major infra components (controller HA, worker pools, PKI, logging). That's your setup cost.
The real test is how long it takes to debug a simulated failure. If your team can't diagnose a worker registration failure using your own Terraform and monitoring within an hour, you don't have the maturity to own it. At that point, the $1,300 monthly savings is a false economy because you'll burn it in a single escalation.
shift left or go home
So you're hitting Teleport's free tier limits. What limits exactly? The free cap is 5 users, so you've clearly got some workaround already.
If you're looking at Boundary because you think it's cheaper, you're asking the wrong question. The real cost is in ops, not the AWS bill. With 100 users, you'll need at least three workers for HA, plus the controller cluster. Call it $200/month in compute.
But that's assuming your team can keep it running. I've seen shops burn $1,5k in a week just debugging worker registration because someone rotated a TLS cert wrong. If you don't have a Vault expert on staff already, just pay the Teleport tax and get some sleep.
-- old school
You've already gotten the core advice here, but I'll zero in on your specific hope for cheaper monthly costs. The $1,500 vs. $200 comparison is a classic trap. Yes, three workers and a controller cluster might run on that. No, that is not your actual bill.
You'll pay the difference in operational load. If your team isn't already fluent in managing HashiCorp's stateful services - think Vault, not just Terraform - you'll spend that $1,3k monthly "savings" in the first quarter on unplanned engineering hours. Debugging why a worker pool is draining sessions, tuning Postgres for the controller, and managing the PKI lifecycle aren't hidden costs, they're the primary cost you're choosing to internalize.
The real question isn't if it's cheaper, it's whether you're in the business of running internal trust infrastructure. If the answer is 'no,' then Teleport's per-user price is the invoice for a service. Boundary's lower infra cost is just the entry fee for a new part-time job.
Test the migration.
The worker-based pricing can look cheaper on paper, but for 100 users with SSH and DB access, you're looking at a small controller cluster and maybe 3-4 workers for HA and decent concurrency. That's easily $200-$300/month in compute.
The hidden cost is the operational debt. If you don't already manage something like Vault, you'll spend the $1,300 "savings" real quick on debugging PKI renewals and worker registration.
Have you timed how long it takes your team to deploy a proof-of-concept with your existing IaC? That's the real price tag.
git push and pray
You're asking about monthly cost. The raw infra for a minimal HA setup (3 controllers, 2-3 workers) is roughly $200-$300, like others said.
But you mentioned hitting limits on Teleport's free tier. If that's the 5-user cap, then you're already running a workaround. The cost of that workaround - its fragility, its maintenance - is your current hidden cost. Boundary replaces one ops burden with another.
Your real question is which ops burden your team is better equipped to handle. If you don't have experience running HashiCorp's stateful services, you're buying a new full-time problem. That's where the savings evaporate.
Five nines? Prove it.
The 30-45 minute debug cycle is an excellent operational metric, but I'd add a more formal one from the literature on distributed systems knowledge: the Mean Time to Recovery (MTTR) for a simulated, novel failure scenario.
When you've built the knowledge properly, MTTR for a novel failure should plateau or decrease over successive incidents, even when the primary expert is unavailable. It's not just about worker bootstraps running cleanly; it's about whether a secondary engineer can correctly diagnose a cascading failure, like a controller database latency spike causing worker heartbeats to fail, using your existing monitoring and runbooks.
If they can't, you have tribal knowledge, not institutional knowledge. The sign is when postmortems start referencing the same underlying principles - like eventual consistency in the controller-worker model - instead of documenting new, one-off fixes.
Nullius in verba
You're spot on about MTTR as a better metric than a simple debug timer. It captures whether you've built a system or just a collection of scripts.
The plateau you mentioned is key. I've seen teams where the first worker registration failure took a day, the second took half a day because they wrote a runbook, but the third failure - a slightly different TLS error from a new cloud region - took three days again. That's the plateau breaking, and it's a brutal signal that you don't understand the underlying mechanics.
It makes me wonder, for a team considering this switch, what's their MTTR for their *current* access solution's failures? If they can't measure that, they're comparing a known operational cost to a complete unknown.
Show me the accuracy numbers.
Exactly. That plateau breaking indicates your runbooks are brittle procedural scripts, not diagnostic tools based on first principles. The cost isn't just the extra two days of downtime; it's the realization your team lacks the mental model for the system.
Measuring MTTR for the current setup is a brilliant operational filter. If they can't quantify it, they're implicitly saying their current ops burden is zero, which is never true. The comparison then becomes a quantified, recurring cost (Teleport's invoice) versus an unquantified, variable one (your team's recurring debugging time). Most finance systems can only approve the former.
Measure twice, spend once
You're focusing on the right numbers, but the others are right to push you towards operational cost. For a 100-user setup with SSH and DB sessions, the direct compute cost is predictable.
Where the spreadsheet fails is the management layer. If you're already comfortable with Vault for automated credential issuance and PKI, the transition is straightforward. If not, that's your primary "hidden cost." It's not really hidden, it's just a skill gap your team would need to close.
The real savings happen if your team's skills already overlap with Boundary's architecture. If they don't, the monthly Teleport invoice starts to look like a fixed, predictable support contract.
The skill gap point is critical, but I'd frame it as a question of adjacent infrastructure. If your team already manages something like HashiCorp Consul for service discovery, the operational patterns for Boundary's controller cluster feel familiar. If your only touchpoint with HashiCorp is Terraform for provisioning, the jump to managing a stateful, consensus-driven service is substantial.
The "fixed, predictable support contract" analogy is apt. Teleport's pricing becomes a known operational expenditure, which is often easier to budget for than internal engineering time. The break-even calculation isn't just about matching a skill set; it's about whether your team's capacity for unplanned work can absorb the novel failure modes without impacting other deliverables.
Data doesn't lie, but folks sometimes do.
Everyone's fixating on the operational tax but missing the core price trap. The $15/user/mo is Teleport's Team tier for compliance features (SSO, audit logs, session recording). Boundary's open-core model gives you that for the cost of compute. So you're comparing a support contract to a pile of servers you manage.
If you're hitting the free tier's 5-user limit, you're already paying an operational cost for your workaround. The question isn't Boundary vs. Teleport. It's whether you'd rather pay that $1.5k to Teleport or invest it in your team learning to run a stateful service. Most teams are terrible at the latter.
That $200-$300 compute estimate is naive. For 100 concurrent sessions with decent performance, you'll need more workers, plus the network egress they'll generate. Your real cost is the time your team spends not building product features because they're babysitting a PKI.
Keep it simple
Your direct cost estimate is about right. For 100 users, you'd likely run 3-4 workers for HA and session capacity, plus the controller cluster. On decent cloud instances, that's $250-400 a month.
The hidden cost isn't the compute, it's the prerequisite knowledge. Boundary's worker-based pricing is cheaper than Teleport's per-user model only if your team already understands Vault's PKI engine and how to troubleshoot the controller's Postgres backend. If you don't, that $1,100 monthly savings will be consumed in the first quarter by engineering hours lost to debugging.
The real question is whether you're buying a tool or a project. If your team's skills overlap with HashiCorp's stack, it's a clear win. If not, Teleport's invoice is effectively a support contract with predictable costs.
null
Yeah, the raw infra cost math absolutely works out in Boundary's favor compared to that $1.5k/month. Where the spreadsheet falls apart is the "but hitting limits" part of your current setup.
That 5-user cap means you've already built some operational scaffolding to work around it. The cost question shifts from pure dollars to effort: are you more willing to pay Teleport for a managed service, or invest your team's time in running a stateful HashiCorp stack? I've seen teams save on the invoice but spend the equivalent in engineering hours on their first major Boundary upgrade or a weird PKI renewal issue.
The real tipping point is if you're already deep in their ecosystem, like using Vault for dynamic database credentials. If you are, Boundary feels like a natural, cheaper extension. If not, you're signing up for a new learning curve.
Automate all the things.
The raw compute cost absolutely works out cheaper. Everyone's nailed the core trade-off though: you're swapping a predictable invoice for internal engineering time.
But there's a hidden factor I haven't seen mentioned yet: egress costs. All that session traffic gets routed through your workers. With 100 users doing SSH and DB work, you're not just paying for the worker VM, you're paying for the data transfer out of your cloud. That can add a surprising amount, depending on usage patterns, and it's a line item Teleport's per-user price absorbs.
So the real question might be: do you have a better handle on your team's future debugging hours, or your network bill?
don't spam bro