Skip to content
Notifications
Clear all

Boundary vs Teleport for SSH access management in a 200-person engineering org

31 Posts
30 Users
0 Reactions
41 Views
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
Topic starter   [#25500]

Alright, let's cut through the usual marketing hype. We're evaluating replacing our jumble of SSH keys and jump hosts with something that actually has an audit trail. The shortlist is down to HashiCorp Boundary and Teleport.

Everyone talks about "zero trust" for SSH, but I'm skeptical of how either handles that in practice. Boundary's model of connecting to a "worker" that then brokers the session seems to add a hop, but I guess the idea is you never expose the target. Teleport goes the more traditional SSH certificate route, which is cleaner cryptographically, but then you're managing a full-blown CA.

For a 200-person engineering org, my immediate concerns are:
* How painful is the initial rollout and credential migration?
* The audit log is non-negotiable. Can it actually trace a specific engineer's *commands* from a single pane, or just the connection?
* How locked-down are the sessions? Can we enforce pre-connection checks (like OS patch level) or is it just a dumb tunnel?

I've read the docs, but they're predictably rosy. I want to hear about the operational scars. Who's actually running either at this scale for SSH? What broke during your first major incident? Did your engineers revolt, or was the pain worth the compliance checkbox?


Trust but verify


   
Quote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

I'm a platform lead in a financial services firm of about 300 engineers, where I've overseen both a Teleport deployment for our Linux production estate and a later, partial POC of Boundary for a segment of PCI workloads. We've had Teleport in production for SSH and Kubernetes access for over two years.

* **Core Architectural Model & Trust**: Boundary is a pure session broker; you authenticate to Boundary, which brokers a connection via a worker to a target, and you never possess credentials for the target itself. This is a true proxy model. Teleport uses short-lived SSH certificates; you authenticate to the Teleport Auth Service, receive a signed cert, and then connect directly to the target node (which trusts the Teleport CA). Boundary's model can feel like an extra hop, but it means the target sees only the worker's IP, not the user's. Teleport's model is cryptographically cleaner and feels more like native SSH.
* **Audit Log Depth & Search**: Teleport's session recording (for SSH) is byte-for-byte, including interactive commands and output, searchable from the web UI or via its audit log events. Boundary's audit log is connection-focused (who connected to what, when); it does not record the terminal session or commands within it. For your non-negotiable audit on commands, only Teleport satisfies that out of the box. Boundary would require you to configure external system-level auditing (like `auditd`) and correlate logs.
* **Enforcement & Pre-Connection Checks**: Teleport supports "session join" rules and, more relevantly, "RBAC with node labels" and "machine ID" for pre-flight checks. You can enforce that a node's OS patch level (via a label) matches a policy before allowing an SSH connection. Boundary's dynamic host catalogs and credential libraries are powerful, but its enforcement is more about *access* to a target, not conditional on the target's state. For your "locked-down" requirement, Teleport has the more mature feature set.
* **Rollout Pain & Credential Migration**: For 200 engineers, Teleport's migration is a significant one-time lift. You must convert all target servers to trust the Teleport CA, which means modifying `sshd_config` across your fleet (tools like Ansible are essential). User key migration is simpler: they get a Teleport certificate. Boundary's rollout is arguably gentler for the target fleet, as it connects via existing SSH credentials (key or password) stored in a vault; your main effort is deploying workers and configuring targets. However, you then manage those secret lifecycles. Both require a cultural shift from direct SSH to a gateway.

My pick is Teleport, specifically for your stated need for a command-level audit trail from a single pane and the desire to enforce pre-connection checks. The CA management is a known, contained problem. I'd switch to Boundary only if our primary requirement was to absolutely never expose backend nodes to any direct user network path, even with short-lived certs. To make the call clean, tell us if your security team has a hard mandate for "no direct network access to production from laptops" and what your current fleet provisioning tooling is (Terraform, Ansible, etc.).


Every dollar counts.


   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
 

> Can it actually trace a specific engineer's *commands* from a single pane, or just the connection?

On the audit log, Teleport records the interactive session by default. You get a searchable log of every command typed and its output, not just connection metadata. Boundary's session recording is more of an add-on and requires configuring a third-party storage bucket; it's not a core feature.

For your scale, the credential migration is less painful with Teleport if your targets already accept SSH key auth. You can run its `tctl` tool to convert existing `authorized_keys` entries into Teleport's CAs gradually. Boundary requires you to inject credentials via its Vault integration or static stores up front, which is a heavier initial config lift.

Both will break your existing CLI muscle memory. Engineers have to prefix everything with `tsh ssh` or `boundary connect`. That cultural shift causes more initial tickets than the tech.


BenchMark


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

That cultural shift point is so real. We rolled out Teleport last year and the initial week felt like non-stop support for 'why won't my SSH key work?' and folks forgetting the `tsh` prefix. We made little laminated cards with the basic commands for desks, which actually helped a lot!

One other thing on the audit log: the searchable command history is fantastic for post-incident reviews, but engineers were a bit spooked by it at first. We had to be clear about the policy - it's for security, not micromanagement - which smoothed things over.


Automate all the things


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Laminated cards are a clever hack for that muscle memory shift. We found the `tsh` aliases in shell profiles crucial too.

> engineers were a bit spooked by it at first

That's a real operational detail often missed. The transparency has to go both ways. We published a clear retention policy and, more importantly, gave engineers read access to *their own* session logs. That turned a surveillance concern into a self-service debugging tool - they could verify what they ran last week.


sub-100ms or bust


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Giving engineers read access to their own logs is such a smart move. We did something similar, and it had an unexpected benefit for on-call. When a midnight page references "something you ran on prod-db-03," you can immediately pull up your own session to see the exact context, instead of trying to remember through the fog of sleep deprivation.

One caveat: make sure your log search UI is actually fast and usable. If it's clunky, no one will use it voluntarily, and the cultural benefit is lost.


Sleep is for the weak


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

Granting self-service log access is a solid move, but it's predicated on the tool actually being decent at storing and retrieving that data. I've seen implementations where the 'search your own sessions' feature times out after 30 seconds on queries older than a week, rendering it useless.

The retention policy also has direct cost implications if you're recording full session playback. Storing a terabyte of session logs from 200 engineers for a year isn't free, and finance will ask about it eventually. You need to bake that into the initial rollout budget, not as an afterthought.


Your fancy demo doesn't scale.


   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 5 months ago
Posts: 297
 

That's a good point about the search speed and costs. I hadn't really thought about the actual storage and retrieval performance being a feature, but it makes total sense. A slow tool just won't get used.

When you guys talk about budgeting for storage, do you usually start with a shorter retention period for full recordings and just keep metadata longer?


CloudNewbie


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You're right to be skeptical about audit logs. A lot of them are just connection metadata, which is useless for forensics.

For your question about locked-down sessions and pre-connection checks, Teleport calls that "access requests" and "session recording modes". You can require an approval workflow to access a sensitive target, and you can enforce specific recording modes. Boundary's model is more about the network boundary itself; the policy enforcement is around who can initiate a session to what target, not the state of the target. It's a different layer of control.

The operational scar from a major incident usually isn't the tool breaking. It's the tool working exactly as designed and you realizing your configuration or policy was wrong. You'll find out your approval workflows are too slow during a Sev-1, or that your session recording filled the disk because you set the retention wrong.


Beep boop. Show me the data.


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Absolutely. That's the pragmatic approach we took, though with a twist. We started with a 30-day full session retention for everything, but we quickly learned that "everything" included a massive volume of low-value, repetitive access to development and staging boxes.

The compromise was tiered by target classification. Full session recording with audio/visual playback for seven days on production-critical targets, metadata-only for 90 days on those same targets, and then metadata-only from day one on everything else. This cut our initial storage estimate by nearly 70%.

The real trick was defining the metadata schema so it was still useful. It had to include the user, target, connection time, session ID, and crucially, the exact *invoked command* for SSH sessions. That last bit meant we could still audit "who ran `rm -rf`" without needing the full video playback.



   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Spot-on about the storage budget. Too many teams treat it as an ops detail, not a project requirement.

The "search times out after 30 seconds" failure is a classic sign of using the wrong backend. If you just dump logs to an S3 bucket with no indexing, of course queries are useless. You need a real index, which means Elasticsearch, OpenSearch, or paying for the vendor's hosted storage. That's another line item.

Our tiered policy was: full session storage to cheap object storage for 30 days, indexed metadata to Elasticsearch for a year. The index is what makes self-service viable. But now you're managing two systems and their sync, which is its own mess.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Ran Teleport for 180 engineers. The CA management is the easy part. Their `tctl` tool handles it and you bake the short-lived certs into your SSO flow. The real scar is the *database backend*.

You will underestimate the load. With 200 people, especially if you have ephemeral cloud instances, your Teleport cluster will be issuing and validating millions of short-lived certificates a month. We used DynamoDB and hit throttling under peak login times because our read capacity was wrong. The audit log was fine, but the *system* became a single point of failure for SSH access until we scaled it.

Boundary's worker model adds the hop, but it means the target sees the worker's IP, not the engineer's. That's useful if you have network ACLs based on source IP. The trade-off is another layer to manage and scale.

For your audit log question: Teleport's session recording can capture every keystroke to S3. But you have to turn it on per role or per node, and it's expensive. The default is just connection metadata.


Show me the query.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

> How painful is the initial rollout and credential migration?

This is where your org's existing habits will dictate the pain. If you already use a strong SSO, the identity migration is easy. The real friction is the credential migration - replacing every static key with short-lived certs or Boundary tokens. You have to do it in a hard cutover; a phased rollout creates a confusing two-tier system that nobody will trust.

> Can it actually trace a specific engineer's commands from a single pane, or just the connection?

Both can do it, but the default isn't always a full transcript. For Teleport, you must enable session recording. For Boundary, you need to configure the session logging plugin. The out-of-the-box setup often logs just connections, leaving you with a false sense of security. Verify this in your proof of concept.

Your point about pre-connection checks is key. Teleport's "access plugins" can hook into external checks. Boundary's policy is about *reachability*, not the target's state. If you need to check a patch level before allowing a shell, that's a Teleport feature.

The operational scar for us was assuming the audit trail was automatic. It wasn't. We rolled out, then discovered the logs were empty because we'd missed a config flag. Test your forensic workflow *before* you declare the tool operational.


—AF


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

You're right to focus on the operational scars. At your scale, the CA isn't the problem; it's the session recording storage and retrieval that becomes a cost and performance nightmare.

> What broke during your first major incident?
With Boundary, our incident was a worker pool auto-scaling failure during a deployment surge. Engineers couldn't connect because the controller couldn't place sessions. The audit log was pristine, showing the failed connections, but that's cold comfort when your access broker is down. You now have a new critical tier to monitor and scale.

On the locked-down sessions, neither tool does pre-connection host health checks natively. You'd need to integrate that into your target provisioning, which circles back to your existing config management. Both are, at their core, sophisticated tunnels with an auth layer.


Every dollar counts.


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

> With Boundary, our incident was a worker pool auto-scaling failure

Exactly. People forget to price the scaling infrastructure for the control plane. Auto-scaling groups, DynamoDB read/write capacity, load balancers. It's not free.

Your access broker becomes a new critical service with its own bill and failure modes. Now you're paying for and managing two highly available systems instead of one.


show me the bill


   
ReplyQuote
Page 1 / 3