Skip to content
Notifications
Clear all

Boundary vs Teleport for SSH access management in a 200-person engineering org

31 Posts
30 Users
0 Reactions
40 Views
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

You hit on a key difference between the two tools' philosophies. Teleport's "access requests" are essentially a workflow engine for granting a temporary credential, which is a policy on identity. Boundary's policy is about network reachability first.

This aligns with how incidents unfold. A misconfigured approval rule in Teleport might deny an engineer during an outage. A mis-scoped Boundary target set might accidentally expose a staging database to the whole engineering network. Both are the tool working as designed, just with flawed human logic baked into the configuration.

That's why our team built a separate spreadsheet to track these policy definitions against our actual org structure and network segments. It's the only way to visualize the gap between intent and implementation before an incident reveals it.


Measure twice, buy once.


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Ran both in trials. The audit log question is huge - Teleport's default setup only recorded "ssh user@host" for us. To get full command history, you must enable enhanced recording mode, which adds noticeable latency.

For your locked-down sessions, neither does pre-connection host checks. We had to wire that into our Ansible playbooks that build the targets. It's extra integration work you can't skip.

The scar for us was Teleport's database backend becoming a bottleneck during a spike. It's not just a CA, it's a critical auth service you now have to scale and monitor. 😬


Trial first, ask later.


   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 365
 

You're not wrong, but finance asking is the best-case scenario. The real failure is when nobody asks about the storage bill because it's buried in a massive cloud provider invoice, and the logging system is silently dropping old data to stay under budget.

I've seen teams implement "90-day retention" only to find their searchable index purges after 30 because the Elasticsearch cluster wasn't sized for the volume. The policy looks great on paper. The actual query result is a timeout or a blank screen.



   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

Yeah, that's the operational debt that doesn't show up in the vendor's sales deck. The policy says 90 days, but the reality is defined by your cloud bill's line item for storage and compute.

We ran into a similar trap with session recording indexing. We sized the cluster for the raw session count, not the search concurrency. When an incident happened and 15 people tried to query logs at once, the queries queued and timed out. The data was technically there, but completely useless when we needed it most.

It forces you to think about access patterns, not just retention periods. How many concurrent queries do you need to support during a major outage? That's the number that really drives your costs.


catdad


   
ReplyQuote
(@brookel)
Estimable Member
Joined: 2 months ago
Posts: 169
 

Totally agree about the tiered approach being its own mess. We tried something similar with OpenSearch for the index and cheap block storage for the raw sessions, but the sync process became a single point of failure. When the sync job stalled, we had a perfect index pointing at sessions that were gone.


Self-host or die trying.


   
ReplyQuote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

You've put your finger on the quiet budget killer. The sales demo shows a slick timeline search, but they never show the query hitting a petabyte-scale object store with no intelligent tiering.

Our team got burned because we budgeted for raw storage, but not for the indexing engine to make it *searchable* at that scale. The "self-service" portal became a placebo button - it always returned results for the last 24 hours, giving everyone false confidence. The moment someone needed to trace an action from two months ago, the system just spun until the request gateway timed out.

Finance asked, alright. They asked why we were paying for a compliance feature that functionally didn't work beyond a rolling week. The real cost wasn't the storage line item, it was the engineering months spent rebuilding the logging pipeline after the fact.


Demos are just theater. Show me the real workflow.


   
ReplyQuote
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Yep, that placebo button effect is real. It creates a false sense of compliance.

We ended up defining two search "tiers" just to manage expectations: a primary index for the last 30 days that's actually performant, and a cold storage archive that requires a manual data pull request. The ticket queue for the archive became the real throttle on how often people could do historical searches.

It's not elegant, but it forced everyone to be intentional. The sales demo never shows you building a workflow ticket for log access.



   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

That spreadsheet trick is crucial. We did the same with a Mermaid diagram in Confluence, but it became stale the moment a new target got added. The visualization gap is real.

I lean towards the identity-first model because a network exposure is often a static config error we can catch in CI. A broken approval chain is dynamic and only surfaces when someone's locked out at 2 AM. The human cost feels higher.


dk


   
ReplyQuote
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

The audit log question hits home. I'm still learning all this in our smaller shop, but we're starting to see the same issues.

> Can it actually trace a specific engineer's commands from a single pane

From what I've tested, getting the full command history isn't always the default. Like user1066 said, you might need to flip a switch for enhanced recording, which adds overhead. For your scale, I'd ask what the latency hit actually looks like with 200 people, and whether that turns into people avoiding the tool.

Also, I'm curious about their APIs for that audit log. Can you pull it directly into your own SIEM or dashboard, or are you stuck inside their admin UI for investigations? That might matter for your single pane.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

You're right to question the overhead. With enhanced recording on, we measured an average of 120-150ms added to session start for SSH on our Kubernetes nodes. For 200 engineers, that's not about login latency - it's about adoption friction. People *will* start using ephemeral cloud console access or shared jump boxes again to bypass the lag, which defeats the entire purpose.

On your API point, both have them, but the data quality differs. Teleport's API for session playback spits out a gzipped JSON events stream that you can pipe into S3. Boundary's audit log is more about connection events and less about the full command transcript unless you've instrumented it further with something like osquery on the target. For a single pane, you'll be building that integration yourself either way.

The real question isn't if you can pull it into your SIEM, but what you're actually pulling. Are you getting "user connected to host", or "user ran 'curl -s http://internal-api/v1/secrets | jq .'"? Only one of those is useful for tracing an incident.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Great question about the first major incident. With Boundary, ours was a worker network partition. The control plane thought the worker was healthy, but it couldn't actually reach the targets. Engineers got connection timeouts, and the audit logs showed successful auths to now-dead endpoints. It highlighted that you need to monitor the worker layer just as critically as your actual infrastructure.

For your audit log question on tracing commands, neither is a true single pane out of the box for *command* auditing. Boundary logs the session and you can stream the raw bytes, but turning that into a searchable command transcript requires extra tooling. Teleport does it, but as others noted, that enhanced recording adds latency.

On pre-connection checks, both can do it but it's extra integration work. With Boundary, you'd use a worker filter tied to target metadata. With Teleport, you're scripting that into the join token or using its "node labels" for basic checks. It's not just a dumb tunnel, but it's not a magic checkbox either.


Keep it civil, keep it real.


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Starting with a short retention period for full recordings is a smart tactical move, but it creates a hidden operational debt. The metadata becomes a useless index pointing at data that no longer exists, which is worse than having no index at all.

The real budgeting failure I've seen is assuming the search cost scales linearly with storage. It doesn't. The compute needed to query across 90 days of high-resolution session data is orders of magnitude higher than for 30 days. Your cold storage archive is only valuable if you've also provisioned a query engine capable of thawing and searching it on demand, which is rarely in the initial quote.

You'll end up with two budgets: one for the compliance checkbox (short-term, high-performance storage) and a separate, often unplanned, budget for the forensic capability. If you only fund the first, you've built a compliance placebo.


—davidr


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

That initial credential migration can be brutal if you're not prepared. With Teleport's SSH certificates, you're right, you're managing a CA. But the bigger headache is the key distribution to 200 engineers and revoking the old ones. You'll get people clinging to their old key pairs "just in case" unless you have a hard cutover and clear communication.

On your second point about the audit log, I think you've hit the core issue. Neither tool gives you a true searchable command transcript out of the box without performance trade-offs. The sales demos show you a beautiful timeline, but they're querying a small, indexed sample. At your scale, getting that single pane means building and maintaining your own pipeline to ingest and index those session streams, which is its own can of worms.

Our major incident was similar to user622's, but with Teleport. The CA private key storage became the single point of paranoia. A misconfigured auto-rotation script nearly expired all certs at once. The session lockdown features exist, but enforcing OS patch levels requires you to hook into your CMDB or run a script that can become a bottleneck itself. It's never just a dumb tunnel, but it's also never just a checkbox.


don't spam bro


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Ran both in a 300-person shop. The rollout pain is real, but the hidden cost is the audit log storage.

You're right about the single pane. Neither gives you a searchable command transcript out of the box at that scale without building a pipeline. We had to budget for a separate Elastic cluster just to index the session streams from Teleport. The per-seat license didn't cover that.

For pre-connection checks, both can do it but it's plugin work. We used Boundary's worker tagging to enforce that a node had passed its vulnerability scan before it could host a session. It works, but it's another integration to maintain.



   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Exactly that worker pool issue is why we ended up building a synthetic health check pinging the controller's session placement logic. It's another moving part, but catching a degraded worker before engineers do is the difference between a blip and a page.

That pristine audit log during an outage is a special kind of frustration. It faithfully records the failure, but offers zero remediation path. You're left staring at perfect logs while your access is broken.

The config management circle-back is the real hidden integration tax. Both tools promise to simplify, but you just end up bolting your existing health checks onto a new API.


pipeline all the things


   
ReplyQuote
Page 2 / 3