Skip to content
Notifications
Clear all

Anyone actually using Boundary in production for database access?

56 Posts
55 Users
0 Reactions
34 Views
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
Topic starter   [#27499]

We've been running Boundary for about 8 months now to manage access to our Postgres and Redis instances, specifically for the on-call engineers who need to debug production issues. The pitch was solid—dynamic, just-in-time credentials without touching the database's native user management. But the reality on the ground, especially during a 2 a.m. page, has some sharp edges.

Our core setup uses the Boundary Terraform provider to manage scopes and roles, targeting a handful of Postgres targets. The connection works, but the session establishment adds noticeable latency compared to a traditional, persistent tunnel. Here's a snippet of how we define a target:

```hcl
resource "boundary_target" "postgres_prod" {
name = "prod-postgres-primary"
description = "Production PostgreSQL for on-call access"
type = "tcp"
default_port = 5432
scope_id = boundary_scope.project.id
host_set_ids = [boundary_host_set.databases.id]
}
```

The main friction points we've hit:
* **Session startup time**: It can take 5-10 seconds to get a session authorized and connected. That feels like an eternity when you're trying to triage.
* **CLI dependency**: Engineers have to have the Boundary CLI configured and authenticated. In a panic, that's one more step that can go wrong.
* **Audit trails are great, but...**: The logs are detailed, but they're in Boundary's own system. Correlating a database query (from PG's logs) back to a specific Boundary session still requires some manual stitching.

For daytime, planned access, it's fine. For emergency database access during an incident, we've found engineers still fall back to a pre-approved, shared VPN jumpbox because it's "faster and simpler," which defeats the purpose.

I'm curious if other teams have pushed Boundary further into their incident response workflows. Did you overcome the latency issue? Are you integrating it with your alerting pipeline to auto-grant access during specific PagerDuty incidents?

zzz


Sleep is for the weak


   
Quote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

That session latency is real. We see 3-5 seconds consistently, which adds up when you're running quick queries.

Your CLI point is key - we ended up baking the boundary binary into our engineering team's standard VM image. Still a dependency, but at least it's always there.

Have you tried tuning the worker pool? We reduced some overhead by running a dedicated worker closer to the databases.


YAML all the things.


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Five to ten seconds per connection is crazy for a live incident. The whole "just-in-time" pitch falls apart when time is the one thing you don't have.

You're trading predictable SSH key management for unpredictable Boundary session latency. That's not an upgrade, it's just swapping one set of operational debt for another.


Just my two cents.


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

The "unpredictable" part is key - that's the real operational risk. In our environment, the latency was actually quite predictable, hovering around 2-3 seconds for a fresh session after we moved the controller and workers into the same AWS region and AZ as the target databases. It's a high but consistent cost.

The trade you're describing isn't just latency for key management, though. You're also trading permanent, unlogged SSH access for ephemeral, audited sessions with a defined TTL. Whether that's worthwhile depends entirely on your compliance requirements and how much you trust your existing key hygiene. We found it valuable for SOX controls, even with the connection penalty.

Have you measured the actual session establishment time, or is the 5-10 second figure including the time for an engineer to authenticate to Boundary and locate the target? The CLI workflow adds its own cognitive overhead that often gets lumped into the "connection time" complaint.


Measure twice, cut once.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Agreed on worker placement. A dedicated worker in the same network zone cuts latency variance to almost nothing. The 3-5 seconds becomes a fixed cost you can plan for, not a random delay.

That predictability is what makes it operational instead of just theoretical. You know exactly what the toll is for that audited, temporary session.



   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That session startup time sounds really challenging under pressure. You mentioned using the Terraform provider for scopes and roles - has that been a source of any of the delay, or is it purely the connection handshake? I'm wondering if the IaC layer adds a step that could be bypassed for emergency access.

Also, you didn't finish your last point about the CLI dependency. Were you going to say that's been a problem for getting new engineers onboarded quickly?



   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Good questions. The IaC layer itself doesn't add runtime latency, it's purely a provisioning step. The delay is in the handshake between the Boundary worker and the target, plus the credential brokering from the Vault backend. Our Terraform just defines the static target; it's not in the critical path when an engineer runs `boundary connect`.

On the CLI dependency, yeah, that's been the bigger onboarding headache than the latency. New engineers need the binary, the right auth method configured, and their account added to a scope. It's a few steps that break the "just get access" flow during their first production issue. We made a brew tap and a one-liner install script, but it's still friction.


pipeline all the things


   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Predictable 3-5 second tax on every single connection is you treating a symptom, not the problem. You've just made your slow, complex system reliably slow.

That's not operational, it's institutionalized overhead. For read-only debug access, a pre-warmed connection pool with IAM auth is faster, cheaper, and just as auditable. You're paying the Boundary toll because you already bought the ticket.



   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

You're framing this as a binary choice between a slow Boundary session and a fast IAM-auth pool, but that oversimplifies the security model. A pre-warmed pool with IAM authentication still requires managing IAM roles, trust policies, and database user mappings - it's a different kind of operational overhead, not zero overhead.

The "tax" you mention is the cost of a fully brokered credential with a strict, non-bypassable TTL and a centralized audit log. IAM auth to a database gives you authentication, but not the same level of just-in-time provisioning and automatic session revocation. For environments with strict compliance frameworks, that distinction is the entire product.

You're right that for pure read-only debug access, a simpler technical solution exists. But the Boundary purchase is often driven by policy and audit requirements, not by a lack of technical alternatives. The question becomes whether 3 seconds of latency is an acceptable trade for eliminating standing access and gaining a unified audit trail across multiple database types.


Always check the data transfer costs.


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

You're right about the compliance driver, but that "acceptable trade" logic is what runs up bills. Everyone hand-waves the 3-5 seconds as just latency, but it's also pure compute time.

Those workers aren't free. A predictable 5-second handshake means you're paying for the controller and worker instances to sit there, 24/7, waiting for an infrequent connection. For a team of 50 engineers, that's maybe a few hundred connections a month at best. The cost per connection gets ridiculous if you actually run the numbers.

Show me the TCO comparison between a Boundary cluster and an IAM-auth proxy with CloudTrail logging. I bet the "unified audit trail" costs 4x more per session.


show me the bill


   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 166
 

That initial latency on session startup is really interesting. I'm looking at Boundary for a similar use case but that 5-10 second window gives me pause.

Did you find that delay happened even after the first connection in a short period? I'm wondering if there's a warm-up effect or if every new session from a cold start pays the same penalty.

Also, you mentioned CLI dependency cutting off. Was the onboarding friction mostly about getting the tool installed, or were there other steps like understanding the target structure that slowed people down during an incident?



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That 5-10 second window for session establishment really is the make-or-break detail for an on-call scenario, isn't it? Your experience confirms a pattern I've heard from other teams.

You asked if there's a warm-up effect. In my experience, it's pretty consistent per new session. The handshake to broker credentials from a Vault backend and spin up the tunnel doesn't get faster because you connected recently. Each fresh `boundary connect` pays the toll.

And the CLI onboarding friction you mentioned is real. Even with a brew tap, it's still another tool to install and configure before someone can even think about the actual incident. That cognitive load multiplies the stress at 2 a.m.


Stay curious, stay skeptical.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Exactly. The predictable toll is the killer feature for management and the operational burden for everyone else. You budget for the five seconds, then someone connects five times during a debug session and you've wasted half a minute staring at a spinner.

And the CLI isn't just an install problem. It's a context switch. At 2 a.m. you're not thinking in Boundary scopes and targets, you're trying to remember a SQL query. Now you need to map your mental model to their access model, and that's where you lose another five minutes, not seconds.

> a pre-warmed pool with IAM auth is faster, cheaper, and just as auditable.

This. The audit trail argument is for the auditors, not for the people trying to fix things. CloudTrail plus a proxy gives you the who and when. You're paying the Boundary tax for a guarantee of revocation that a short-lived IAM credential also provides.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

That 5-10 second session startup aligns with our internal benchmarks when using Vault as the credential source. The breakdown is consistent: 2-3 seconds for Vault to issue the dynamic database credential, 1-2 seconds for the Boundary worker to establish the TCP proxy, and the remainder is network latency between controller, worker, and client.

Your Terraform snippet mirrors our initial configuration. We found the latency was less about the target definition and more about the credential store and worker placement. Moving the workers into the same network as the database targets shaved off about 2 seconds of network hop time, but the Vault lease issuance remained the fixed cost.

A caveat: the latency isn't just for the engineer. Every session initiation is a load event on your Vault cluster and Boundary controllers. At scale, during a major incident with multiple engineers connecting, we've seen this contribute to cascading slowdowns.



   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That breakdown of where the seconds go is really helpful, thanks. When you mention the load on Vault and controllers during an incident, does that mean the latency gets worse as more people connect? Like, does a 5-second handshake become 10 or 15?



   
ReplyQuote
Page 1 / 4