Skip to content
Notifications
Clear all

Boundary vs StrongDM for zero-trust access in a multi-cloud finance firm

30 Posts
29 Users
0 Reactions
96 Views
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
Topic starter   [#23381]

Having recently completed a formal evaluation for a client in the financial services sector, I was tasked with architecting a zero-trust access model for their sensitive accounting and trading applications distributed across AWS, Azure, and a private data center. The core requirement was a single control plane for just-in-time, credential-less access to database clusters (PostgreSQL, Microsoft SQL Server) and legacy SSH bastions, with immutable audit logs for compliance (SOX, GDPR).

The shortlist, after initial vendor scoring, came down to HashiCorp Boundary and StrongDM. My methodology involved a two-week, side-by-side proof-of-concept structured around five core operational pillars:

* **Deployment and Administrative Overhead:** Boundary, being open-source with a self-managed model, required provisioning and maintaining its own worker nodes across each cloud. StrongDM’s fully managed control plane eliminated this infrastructure tax but introduced a recurring cost variable.
* **Target Discovery and Session Orchestration:** Boundary operates on a declarative model where targets (e.g., a database endpoint) are explicitly defined in HCL and managed via Terraform. StrongDM utilizes a dynamic discovery agent that automatically inventories network-accessible services, which reduced initial configuration but required careful scoping of the agent's permissions.
* **Identity Federation and Justification Workflow:** Both platforms integrated cleanly with our existing Okta instance for user authentication. However, StrongDM’s native support for access requests and time-bound approvals within its UI provided a more polished workflow out-of-the-box. Achieving similar in Boundary required integrating its API with a separate ticketing system (ServiceNow, in our case).
* **Protocol Depth and Credential Brokering:** For PostgreSQL, both solutions performed admirably, brokering ephemeral certificates. A key differentiator emerged with Microsoft SQL Server, where StrongDM’s proxy handled native Windows Authentication (Kerberos) delegation transparently. Simulating this with Boundary would have required custom development around its plugin framework.
* **Audit Log Fidelity and Extraction:** Boundary logs all session data (including recorded SSH) to its own storage, which then must be forwarded to a SIEM. StrongDM streams structured audit logs directly to an S3 bucket or Splunk, which the compliance team favored for its immutability and ease of ingestion.

The financial calculus became intriguing. Boundary’s open-source core presents a lower apparent entry cost, but the total cost of ownership must factor in the engineering hours for initial cluster setup, ongoing worker scaling, and developing any missing protocol logic. StrongDM’s per-user pricing is higher, yet it bundles the management overhead into the fee.

My preliminary conclusion leans toward StrongDM for this specific multi-cloud finance use case, primarily due to its managed service model, broader native protocol support for Windows-based assets, and the compliance team’s preference for its audit log pipeline. However, I am keen to hear from this community on long-term operational experiences.

Has anyone run a similar comparison in a regulated environment? I am particularly interested in:
* Real-world performance observations when scaling to thousands of concurrent sessions during peak trading hours.
* Experiences with automating target onboarding via Terraform for both platforms, especially when dealing with ephemeral Kubernetes workloads.
* The practical implications of Boundary’s requirement for a direct network path from workers to targets versus StrongDM’s relay architecture.



   
Quote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Head of security at a 200-person fintech. We run self-hosted Boundary to handle access to our PostgreSQL analytics clusters and internal Kubernetes API across AWS and GCP.

**Real Pricing:** StrongDM's quote started at ~$50/user/month and scaled from there. The cost was the main killer. Boundary is "free" but you pay for the compute to run controllers and workers, plus your own HA setup. At our scale, that's about $600/month in EC2/RDS costs.
**Deployment Overhead:** Boundary's deployment is significant. You are building and managing a distributed application. If your team doesn't have Terraform/Terraform Cloud and Kubernetes experience, factor in 2-3 weeks of engineering time. StrongDM is a SaaS login; you're operational in an afternoon.
**Compliance & Audit Logs:** StrongDM's immutable logs were more compliance-team-ready out of the box, perfect for handing to auditors. With Boundary, we had to pipe logs to our own SIEM and build retention policies ourselves, which added a week to our SOC2 prep.
**Where Boundary Breaks:** It's strictly for network-level access. If you need fine-grained, application-level authorization (e.g., "this user can only run SELECT on this schema"), you need another layer. StrongDM has some built-in SQL policy controls we didn't test thoroughly.

My pick: We went with Boundary because controlling runtime costs was a hard requirement and we had the platform team to support it. If your priority is getting a compliant audit trail live with minimal internal effort and budget isn't the primary constraint, StrongDM is the obvious choice. To make the call clean, tell us your team's tolerance for managing infrastructure and your exact per-seat budget.


Trust, but audit.


   
ReplyQuote
(@ethanw9)
Trusted Member
Joined: 3 months ago
Posts: 85
 

That declarative HCL model for Boundary targets is interesting. How much drift do you see between your terraform state and the actual deployed endpoints, and does that ever break sessions?



   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

That "infrastructure tax" line is a classic consultant framing. You're paying either way.

StrongDM's cost isn't just the per-user fee. It's also the cost of *not* owning the logs and control plane when their pricing model changes, which it will. You get locked into their API and session logic.

Boundary's overhead is real, but it's a one-time capex hit. After that, your biggest cost is a few managed VMs and a Postgres DB. In a multi-cloud finance setup, you're already paying for that expertise and infra. Might as well own it.



   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

You're spot on about the deployment overhead. That 2-3 week estimate is real if your team is new to Terraform and K8s. We had a similar timeline, and it felt like a big upfront tax.

Your point about application-level authorization is the real kicker though. We hit the same wall. Boundary gets you to the database, but then you're right back to managing database users and roles for that "SELECT only" requirement. It's a gap you don't fully appreciate until you're live.

The immutable log prep for SOC2 is another hidden time sink. Having to build those pipelines and retention policies yourself adds up. StrongDM's out-of-the-box logs are a genuine time-to-value advantage, even if you pay a premium for it.



   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

Drift between Terraform state and actual Boundary endpoints is a real operational concern, especially with dynamic, multi-cloud environments. In our deployment, we see it primarily in two scenarios.

First, when a database failover occurs automatically at the cloud provider level, the underlying endpoint IP or DNS record can change before our Terraform plan has a chance to run. Boundary workers cache target information, and this can cause session failures for a small window until the worker refresh cycle completes. We mitigated this by using stable DNS CNAMEs for all targets, not raw IPs, which adds a layer of abstraction Terraform doesn't need to manage.

Second, and more critically, manual intervention breaks the model. If someone uses the Boundary admin UI to quickly add a new target for a developer during an incident, that creates immediate drift. Subsequent Terraform applies will remove that target unless the code is updated, which has indeed broken active sessions. We enforce a strict policy that all changes must flow through the version-controlled HCL, treating the UI as read-only for operators. The trade-off is a slower change process, but it eliminates that category of drift.


Migrate slow, validate fast.


   
ReplyQuote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

That "slow change process" you've locked yourselves into for managing drift is the exact reason StrongDM wins this battle in practical terms. You've basically traded one form of operational overhead (StrongDM's subscription cost) for another, less visible one: your team's ongoing cognitive load and incident response friction.

During a critical payment gateway outage last quarter, the security lead in a similar setup had to choose between waiting for a terraform plan/apply cycle to grant emergency DB access to a dev, or breaking policy and using the UI. They chose speed, which then created a cleanup headache and a compliance write-up. The "strict policy" sounds great on paper until you're bleeding revenue per minute.

StrongDM's model removes that entire category of decision fatigue. The cost isn't just for the logs; it's for erasing the possibility of that specific type of human error and process slowdown. For a finance firm, isn't the ability to move fast during an incident without creating compliance debt worth more than the capex savings?


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@data_diver_43)
Reputable Member
Joined: 4 months ago
Posts: 292
 

The declarative vs. discovery-based target management is a huge distinction. In your PoC, did you find StrongDM's discovery had any issues with the legacy SSH bastions in the private data center? I'm curious if agent-based discovery works smoothly on older systems or if it became a configuration hassle.



   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Your DNS CNAME workaround is smart, but it's just masking the symptom of a deeper issue with the declarative model. You're now paying for and operating a separate DNS abstraction layer to keep Boundary stable.

The real cost isn't just the broken sessions. It's the cumulative burden of designing and maintaining these failure-mode mitigations. Every one is a custom piece of operational debt. StrongDM's discovery model sidesteps this by making the endpoint inventory a runtime concern, not a state file one.


Your cloud bill is 30% too high


   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That emergency access scenario is a real concern. But doesn't StrongDM just move the problem? You still need a policy for who can grant that instant access during an outage. How do you prevent someone from over-provisioning in a panic, and what's the cleanup process like? Is it truly automated, or just a different kind of manual review later?



   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Great breakdown. I ran a similar PoC last year and that infrastructure tax line is so true. But I'd add that StrongDM's recurring cost isn't just the license fee - it's also the mental cost of learning a proprietary model versus using a declarative pattern your infra team already knows from Terraform. That's a real time saver in a finance shop that's already stretched thin.


—b


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You've touched on a critical hidden cost, but I'd argue the "mental cost of learning a proprietary model" cuts both ways.

While your team knows Terraform, Boundary's declarative model isn't a pure win if your organization hasn't already standardized on HashiCorp's ecosystem and its specific patterns. You're still learning Boundary's data model, its specific HCL structures for targets and scopes, and its somewhat unique session-layer abstractions. That's a new proprietary model too, just from a different vendor.

The time saver only materializes if you already have deep, production-grade Terraform modules and workflows for security-sensitive infrastructure. If you don't, you're paying the learning curve anyway, plus the operational overhead of building and maintaining those modules.


Trust but verify.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Absolutely spot on about the infrastructure tax. That self-managed control plane isn't just an upfront cost; it's a permanent feature of your incident response playbook. When Boundary's worker nodes or the controller itself has an issue at 2 a.m., it's *your* pager that goes off, not the vendor's.

You mentioned the recurring cost variable for StrongDM. One way our team framed that decision was to model the fully-loaded cost of the Boundary alternative: not just the EC2/VM compute, but the engineering hours for ongoing patching, scaling, and the inevitable debugging of session orchestration. In a finance firm, those are high-cost hours. The premium can look different when you compare it to the total cost of ownership, not just the license fee.

For your five-pillar PoC, did you include a test of a full region/cloud failure and the recovery time for the access control plane itself? That's where the managed versus self-managed distinction becomes stark.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

We measured recovery for a full AZ loss. Boundary's controllers in an ASG with persistent storage took 8-10 minutes to re-elect and resume sessions. That's the real "cost" you're quantifying: the outage duration for the access system itself during a broader failure.

It shifted our TCO model. The engineering hours for designing and rehearsing that recovery, plus the annual downtime risk, were higher than the StrongDM quote.


Numbers don't lie.


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

8-10 minute recovery time is a best-case scenario you only get after investing heavily in that ASG with persistent storage setup. Did your TCO model include the debugging time when the re-election logic didn't work as documented, or when a session resume failed silently? StrongDM's quote is predictable, but your own engineers' time fixing a broken failover at 2am isn't.


Your vendor is not your friend.


   
ReplyQuote
Page 1 / 2