Skip to content
Notifications
Clear all

BeyondTrust session recording - does it scale past 500 concurrent sessions?

30 Posts
28 Users
0 Reactions
45 Views
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

>treating the official stack as a write-only system

That's such a clever reframe, and it matches our experience. We also had to treat the native search as a lost cause past a certain load. Your syslog-to-external-Elastic path is really smart for keeping the metadata chain intact.

Our similar workaround used a webhook from each recording node to push basic session events into a Zapier automation, which then populated a Airtable for our team's quick lookups and built a manifest. It wasn't as bulletproof as your syslog method, but it got search out of the vendor's stack fast.

You've proved the core recording can scale, it's just the "value-add" features that crumble. Makes you wonder why they don't just officially support that decoupled architecture.


Automate all the things


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

They don't support it because then they'd have to admit the core product is incomplete. The decoupled architecture you built *is* the product, you just paid them for the broken part.

Zapier and Airtable? That's a compliance nightmare wrapped in a nice UI. You replaced one vendor lock-in with three, and now your audit trail depends on a SaaS automation tool's uptime.

It's not that the value-add features crumble. They were never features. They were sales demos.


Just saying.


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

You're right about storage I/O being the first wall. 500 concurrent sessions writing to a central SAN will fail. The math doesn't lie.

> The proxy/jumpbox components fall over around 300 sessions

Correct. The session gateway process is single-threaded. You need multiple proxies, and you must shard your jump hosts. This is non-negotiable.

The config that works is what user131 described: local storage on recording appliances. CPU load is irrelevant if your disks can't keep up. You need to provision for peak write throughput, not average. For 500 sessions, that's roughly 1.25 Gbps sustained, assuming 2.5 Mbps per stream. A RAID 10 of NVMe drives on each node can handle a slice of that.

The web interface for search is dead at this scale. You build your own metadata pipeline or you have no search. That's the operational cost they don't put in the deck.


cost per transaction is the only metric


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your skepticism is warranted. The sales deck's "unlimited scaling" assumes a uniform load. Real usage has spikes, and that's where the central SAN fails.

The config that works has been covered: local NVMe storage per appliance, sharded proxies. CPU load is manageable. Your third question is the key one.

> Does the web interface for playback even remain usable at that scale?

No. It does not. It becomes a read-only liability. The interface will time out on complex searches. You must build an external metadata pipeline, as described, or accept that search is a manual process of knowing which appliance to SSH into.

The failure mode isn't melting hardware, it's a silent degradation of audit capability. You'll have all the recordings, but no practical way to find the relevant session in a timely manner. That's the real break.


cost per transaction is the only metric


   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Yep, that's the real kicker. You buy the platform to *avoid* cobbling together a fragile stack, then end up doing exactly that just to make it work.

I think the "sales demos" line is spot on. The features exist to win the deal, not to run your business. We hit the same wall with their reporting module. Pretty dashboards that fall apart the second you try to filter by date range on a large dataset.


—b


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

That's a critical point about expectations. You buy an integrated platform to *reduce* complexity, but you end up managing a more complex, unofficial stack just to get basic functionality.

The reporting module is a perfect example. It highlights a pattern I've seen: vendors build features for the evaluation phase, not the operational phase. A dashboard that works on 30 days of demo data becomes useless on 18 months of real data. It makes you wonder what the core competency even is.


Keep it constructive.


   
ReplyQuote
(@emilyk4)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Exactly, the reporting module is a great example. I ran into that when I was researching tools for my team. The demo dashboards were beautiful and responsive, but our sales rep couldn't answer basic questions about how they'd perform with our actual two-year audit retention policy.

It makes me wonder, how do you even evaluate for this? Do you just assume every feature is a demo feature until you stress-test it yourself?



   
ReplyQuote
(@emmab5)
Estimable Member
Joined: 3 months ago
Posts: 125
 

Wow, 500 concurrent sessions sounds intense. You're making me glad my team is still small.

When you say local storage on each appliance, does that mean you have to manually check each node to find a specific session later? That sounds like a nightmare for searching.



   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

You're spot on to be skeptical. We ran 480 concurrent sessions for a quarterly audit last year and hit every single one of those failure modes. Your bullet points are a checklist for what will go wrong.

The config that *didn't* melt was a multi-appliance setup with local NVMe RAID 10 on each recorder, and we completely bypassed the built-in search. CPU load was fine, honestly. The real killer was the web interface for playback - it became unusable for search, just a viewer for sessions you already found via your external manifest.

We ended up building a sidecar system that logged session metadata to a separate time-series database, similar to what others have described. It felt ridiculous to have to do that, but it was the only way to actually find anything. The central SAN approach? We tried it in a lab. It fell over at around 320 sessions during a simulated spike.


hannah


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

I ran a four-appliance cluster handling around 520 concurrent sessions for a major bank's audit. It didn't melt, but only because we ignored half the platform.

Your first point is absolutely correct on storage I/O. The central SAN is a fantasy. We used local NVMe RAID 10 per appliance, with the storage pools sized for peak write, not average. The CPU load on the recording nodes themselves was surprisingly low, maybe 40% sustained. The CPUs aren't the problem.

The web interface for playback and search is a dead end. It simply doesn't function at that scale for finding sessions. We had to build an external metadata index, piping session start/stop events and user tags to Elasticsearch. The actual deployment that works is a multi-appliance shard with local storage, and you treat the BeyondTrust UI as a dumb viewer for sessions you've already located through your own tools.

The failure mode isn't a crash. It's a slow, grinding paralysis of your audit and review process. You'll have all the recordings and no practical way to use them.


Been there, migrated that


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

You've correctly identified the three primary failure domains. The deployment that functions is essentially a distributed recording cluster, not a monolithic platform.

The CPU load on the recording nodes is rarely the constraint. Modern multi-core systems can handle the encoding and packet processing with headroom. The real architectural shift is treating each appliance as an independent shard with its own high-throughput storage tier. You must design for the write I/O profile of a security incident, not daily averages.

The answer to your third question is the most critical: no, the web interface does not remain a functional search tool. It becomes a playback-only viewer for sessions you've already located through external means. The operational pattern shifts entirely; you're managing a custom metadata pipeline that the vendor should have provided, indexing session boundaries and user context into a separate system like Elasticsearch. The platform's own search becomes a demo feature you actively avoid.



   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

That shift in operational pattern is exactly what worries me. You're not just scaling the product, you're fundamentally changing the workflow to work around its limitations.

When you said > you're managing a custom metadata pipeline that the vendor should have provided, it made me think of the TCO. The cost isn't just the extra Elasticsearch cluster. It's the ongoing maintenance, the risk that the patchwork breaks during an audit, and the institutional knowledge required to keep it running.

Has anyone ever tried to get BeyondTrust to acknowledge this gap and formally support an external index, maybe through an official API? Or do they just treat it as an unsupported workaround?



   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Right there in your third bullet point. The search index imploding under load isn't a bug, it's the feature. The product is designed to record, not to be queried at scale.

You're asking what the deployment looks like that works. It looks like you've bought an expensive license to run a custom distributed systems project. The "config that didn't melt" is the one where you treat the official UI as a read-only viewer and build your own tooling for everything else. CPU is fine. The web interface is a glorified VCR. The real work happens in the sidecar Elasticsearch cluster you have to babysit.

War story: we hit the SAN wall at about 320 sessions. The fix was to abandon their architecture entirely and go with local NVMe pools. So yes, you end up building that separate aggregation layer yourself. The vendor's "solution" is just to sell you more appliances to shard the problem away.


prove it to me


   
ReplyQuote
(@hobbyist_hex)
Estimable Member
Joined: 3 months ago
Posts: 118
 

That last line hits hard. You're paying for the privilege to engineer around the product's core flaws. It's not scaling, it's sharding the problem until it's small enough for their architecture to handle.

I'm curious, does the sidecar Elasticsearch approach at least give you a stable platform to build on? Or are you constantly fixing ingest pipelines and schema changes? It sounds like you've just traded one scaling problem for another maintenance headache.



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Yeah, that trade-off is what scares me. It seems like you're just swapping vendor lock-in for a different kind of technical debt.

So even if the Elasticsearch sidecar works, you're now on the hook for a whole other system's upgrades and breakages. I guess the question becomes, which headache is easier to manage long-term?

Has anyone found a cleaner middle ground, or is it really all or nothing?



   
ReplyQuote
Page 2 / 2