Hey everyone! 👋
I've been diving deep into our CyberArk PSM setup for database access, and while the security benefits are a no-brainer, I keep circling back to a performance question we haven't fully answered. We have a handful of business-critical, high-traffic SQL Servers (think thousands of queries per minute, some quite complex) that our analysts need to access via PSM. We're using the standard RDP proxy connection to a dedicated PSM server, which then connects to the target DBMS via SSMS.
Lately, I've been hearing some grumbling from our power users about "lag" or "feeling slower" when running large reports through the PSM session compared to their old, direct (and now forbidden!) connections. It's not quantifiedβjust a feeling. But in my world, feelings are just unmeasured metrics! 😄
So, I'm on a mission to benchmark this properly. Before I reinvent the wheel, I wanted to ask:
**Has anyone here done a formal performance comparison for SQL workloads over PSM proxy vs. direct?**
I'm especially curious about:
* **Methodology:** What did you measure? Average query execution time reported by SQL Server? End-to-end latency from the user's click to data appearing? Network latency *through* the proxy? All of the above?
* **Findings:** What was the average overhead or "performance hit" percentage you observed? Was it negligible (10%)? Did it vary dramatically based on result set size or query complexity?
* **Configuration Tweaks:** Did you find any PSM or connection broker settings that helped mitigate latency? For instance, anything around session timeouts, RDP optimization settings, or resource allocation on the PSM servers themselves?
* **Scale Context:** How many concurrent PSM sessions were hitting your SQL boxes during your tests? Our peak could be 30-40 concurrent sessions on a single target.
I'm planning to set up a controlled test using a repeatable set of stored procedures (some simple selects, some heavier joins and aggregates) and capture timestamps at various points. But if someone has a script or a report template they'd be willing to share, I'd be eternally grateful! It would save me so much time.
Also, if your conclusion was "the hit is real, and here's how we justified it to the business," I'd love to hear that story too. Balancing security and performance is always the fun part, right?
Thanks in advance for any data, war stories, or config wisdom you can offer!
test everything twice
Excellent question, and you're absolutely right to treat that "feeling" as a signal demanding quantification. I haven't benchmarked PSM specifically for SQL, but the architectural overhead is predictable.
You're adding two network hops and a session proxy's CPU overhead for RDP packet processing. The measurable impact will be on the *perceived* latency for the user's interactive session - mouse/keyboard echo, screen refreshes for large result sets - more than the actual SQL execution time on the server. SSMS running on the PSM box should see near-native query times.
Your benchmark should isolate these layers. Capture SQL execution times from the PSM server itself via Profiler or Extended Events for a controlled query set. Then compare that to the full end-to-end time a user experiences, which includes RDP's compression and transport. The delta is your PSM tax, and it's almost entirely in the presentation layer.
What's the network latency between your users and the PSM farm? That's often the primary multiplier for that "lag" feeling, especially with large graphical result sets.
CostCutter
Exactly, the "presentation layer" is where it bites you. I've clocked near-identical query times from the PSM server itself, but the RDP rendering for big SSMS result sets is a real bottleneck.
The graphic compression isn't optimized for dense data grids. Ever notice how scrolling a large Results pane feels choppy? That's the extra tax. What's your round-trip latency to the PSM? Even 30ms gets painful fast when painting cells.
Automate everything.
We haven't run a full formal benchmark, but we did some quick instrumentation on a similar high-volume MSSQL setup. I can tell you where our biggest numbers came from.
We logged query duration from SQL Server itself (using Extended Events) and compared it to the total time from the user's "Execute" click to the "Query completed" message in their SSMS session via PSM. The database execution time was almost identical. The entire added latency was in that RDP channel, especially for returning large result sets. We're talking an extra 150-300ms on queries returning just 10k rows, just for the grid to populate visually. For analysts running big reports, that adds up fast.
Your approach is spot on. Start by measuring the SQL execution time directly from the PSM server to isolate the database performance. Then measure the full user journey. The gap is your PSM/RDP overhead. What's the network latency between your users and the PSM box? That's your baseline floor for the added lag.
terraform and chill
You're absolutely correct about isolating the layers. I'd add that the benchmarking method itself can skew results if you're not careful.
While a Profiler trace from the PSM server gives you execution time, it misses the SQL Server's own resource contention from handling dozens of simultaneous proxy connections that all appear to come from the same PSM host. You might see similar execution times in a vacuum, but under load, the DBMS's response to a single source IP making hundreds of connections can differ from distributed user IPs. This can subtly increase wait times that then get amplified by the RDP layer.
The network latency multiplier is key, but also consider the PSM server's own resource limits. Its CPU handling the RDP graphics encoding for multiple high-resolution SSMS sessions can become the bottleneck, independent of network quality. Have you monitored the PSM host's performance counters during peak usage?
Migrate slow, validate fast.
That's a really good point about the PSM host's own limits. I hadn't considered that. So even if the network and SQL server are fine, the PSM server's CPU could be maxed out just from encoding all those screens.
For monitoring, are you checking standard Windows Performance Monitor on the PSM server, or something else? I'm curious what specific counters to watch for that graphics encoding load.