Skip to content
My results after st...
 
Notifications
Clear all

My results after stress-testing 3 ZTNA gateways with k6.

2 Posts
2 Users
0 Reactions
31 Views
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
Topic starter   [#20090]

Hey folks! Been down a rabbit hole the last few weeks evaluating ZTNA gateways for a potential migration from our old VPN. We're a cloud-native shop, so the "never trust, always verify" model fits like a glove.

But specs on paper are one thing. I wanted to see how they *actually* perform under load, especially during auth storms (think: everyone logging in at 9 AM). So I spun up k6 and stress-tested three major vendors' ZTNA gateways (keeping them anonymous as Vendor A, B, and C). Focused on the data plane after the initial authentication.

**The setup:**
- Simulated 500 concurrent "users" ramping up over 2 minutes.
- Each user makes HTTP requests to a simple internal app through the gateway every few seconds.
- Measured HTTP req duration, failure rate, and gateway CPU/mem via their admin APIs.
- All deployed in AWS us-east-1.

**Key finding that surprised me:** The agent-based gateway (Vendor B) crushed it on latency under load, while the agentless ones (A & C) showed higher tail latency. Trade-offs, right?

Here's a snippet of the k6 config I used for Vendor B's test:

```javascript
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
stages: [
{ duration: '2m', target: 500 },
{ duration: '5m', target: 500 },
],
};

export default function () {
const res = http.get('https://gateway-b.company.com/internal-app/api/health', {
headers: { 'Authorization: Bearer ${__ENV.APP_TOKEN}' },
});
check(res, { 'status was 200': (r) => r.status === 200 });
sleep(Math.random() * 3);
}
```

**Quick results summary:**

* **Vendor A (Agentless):** Highest throughput, but 95th percentile latency jumped to ~850ms under full load. A few auth timeouts.
* **Vendor B (Agent-based):** Most consistent. 95th percentile stayed under 200ms. Agent overhead was negligible in this test.
* **Vendor C (Agentless):** Lowest resource use on the gateway node, but latency spiked more than expected – some TCP connection drops under sustained load.

The big takeaway for our team? If you need rock-solid predictable performance for a known set of devices, the agent path is solid. But if you need broad, unmanaged device access and can tolerate a bit more latency, agentless has its place.

Has anyone else run similar load tests? I'm especially curious about how ZTNA gateways handle WebSocket/Secure Shell traffic under strain. My next weekend project!

Keep deploying!


Keep deploying!


   
Quote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That's a fascinating finding about agent-based vs agentless latency under load. I've seen similar trade-offs in marketing automation platforms where a lightweight connector outperforms a full API-based integration during peak send times, but it introduces more deployment complexity.

Did you happen to measure how each gateway handled connection persistence? Sometimes that tail latency difference comes from how frequently the agentless gateways re-establish or validate sessions, not just raw throughput.

Would love to hear how the CPU/memory consumption compared between Vendor B and the others. An agent chewing resources on endpoints might win the latency battle but lose the war on operational overhead.


Cheers, Henry


   
ReplyQuote