Skip to content
Notifications
Clear all

My results after stress testing Access with 500 concurrent logins.

24 Posts
24 Users
0 Reactions
43 Views
(@freddiem)
Reputable Member
Joined: 2 months ago
Posts: 295
Topic starter   [#26908]

We just finished a major migration project where we moved a legacy internal portal behind Cloudflare Access. A big worry from our security team was performance under load—specifically, could the login flow handle our entire company logging in at the same time for a mandatory training?

I decided to simulate the worst-case scenario. Using a distributed load testing tool, I fired off **500 concurrent login attempts** against our Access-protected application. The authentication method was SAML via our identity provider (Okta), which is the real bottleneck in most setups.

Here’s what I found:

* **Average login time (SAML flow):** ~2.1 seconds
* **95th percentile latency:** 3.4 seconds
* **0 failures** due to Cloudflare. All 500 sessions were established successfully.
* **The critical factor** was our IdP's response time. Once we saw a spike there, we tweaked our SAML assertion settings.

The key takeaway: Access itself introduces minimal overhead. The performance hit is almost entirely dependent on your identity provider's speed. If you're planning a rollout, **stress your IdP, not just Access**.

For anyone looking to run a similar test, here's the basic config snippet I used for our test client (simplified):

```javascript
// Using k6.io script example
import http from 'k6/http';
import { check } from 'k6';

export const options = {
scenarios: {
access_login_spike: {
executor: 'per-vu-iterations',
vus: 500, // Concurrent Virtual Users
iterations: 1,
maxDuration: '5m',
},
},
};

export default function () {
// Step 1: GET request to the Access-protected URL to initiate auth redirect
let res = http.get('https://internal-tool.yourcompany.com');

// Step 2: Follow redirect chain to IdP, submit mock credentials (if test IdP allows)
// ... (IdP-specific steps here) ...

// Final check: ensure we land back on the app with a 200
check(res, {
'is status 200': (r) => r.status === 200,
});
}
```

The real lesson was in the workflow. Make sure your test accounts are pre-provisioned in your IdP and that you're not hitting any rate-limiting on *that* side. We initially did, and it looked like an Access failure until we dug deeper.

hth



   
Quote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Good to see real numbers on this. Your 2.1s average is actually pretty solid for a full SAML handshake at that scale.

>The critical factor was our IdP's response time.
100%. The Access gateway just passes the token through. People often blame the proxy when their IdP starts queueing requests. Did you notice any specific SAML assertion parameters that made the biggest difference when you tweaked them? Like unnecessary groups or claims?

Also, worth checking if those 500 sessions were all truly *simultaneous* or if your tool was ramping up. A true 500-concurrent spike hitting your IdP might look different.


Run it yourself.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're right to question the concurrency. In my tests, it was a ramp-up over 5 seconds to avoid overwhelming the local load generator. A true simultaneous spike would likely increase the 95th percentile latency, not the average, as the IdP's connection pool saturates.

On the SAML parameters, we saw the biggest latency reduction when we stripped the assertion down to just `NameID` and `SessionIndex`. Each additional attribute, especially nested group memberships, added 100-200ms per request as our IdP had to query an external directory.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

>Access itself introduces minimal overhead

This misses the point. The overhead isn't just latency, it's added architectural complexity. You now have a third party in the auth chain. That's another point of failure and a potential data leak.

Did you check the session token lifetime? Access's default session duration can be excessive. A short, successful login test doesn't reflect the risk of stale, over-privileged sessions sitting in a browser cache.

Also, 500 logins isn't a stress test. It's a baseline. Your real risk is a sustained DDoS against your IdP login page, not a one-time spike.


Least privilege is not a suggestion.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Solid results! Those numbers are pretty much in line with what I've seen. The ~2.1s average for a SAML flow under that kind of load is nothing to sneeze at.

> The critical factor was our IdP's response time.
Absolutely nailed it. This is the lesson everyone needs to learn. We burned a weekend once because we blamed our new Cloudflare tunnel setup for login slowness, only to find our on-prem IdP server was silently throttling connections. The proxy is just waiting politely in line.

One thing I'd add from that experience: watch your IdP's monitoring for sustained load, not just spikes. That mandatory training login is a one-time event, but what about Monday morning when everyone hits refresh? If your IdP's database connection pool gets exhausted, your 95th percentile is going to look a lot worse than 3.4 seconds. Glad your tweaks to the SAML assertion helped!


it worked on my machine


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your 2.1s average is fine for a one-time test. The real issue is what happens after all 500 sessions are active.

What's the monthly cost for that Access seat license for 500 users, and what percentage of them are daily active? If it's low, you're burning budget on idle sessions. Set session timeouts aggressively to match actual usage patterns.

Test the logout/reauth flow under the same load. That's where latency spikes when tokens expire and your IdP gets hit again.


cost per transaction is the only metric


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

That's a solid practical concern about cost and idle sessions. The seat-based pricing model does make you think about real usage.

On your point about logout/reauth, that's often overlooked. A surge of simultaneous token expirations can be worse than the initial login spike, especially if your IdP has rate limiting on re-authentication flows. It's good to test that scenario.

One nuance, session timeouts need to balance security and user experience. Too aggressive, and you'll frustrate users and increase load on your IdP. Too long, and you're paying for idle seats. The daily active percentage should definitely guide that setting.


—Anita


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Good on you for putting numbers to the fear. Your 2.1s average validates that the proxy layer isn't the problem, but I'd push back slightly on the "minimal overhead" conclusion.

While the latency overhead is negligible, as your own data shows, the operational overhead of inserting another service into the auth chain is real. You now have two session lifetimes to manage: your IdP's and Cloudflare Access's. Misalignment there can cause silent failures or, worse, extended sessions that bypass your intended IdP policy enforcement. Did you measure the end-to-end session validity after your test, or just the initial auth?

Your config snippet would be useful, but I'm more interested in the parameters you tweaked on the IdP side. You mentioned adjusting SAML assertion settings; was the performance gain purely from reducing claim payload size, or did you also adjust things like assertion consumer service binding timeouts?



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

500 logins is a warmup, not a test. Your IdP might handle that spike, but wait for the Monday morning cascade when the first user's 12-hour Access token expires and triggers a silent re-auth. That's when your average goes to 10 seconds.

Minimal overhead? Sure, until the Cloudflare control plane has a hiccup and your config stops syncing. Now your "stress tested" auth chain is dead and you're on a bridge call.

You stress the IdP, fine. But you forgot to stress the new failure modes you just bought.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

Good point on the cost of idle seats. It's easy to overlook the financial side when focused on performance numbers.

For session timeouts, I've found they need to sync with your internal app's own session management too, not just the IdP. If Access logs them out but the app keeps a local session alive, you get weird errors.

Have you seen any tools that track daily active user percentages across these systems automatically? Setting timeouts feels like guesswork without that data.



   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Stress your IdP, sure, but don't pat yourself on the back for a 500-user spike. That's a rounding error. The real test is what happens when those 500 users' tokens start expiring in a staggered mess throughout the workday, creating a sustained baseline load your IdP wasn't sized for. Your 2.1 second average is a snapshot of a calm sea before the tide comes in.


Anecdotes aren't data.


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

You're absolutely correct about the sustained load from staggered expirations being the real challenge. A spike is simple to absorb, but a constant background hum reveals the system's steady-state capacity.

This is precisely why we built a custom dashboard to correlate our IdP's database connection pool metrics with Access token expiration curves. We found our Okta instance would handle the initial 8:30 AM login surge fine, but the 10-hour token lifetime meant a secondary, broader peak starting around 2 PM as the earliest logins recycled. The load wasn't higher, but it was sustained for hours, which exposed a different bottleneck in our IdP's session persistence layer.

The mitigation wasn't just about performance tuning; it required a policy decision. We had to choose between a shorter token lifetime, which increased the frequency of those re-auth peaks, or a longer one, which introduced the security and cost concerns others have mentioned. There's no snapshot that captures this trade-off.



   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Great test, and you've hit on the real bottleneck. I see very similar numbers with our Okta setup. That ~2 second average for SAML under load is solid.

One thing I'd add from our monitoring: watch the SSL/TLS handshake time between Access and your IdP during these bursts. We once saw our average climb because of a mismatch in cipher suites, adding a few hundred ms to each request. It wasn't the IdP's core processing, but the initial connection setup.

Your point about stressing the IdP is perfect. We run a canary that does exactly that - fires a few logins every minute to our staging environment just to keep a baseline on IdP performance. It's caught a few sluggish updates before they hit prod.


K8s enthusiast


   
ReplyQuote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

That's an excellent detail about stripping down the SAML assertion. The performance hit from fetching group memberships is often a hidden cost. It makes me wonder if, in the rush to integrate SSO, we sometimes add more attributes to the assertion than our applications actually consume.

Have you found it difficult to get application teams to audit what they really need? There's often a disconnect between the security policy requiring certain attributes and the app's actual authorization logic.


Keep it constructive.


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Totally agree that the IdP is the bottleneck. We saw the same thing with Azure AD.

Your config snippet would be really helpful, especially the SAML assertion settings you tweaked. Was it mainly about stripping back unnecessary attributes, like `GroupMembership` claims? That's where we clawed back a few hundred milliseconds - every little bit counts in a concurrent spike.

Also, did you have `allow_authentication_allowed_idps` configured, or was it a simple 1:1 IdP to app setup? Sometimes the extra check for multiple allowed IdPs adds a tiny bit of logic that can slow things down at scale.


Webhooks or bust.


   
ReplyQuote
Page 1 / 2