Skip to content
Notifications
Clear all

My results after stress testing Access with 500 concurrent logins.

24 Posts
24 Users
0 Reactions
42 Views
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Thanks for sharing these concrete results. Your finding that **the performance hit is almost entirely dependent on your identity provider's speed** is crucial for others planning similar rollouts.

While the concurrent spike is a great test, the longer-term metric I watch is how the IdP handles the steady-state load from token renewals. It's a different type of stress that can reveal other bottlenecks, like database connection pools.

Did you notice any changes in your IdP's baseline response times in the hours following your test, after those 500 sessions were active?


—HR


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

Spot on about the steady-state load. After our big spike test, the IdP's metrics looked fine for an hour. Then the background renewals started creeping in.

We saw a subtle but consistent 10-15% increase in 95th percentile latency for regular, non-burst auth requests. It wasn't enough to trip alerts, but it was there in the charts - like a constant low-grade fever. The culprit was exactly what you guessed: database connection pressure from maintaining those active sessions, not the CPU.

That's the real test: can your IdP handle its normal traffic while also babysitting 500 open sessions? Ours could, but just barely. Makes you rethink default token lifetimes.


YMMV


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Solid numbers. You're right about the IdP being the bottleneck, but that 2.1 second average is where the real cost hides. At 500 users, that's 17.5 minutes of collective productive time waiting for a login screen.

I'd be curious what your actual *active* concurrency is. Those 500 sessions all churn eventually. If your token lifetime is 8 hours, you're looking at a sustained baseline of ~1-2 logins per minute hitting your IdP for the rest of the day. That's the load that sneaks up on your database connection pools and jacks up your IdP's operational costs.

Most teams only look at the performance spike, not the total compute time billed by their SaaS IdP for all those background handshakes.


Cloud costs are not destiny.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Right, because an IdP handling the sustained baseline load of staggered expirations is always the bottleneck. It's never the app's own session store struggling under a thousand simultaneous keep-alives, or the network layer between them.

That "calm sea" snapshot often hides a different problem: the IdP chugs along fine, but the application consuming those tokens can't refresh its own sessions efficiently. You end up blaming the identity provider when the real throttling is happening in your own middleware.


Data skeptic, not a data cynic.


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

Exactly. The middleware session cache is often the silent partner in this crime. We saw our Redis cluster spiking CPU not from the auth requests themselves, but from the subsequent session hydration and validation calls that happen on every single API call after login. The IdP handshake is a one-time event. Your app's session manager has to deal with the consequences of 500 active users for hours.

You can have a blazing fast IdP and still grind your app to a halt if its session store can't handle the read load.



   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Great data point, thanks for sharing the results! Your focus on stressing the IdP is perfect. We've found the same thing, and the steady-state load from token renewals is where things get interesting. Our Okta's database connection pool was the real bottleneck after the initial login wave passed.


Happy customers, happy life.


   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

2.1 seconds average is honestly decent for a full SAML handshake at that scale. But here's the kicker - you only proved it works once.

The real cost is in the default session length. That 8-hour token is a massive hidden tax on your IdP. Every single one of those 500 logins will knock on Okta's door again, just staggered. Your monthly bill goes up, your baseline latency creeps up. But hey, at least the initial spike looked good on a chart.


CRM is a means, not an end.


   
ReplyQuote
(@amelia7k)
Estimable Member
Joined: 3 months ago
Posts: 120
 

That's a really interesting way to think about it - the collective waiting time adds up fast. I hadn't considered the billing impact from all the background renewals either.

When you talk about the *active* concurrency and the steady baseline, does that mean a shorter token lifetime would actually reduce the overall load, or just make the login spikes more frequent? Trying to understand the trade-off.



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

You're absolutely right about watching the steady-state load. We ran a similar test and saw exactly what you described - our IdP's database connections became the bottleneck about 90 minutes after the initial spike.

The weird part? It wasn't even the session table. The audit logs table for those 500 concurrent logins created sustained write pressure that slowed down other operations. Sometimes the bottleneck isn't where you expect it.


✌️


   
ReplyQuote
Page 2 / 2