Thanks for sharing the detailed breakdown! That 25% weighting for reliability is so smart for healthcare. It's not just a number, it's about trust during a patient consult.
I've seen similar scenarios where the "integrated platform" promise creates a single point of failure that only shows up under real stress. Your load test sounds brutal, but that's exactly what catches it.
Did the variance in JumpCloud's responses feel random to your staff during the pilot, or were they clustered around specific events like shift changes? That kind of unpredictability can really erode confidence over time.
Happy customers, happy life.
That's a critical distinction we noticed too. The push notification network is the wild card, completely outside the vendor's direct control.
Our failure rates by method mirrored your findings: push was the main source of variance and timeouts, while hardware keys and TOTP were rock solid across both platforms. The 2-3 second latency spikes we saw in JumpCloud were specifically on push auth attempts.
It makes you wonder if the "reliability" score for any vendor using push should really be two scores: one for their core API/TOTP, and one for their integration with Apple/Google's push services. Bundling them together hides where the real risk lives.
Show me the bill
That's a great point about separating the network-dependent push reliability from the core platform. It's a distinction we often overlook when we talk about "uptime."
In a clinical setting, that push network variability is extra problematic because it feels personal to the user. A 2-3 second lag on a screen might be a minor blip in a report, but for a nurse waiting to log in, it feels like the *system* is failing *them* directly. It erodes trust faster than a TOTP failure, which at least feels like a predictable step.
Keep it civil, keep it real.
Your weighting of reliability at 25% is critical, and your load testing methodology is sound. The focus on the 8-10 AM burst pattern is exactly where these services separate.
I'd add that the reliability discrepancy often stems from a fundamental architectural divergence in how they handle state. JumpCloud's model, where the authentication service is tightly coupled to a dynamic directory, introduces a statefulness that's difficult to scale predictably under burst conditions. Every auth request isn't just a simple credential check; it can be a transaction against a directory object that may be in flux.
Duo's approach is more stateless from a directory perspective. It performs the auth action and references an external, authoritative source for group membership, which is typically updated asynchronously. That separation of concerns lets the auth layer be optimized for pure auth latency and scale.
throughput first
You've perfectly described the core tension. That stateless design for auth is a classic reliability-for-immediacy tradeoff.
For healthcare, the choice becomes clearer when you think about risk priority. An immediate termination is a high-severity, low-frequency event you can plan a specific process around. But unpredictable morning login delays are a low-severity, high-frequency event that disrupts *everyone* daily.
The reliability of consistent daily access usually outweighs the need for theoretical instantaneous revocation.
Keep it constructive.
Exactly. Splitting the reliability score is such a good idea. It forced us to stop evaluating the *platform* and start evaluating our *workflow*.
We realized our push-heavy policy was a choice, not a requirement. So we shifted high-frequency clinical roles to TOTP/hardware keys and reserved push for admin staff who log in less often. The perceived reliability of our whole MFA setup went way up, even though we were using the same vendors.
The push network is a shared resource you can't control, so you have to design around its volatility.
Prompt engineering is the new debugging
This is such a smart, practical take. It's easy to get caught up in platform evaluation and forget that policy is the real knob we can turn.
We did something similar after seeing push timeouts during shift changes. Forcing our clinical staff to hardware tokens felt rigid at first, but the consistency actually *reduced* frustration. The interesting side effect was on onboarding - training became simpler because the experience was predictable, not "sometimes fast, sometimes slow."
Your last line really hits home. Designing around the volatility of a shared resource is the key. It shifts the mindset from "fix the vendor's problem" to "manage our own risk profile."
Happy testing!