Has anyone else run into this? We switched to 1Password Business about six months ago for our ~150 person remote team, and the SSO piece has been a consistent headache. Specifically, the login page itself (the one where you pick your IdP) just... spins and times out. It happens to me, our HR team, and random employees report it weekly.
We’re using Okta as our identity provider. The weird part is it’s inconsistent—maybe 50% of the time it loads instantly, the other half it hangs until you get a timeout error. Refreshing sometimes works, sometimes doesn’t. Our IT team is stumped because our Okta dashboard shows no failed attempts during those timeouts.
A few quick details on our setup:
* We enforce SSO for all web logins.
* The timeout happens before we even reach the Okta login screen.
* We’ve tried it across different browsers and networks (home, office VPN).
It’s becoming a real onboarding pain point—new hires get anxious when they can’t even reach the login step. 😅 I’m hoping someone here has seen this and has a practical fix or even just a direction to point our IT folks in. Is this a known issue with certain configurations, or maybe a DNS/caching quirk?
—Emma
That timeout before you even hit Okta's login is a classic symptom of a DNS or network path issue, not an IdP config problem. Since your IT team sees no failed attempts on Okta's side, the hang is likely in the initial handshake from the user's browser to the SSO service.
We had similar flaky timeouts that turned out to be our DNS load balancer misbehaving. It was routing requests to a geographically distant gateway half the time. A quick way to rule this out is to have someone experiencing the timeout run a traceroute to your SSO domain and compare it to a successful attempt.
Also, check if you have any aggressive client-side security software or browser extensions that might be intercepting the initial redirect unpredictably.
Good call on the traceroute idea. I've seen similar flakiness that ended up being a CDN node with high latency, but only during certain hours.
Question though - wouldn't the DNS load balancer issue affect all traffic to that domain? The OP says it's only happening about half the time on the SSO page specifically. Could it be something messing with the initial SAML auth request, but only when certain conditions are met?
"Spinning before you even hit Okta" sounds familiar, unfortunately. I've had the same headache with a past setup.
Your IT team is checking the Okta dashboard, but are they looking at 1Password's side? The handshake can fail in their cloud gateway. Inconsistent failures often point to an overloaded or poorly configured SAML service provider endpoint. 1Password's SSO service isn't exactly bulletproof - we saw similar timeouts that magically stopped after they migrated our account to a different infrastructure cluster.
Have your team open a support ticket and ask for a deep dive on their SP-initiated SSO logs. It's probably their problem, not yours.
been there, migrated that
That's a solid point about DNS affecting all traffic. But it could still be a routing issue specific to how their SSO gateway is set up. Maybe it's hitting a different subdomain or API endpoint than the main app, and that path has the problem.
The "certain conditions" thing makes me think of time based rules in firewalls or security groups. Maybe something's throttling SAML traffic during peak hours?
Yeah, this tracks with what I've seen. The "magically stopped after they migrated our account" part is super telling. We had spotty SSO timeouts with another service that cleared up instantly when support moved us to a different regional endpoint. It was literally a flip of a switch on their side.
I'd still pair that support ticket with checking your own network egress points though. If your traffic is leaving from multiple offices or cloud VPCs, maybe only one path is hitting the problematic 1Password cluster. That could explain the 50% hit rate.
K8s enthusiast
Emma, you're right to zero in on the "before we even reach the Okta login screen" detail. That's the critical clue. Since your IT team sees nothing in Okta, the failure is happening entirely in the handshake between 1Password's login page and their SSO gateway.
The "different browsers and networks" test you did is good, but it's not a silver bullet. We chased a similar ghost where the timeout was tied to specific ISP peering arrangements with one of our cloud provider's regions. The fix wasn't on our end at all.
You need to push your team to stop looking at Okta and start looking at 1Password's SP logs. Open a ticket with them and demand the raw request logs for both successful and timed-out attempts. I've seen this exact 50% pattern twice before, and both times it was the service provider's load balancer or gateway silently dropping packets for a subset of their IP ranges.
While you wait on that ticket, have your team map exactly which 1Password endpoint IPs the login page is hitting on a success versus a failure. The inconsistency screams of a routing problem on their infrastructure, not yours.
Migrate once, test twice.
Everyone's jumping to blame the network or DNS. They're missing the obvious. You just said it's a "real onboarding pain point." That means it's new hires hitting this right away, which tells you it's not your complex internal setup causing it.
It's likely 1Password's SSO gateway can't handle concurrent auth requests during high traffic periods, like when a batch of new accounts gets activated. You'll see it "randomly" because your team's logins aren't coordinated. Open a ticket and ask them point blank about concurrency limits on their SP endpoints. I bet they have them and they're throttling you.
Just saying.
I like where you're going with the concurrency angle, it's a sneaky one that gets overlooked. It reminds me of a similar issue we had with another vendor's OAuth flow where the SP's token endpoint had a hidden request-per-second cap.
But I'm not fully convinced it's *only* that here. If it were pure concurrency throttling, you'd expect the failures to cluster around predictable times, like 9 AM login rushes. OP's "inconsistent, 50% of the time" pattern feels more like a routing or session affinity problem within 1Password's gateway clusters.
Maybe it's both? The concurrency limit hits, causing a queue, and then their load balancer bounces you to a healthy node half the time and a stuck one the other half.
Either way, you're right that the ticket needs to ask specifically about SP logs and limits. Just asking for "help" will get a generic response.
null
That's a good point about the concurrency not matching the pattern, but I think you're both overcomplicating it. The "50% pattern" is almost always a load balancer or health check failing over to a bad node.
> routing or session affinity problem within 1Password's gateway clusters
Exactly this. I'd bet money their internal health checks are too coarse, marking a degraded node as "healthy" half the time. When the LB sends you there, you timeout. It's the classic symptom of a platform team trying to avoid alert noise by making their checks too lenient.
The real question is why anyone accepts paying for enterprise SSO that has this kind of flaky infrastructure. It's the cloud tax - you pay a premium for a black box you can't debug.
-- cost first
> It's becoming a real onboarding pain point
That's the signal. New hires hitting this on day one rules out your internal config. It's a 1Password platform problem.
The 50% pattern is a classic symptom of a bad load balancer health check. One of their SP gateway nodes is degraded, and their LB is sending half the requests there. When you land on the good node, it's instant. On the bad one, you timeout.
Open a ticket and demand they check the health checks on their SP gateway clusters. Ask for the percentage of requests routed to each node over the last week. They'll see one node with a 100% failure rate on the SAML handshake.
Ship it, but test it first
Good call on the multi-office angle. We had that with our VPN setup. Traffic from the west coast office failed 90% of the time, east coast was fine. It turned out their SP endpoint in one region was just busted.
Makes support tickets tricky because they see a "working" global average. You have to tell them *which* of your IPs is failing.
You're right about that different subdomain or API path angle. I've seen similar weirdness where the SSO flow uses a dedicated `sso-api.service.com` endpoint that has its own DNS zone and routing rules separate from the main app.
The time based rules idea could explain the randomness, but I'm leaning away from it. Firewall rules usually fail 100% of the time when they hit, not 50/50. That half-and-half pattern screams a load balancer issue to me, like one unhealthy node is still passing health checks intermittently.
Have you tried checking if the page times out when hitting the main 1Password web UI versus a direct link to the SSO login? Sometimes they're served from totally different clusters.
Dashboards or it didn't happen.
Agreed on the load balancer guess, it's often the culprit with that flip-a-coin pattern. The comment about health checks being too coarse to avoid alert noise is painfully accurate - I've seen teams do that, and it just pushes the pain onto the customer.
But your last point about the "cloud tax" gets to the real frustration. Even when you're paying for an enterprise SLA, you hit these opaque infrastructure walls. The support ticket process becomes a game of "prove it's your side," when sometimes it's just a bad node in their fleet that their own monitoring won't flag.
Review first, buy later.
Your IT team needs to stop looking at Okta. The failure happens before the request gets there. Open a ticket with 1Password and ask for the service provider logs for both successful and timed-out SAML requests. Compare them. You'll see the request dying at their gateway.
It's likely a bad node in their cluster, but you need their logs to prove it.
Beep boop. Show me the data.