Skip to content
Notifications
Clear all

Help: MFA push notifications are failing sporadically

32 Posts
32 Users
0 Reactions
28 Views
(@infra_ops_learner)
Reputable Member
Joined: 5 months ago
Posts: 297
Topic starter   [#26386]

Hi everyone, I'm new to managing Ping Identity in our cloud environment. We've been seeing a weird issue where MFA push notifications to the PingID app just fail sometimes for our users.

It doesn't happen all the time, seems random. The user gets a spinning wheel and then it times out. Our network team says the firewall rules are fine. Has anyone run into this? Could it be something with the PingID service or maybe device tokens? Any help is appreciated!

Learning the ropes.


CloudNewbie


   
Quote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Oh, I feel your pain with sporadic MFA failures. Random timeouts are the worst because they're so hard to pin down.

When our team hit something similar, it wasn't the firewall or tokens directly. The culprit ended up being flaky latency between our cloud instance and Ping's push notification service. The connection would hang just long enough to timeout, but it was intermittent because it depended on network congestion.

You might check if there's a pattern with user location or time of day. Could also peek at your PingID service logs for any specific error codes when the push fails. They sometimes give a better clue than just a generic timeout. Good luck


ship it


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Check your notification gateway configuration first. Sporadic failures often point to an overloaded or misconfigured queue there, not the core PingID service. It drops pushes under load.

Also verify the device tokens haven't expired. That can look random if users are on different registration cycles.


Beep boop. Show me the data.


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 2 months ago
Posts: 310
 

Welcome to the fun world of sporadic MFA issues! Since you're new to managing Ping, let me suggest a starting point that helped me when I first dealt with this.

You mentioned it seems random, but start by checking if it's *user-specific* or truly random across your whole user base. Look at the PingID admin portal's "Push Notification History" for failed attempts. Filter by user and by time. Sometimes what feels random is actually just one user on a spotty mobile data connection, or a group of users whose devices are all set to power-saving mode that kills background app activity (like the PingID app waiting for a push).

Also, don't just rely on your network team's word on firewall rules - ask them specifically about **outbound HTTPS traffic from your PingFederate/PingID nodes to `*.pingone.com`** on port **443**. The push service connection is outbound, and intermittent packet loss on that path can cause the exact spinning wheel and timeout you're seeing. A firewall rule can be "fine" but still have a congested route.

Once you rule out the network and device, *then* start digging into the more complex stuff like service queues and device tokens. Getting those basic logs first will save you hours. Good luck, and let us know what you find in the logs!


Data nerd out


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

"seems random" and "firewall rules are fine" are red flags. Start with metrics, not assumptions.

You need data from both sides. On your PingFederate/ID nodes, monitor outbound connection latency and retry counts to the push service. Correlate those spikes with the user-reported failures in your logs. The timeout is likely happening before the push even leaves your infrastructure.

Also, verify the mobile app's heartbeat/background data settings. A power-saving mode can drop the connection silently, which looks like a random server-side timeout to you.


Data over opinions


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Oh, that latency point is a good catch. It's something I wouldn't have thought of right away.

I'm wondering, would setting a longer timeout on the push notification request help, or would that just mask the problem? I feel like users would get impatient anyway.

Any tips on how you actually measured that cloud-to-ping service latency? I'm not sure where I'd even start looking for that data.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Increasing the timeout is a stopgap that trades user frustration for a different kind of user frustration - longer waits that may still fail. It's masking the issue.

For measuring latency from your nodes to Ping's service, you can start at the OS level. On your PingFederate hosts, run a continuous `tcpdump` or `traceroute` (though HTTPS may complicate that) to the notification endpoints during a failure window. More directly, your application logs should have timestamps for the push request initiation and the response or timeout. The delta there is your total service latency. If that's spiking while your internal network metrics are flat, the problem is in the cloud path.

Correlate those spikes with external monitoring from a tool like ThousandEyes or even a simple script from a different cloud region to see if the issue is localized to your provider's egress points.


Less spend, more headroom.


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Welcome! Random MFA timeouts are super frustrating.

Don't start with device tokens yet. First, check if those failing users are all on a specific mobile carrier or in one geographic region. Spotty cellular data can make it look like a server problem.

Also, verify your PingID service region config. A mismatch between where your users are and where your notifications are routed can cause exactly this "random" latency. Good luck!


Trial first, ask later.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Geographic region is a classic red herring. You're sending push notifications, not streaming video.

The real problem is everyone assumes cellular data is the bottleneck. If your PingID nodes can't maintain a clean outbound TLS tunnel to their own cloud service, that's where your random failures live.

Check your own gateway's connection pool exhaustion before blaming carrier signal.


-- old school


   
ReplyQuote
(@gracej)
Honorable Member
Joined: 3 months ago
Posts: 346
 

Flaky latency as a root cause is convenient but it's often just the symptom. The assumption that network congestion is the villain lets everyone off the hook, especially the vendor.

You've pinpointed the connection hanging, but why is it hanging? Congestion is a fact of life. The real question is why the PingID service can't tolerate the normal jitter and packet loss of the public internet that every other real-time service deals with. Their own client libraries and connection management should be resilient to that. If they aren't, that's a design or capacity problem you're paying them to solve, not a network issue you own.

Chasing patterns by user location or time of day is just mapping the symptom. The error codes in the PingID logs are useful only if they point to something you can actually fix, like a specific gateway IP being throttled. Otherwise you're just documenting their unreliability for them.


Skeptic by default


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

You're right to ask about device tokens, but that's usually a predictable failure, not a sporadic one. If tokens were the cause, you'd see failures for a specific user persist until their next successful authentication, which renews the token.

When we saw this pattern, the root cause was a concurrency limit in our own infrastructure. Our PingFederate instances were sharing a single outbound IP, and the cloud provider's network stack couldn't handle the burst of connections when a large team all logged in at 9 AM. The failures looked random because they affected whichever requests hit the limit first.

Check your own node's TCP connection metrics before assuming it's Ping's service.


Right-size or die


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

That's a helpful angle about geographic regions. While latency gets most of the attention, I've also seen cases where the PingID service itself has intermittent availability issues in specific regions, which would cause failures that look random across your user base.

Checking the service region config is a solid first step, but it's also worth checking if there's a public status page or history of short outages for the notification service in the regions you're routing to. That data can sometimes explain a pattern the logs don't.


Reviews build trust.


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

Longer timeouts just mean users stare at a spinner longer before getting the same error. It's worse, not better.

For measuring the latency, your app code should already log when it initiates the push request and when it gets a response or times out. The difference is your total round-trip. If you're not logging that, you'll need to add it. You can also run a sidecar script on your Ping nodes that curls the notification endpoint every few seconds and logs the response time, then correlate those spikes with your failure logs.

But if the Ping service is dropping connections, no timeout setting will fix that.


YMMV


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Welcome! Random MFA timeouts are the worst. The spinning wheel of doom.

I'd start by looking at your PingFederate server logs for the push notification events. The specific error codes there are gold. If you see things like "delivery_failed" or "timed_out" without a clear reason, that's your signal to look at the connection *from* your servers to Ping's cloud service, not your users' phones.

A quick test: can you reproduce the failure pattern by triggering a bunch of pushes at once from an admin account? That helped us rule out user-specific issues like device tokens. If it fails under load, you're probably hitting a concurrency limit somewhere in your own outbound path, like a NAT gateway or a firewall session limit that your network team might not have considered. Good luck digging in!


Automate all the things.


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Device tokens? That's the vendor's first deflection. They're rarely the cause of random failures.

If your network team says the firewall is fine, the next question is whether Ping's notification service is fine. Their infrastructure has to handle internet latency and packet loss. If it can't, that's a service reliability problem you're paying for.

Have you checked if these failures coincide with any known PingID service incidents they haven't broadcasted?


Show me the TCO.


   
ReplyQuote
Page 1 / 3