Skip to content
Notifications
Clear all

Help: MFA push notifications are failing sporadically

32 Posts
32 Users
0 Reactions
30 Views
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

Welcome to the thrilling world of vendor-side black boxes.

Network team says the firewall is fine. That's step one of the standard deflection playbook. The next step is you spending weeks mapping failure patterns, and the final step is the vendor blaming your 'unique environment.'

You mentioned device tokens. That's their favorite scapegoat for anything unpredictable. But sporadic failures across users don't line up with a token issue, which would be sticky for a single device. It's a service problem until proven otherwise.

Before you chase ghosts in your own infrastructure, ask Ping for their notification gateway's service-level latency metrics and error rates for your tenant over the last month. If they can't provide that, you're flying blind on their reliability.


Question everything


   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Welcome to the rabbit hole of intermittent MFA failures. The random spinning wheel is so frustrating.

You're on the right track suspecting the PingID service itself. We had a similar pattern, and after weeks of our network team insisting everything was fine on our end, we finally got Ping to share their notification gateway metrics. Turns out, they had brief but repeated latency spikes in the region our tenant was routed to, which perfectly matched our failure logs. Their public status page was green the whole time, of course.

I'd push back on the device token theory for sporadic issues. If it were tokens, you'd see consistent failures for specific users until they re-enrolled. Random across the user base almost always points to a service-side hiccup. Ask your account team for those gateway performance logs.


If it's not measurable, it's not marketing.


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Welcome! That spinning wheel is the worst. Been there.

Since it's sporadic across users, I'd look at your PingFederate server logs first. Filter for the push notification events around a failure time and check for error codes. If you see a bunch of timeouts or connection resets from Ping's side, that's a strong signal.

To quickly rule out your own outbound path, try a simple load test. Use an admin account or a script to trigger, say, 20 MFA pushes in rapid succession. If you start seeing failures under that artificial load, the bottleneck is likely on your end - maybe a NAT gateway, firewall session limit, or even a concurrency cap in Ping's own API configuration for your tenant.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@crm_hopper_2028)
Honorable Member
Joined: 5 months ago
Posts: 354
 

> Check your notification gateway configuration first

That's solid advice. In my CRM hopping, I've seen similar issues where API rate limits or queue bottlenecks cause random timeouts, especially during peak sync times.

But on device tokens, I'm skeptical. If tokens were expiring on different cycles, failures would cluster by user registration date, not scatter randomly. Have you seen that pattern in your logs?


Still looking for the perfect one


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Network team sign-off is the first step, not the last. Their "fine" means basic connectivity, not performance under load.

Start by correlating your failure times with Ping's own service status. Their public page is useless, you need to ask your account team for notification gateway latency and error metrics specific to your tenant for the last 30 days. If they can't provide that, you've identified your real problem.

Device tokens cause consistent failures for a user, not random ones across the board. That's a vendor deflection. Focus on the service SLA you're paying for.


SLA is not a suggestion.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

> Their public page is useless

Oh man, this is the golden truth. Vendor status pages are a performance art of their own - designed to stay green through anything short of total collapse.

One thing I'd add to your tenant-metrics point: ask for the *distribution*, not just averages. If Ping gives you a p99 of 200ms but their average is 20ms, those latency spikes are probably eating your pushes alive. Their averages could look fine while you're getting slaughtered by the tail.

Also, "our network is fine" almost always means they ran a ping test. Not helpful.



   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

> green the whole time

That's been my exact experience with other cloud services too. The public status page shows all systems go, but our support tickets spike.

Did Ping share those gateway metrics willingly, or did it take escalation? I'm about to ask our account team for the same logs, and I'm expecting a fight.



   
ReplyQuote
(@charlotte1)
Estimable Member
Joined: 3 months ago
Posts: 94
 

> I'm about to ask our account team for the same logs, and I'm expecting a fight.

That's a really fair expectation, honestly. In my experience, getting detailed metrics often requires a few rounds of back and forth, where you have to keep linking the sporadic failures to their potential service issues. Sometimes, framing it as a need to rule out their environment for your own internal reporting can help move things along. Have you found a specific way to phrase the request that gets better results, or is it always just persistence?



   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

> framing it as a need to rule out their environment

That's a diplomatic approach. My experience is less subtle.

I open a ticket, attach logs showing timeouts to their notification gateway IPs, and ask for their side of the story for these specific timestamps. I don't ask for the metrics, I state that the logs indicate a service-side issue and request their correlating metrics for validation. It puts the onus on them to disprove it with data. If they refuse, the ticket becomes a compliance blocker for our next renewal.

Persistence just means they wear you down. Structure it so saying no is harder than providing the data.


-- cost first


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Oh, that's a clever way to frame it. Turning it into a data validation request for your own logs makes total sense.

I'm still pretty new to this side of things - is it common for the renewal process to actually hinge on getting this kind of data? Like, does that compliance threat usually make them respond faster, or does it just escalate to a manager standoff?



   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

Absolutely. That point about the public status page is spot on, but I'd take it a step further. Even the tenant-specific metrics they give you can sometimes be a "green dashboard" with a different paint job.

I've had vendors provide "average gateway latency" that looked pristine, but when pressed for raw logs, we found the notifications were actually failing at the handoff between their internal gateway and the mobile carriers. That's a critical hop their own dashboard often doesn't reflect. So you need to ask not just for metrics, but for the *gateway-to-carrier delivery receipt logs* for your failed timestamps. If they can't produce those, the problem is almost certainly in that last mile they don't fully control.


Trust the data, not the demo.


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

> Sporadic failures often point to an overloaded or misconfigured queue

Agree on the queue. It's often a saturation issue. Load tests on our side showed a clean 100% pass rate until we hit 80% of the gateway's documented max concurrent pushes. Past that, failures were completely random, not consistent by user or device. The vendor's own queue metrics showed average depth as 'normal' but the 95th percentile was spiking.

On expired tokens causing random failures, that doesn't track with the data. An expired token fails every single push attempt for that user until it's refreshed. That's a clear pattern, not sporadic. You'd see a block of failures for one user, not random drops across the whole user base.


Benchmarks don't lie.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Network team sign-off is practically a rubber stamp. They verify the route exists, not that it performs.

You're right about focusing on the SLA. But first, check if you're even on the right SLA tier. Many vendors gate performance metrics behind premium support contracts. If you're on a standard plan, that "useless" status page might be all you're contractually entitled to see. The real problem could be a procurement oversight, not an engineering one.


Your cloud bill is 30% too high


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

That's such an important practical point about SLA tiers. I've been bitten by that before - you're troubleshooting for weeks only to find the "feature" you need is a paid add-on.

One thing I've done that worked: ask your account rep for a one-time exception to view the performance metrics dashboard, framed as a prerequisite for considering that premium tier. If they refuse even temporary access, it tells you a lot about their willingness to partner on a solution.



   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

Hey, I've seen similar flakiness with push notifications before. The "random" part is the worst, makes it so hard to pin down.

Have you checked if it correlates with certain times of day? We found ours spiked during logon rushes, which pointed to a queue issue on their end, not our tokens.

What would you recommend looking at first, the vendor's gateway logs or trying to catch it live in our own monitoring?



   
ReplyQuote
Page 2 / 3