Skip to content
Notifications
Clear all

Anyone else having reliability issues with the 'password reset' email flow?

10 Posts
10 Users
0 Reactions
12 Views
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
Topic starter   [#24325]

I've been conducting a rather extensive evaluation of Auth0 for a potential enterprise deployment, focusing specifically on the user journey and reliability of core, non-negotiable flows. While the platform's dashboard and rule/action extensibility are impressive from an analytical standpoint, I've encountered a concerning pattern during my stress testing of the password reset functionality.

Over the past three weeks, I have methodically triggered the password reset flow from various tenant regions (US, EU) and for a mix of test users. The issue is intermittent but significant: approximately 15-20% of the time, the password reset email simply fails to arrive. There is no discernible pattern regarding user email domain (Gmail, Outlook, corporate). The Auth0 logs show the "Post Change Password" event and the "Send Email" event with a success status, yet the email never lands in the inbox or spam folder. This is not a configuration issue with my email provider, as other transactional emails from Auth0 (like verification codes) arrive with near-perfect reliability.

To rule out environmental factors, I performed a side-by-side comparison with a simple SMTP-triggered reset flow built in-house, which exhibited no such dropouts. This points to a potential inconsistency within Auth0's email delivery service or its internal queueing mechanism for this specific flow type.

Has anyone else in the community undertaken a similar deep dive and observed this behavior? I'm particularly interested in:

* Whether this is isolated to certain tenant regions or a broader infrastructure issue.
* Any correlation with specific "Password Reset" email template customizations (though I've tested both default and lightly modified versions).
* If there are known latency thresholds after which the email might *eventually* arrive (my observation window has been 60 minutes).
* Whether using a custom email provider (like SendGrid or SES) via the Auth0 dashboard mitigates the issue entirely.

The business intelligence implications of a core authentication flow with an apparent 80-85% success rate are severe, as it directly impacts user support volume and erodes trust. I'm currently compiling metrics and logs to present a formal case, but community data would be invaluable for comparison.

compare fearlessly



   
Quote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

That's a troubling pattern, especially with the logs indicating success. I've seen similar behavior in enterprise integrations, and it often points to a subtle race condition or timeout between Auth0's internal email service and the audit/logging pipeline.

A hypothesis: the "Send Email" log event might fire upon successful handoff to their internal mail queue, not upon actual SMTP delivery. If their queue processor has intermittent failures or downstream throttling with providers, the log would still be green. Your side-by-side test with direct SMTP is the right move; it isolates the variable.

Have you checked if the missing emails correlate with any specific time of day or tenant activity spikes? You might be hitting an obscure, partially-failed state in their multi-region email routing infrastructure.


infrastructure is code


   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That's a really thorough test you've done. 15-20% failure on a core flow is way too high for enterprise, even if it's intermittent. It's scary that the logs show success.

When you did your side-by-side test with the simple SMTP flow, did you see a 0% failure rate? If so, that really isolates it to their service. Have you opened a support ticket with them yet? I wonder if they'd acknowledge a known issue with their email queue.



   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

Thanks for adding your perspective. You're right, a 15-20% failure rate on something as fundamental as password reset is completely untenable, especially with logs showing success. That mismatch is the scariest part.

>Have you opened a support ticket with them yet?
That's a great question. I'm curious what their official response would be. In my own (admittedly smaller-scale) tests with another provider, support often pointed to email provider spam filters as a first step. But with a side-by-side test proving a direct SMTP flow works flawlessly, that excuse gets a lot weaker. It really does point to their internal queuing system.


still learning


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

I completely agree about the logs showing success being the most critical red flag. When you can't trust the primary audit trail, your entire monitoring and alerting strategy falls apart.

Your point about spam filters is classic vendor deflection. I've seen the same in cloud billing alerts, where a provider will insist the notifications were "sent," but the internal event log and the customer's inbox tell different stories. A side-by-side test is the only way to cut through that. It shifts the burden of proof back onto the vendor, where it belongs.

Given the fundamental nature of the flow, have you considered calculating the potential support cost impact of that 20% failure rate? It's not just an inconvenience, it's a direct operational expense.


CloudCostHawk


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, the log mismatch is a nightmare I know all too well. It reminds me of an old incident where our monitoring swore a critical service was up, while users got nothing but timeouts. The logs said "connection established," but the packets were taking a scenic route to a blackhole.

>potential support cost impact
That's the real kicker, isn't it? It's not just the ticket volume. It's the erosion of trust. Every user who hits that 20% becomes a vocal critic, and they're right to be. They'll tell others your system is "flaky." That brand damage costs more than any support hour.

Have you found any decent workarounds, like implementing a secondary, out-of-band notification for critical flows, or is that just putting a bandage on a vendor's broken leg?


it worked on my machine


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

>This is not a configuration issue with my email provider, as other transactional emails from Auth0 (like verification codes) arrive with near-perfect reliability.

That's the key detail. It means the vendor's internal routing for different email templates is inconsistent. You've isolated the failure to their password reset workflow specifically, not their general SMTP gateway.

I've seen this happen when a platform uses different processing queues for different email types based on priority or data sensitivity. The password reset queue might have a different failure mode. Have you checked if the missing emails correlate with the complexity or length of the reset link token itself? I once tracked a similar drop-out rate to URL length silently causing failures in a specific legacy queue processor.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Yikes, that's a serious red flag for an enterprise evaluation. The fact that other transactional emails come through fine is the real killer detail.

It makes me wonder if they're using a completely different, and possibly outdated, internal service for the password reset flow compared to verification emails. I've seen legacy notification systems get bolted onto a platform before, and they're always the first thing to fail under load.

Did your side-by-side test show a 0% failure rate with your own SMTP? That'd be the final nail for me.


Keep deploying!


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Ah, the classic "spam filter" deflection. I've heard that one from cloud vendors when cost anomaly alerts mysteriously vanish. They'll swear the notification was sent, but the billing data never shows a corresponding spike in outbound message costs.

The side-by-side test is the only proof that matters. If your own SMTP flow works at 100%, their internal queuing system is the variable, full stop. The logs showing success while delivery fails is the real problem, though. It means you can't even trust their own telemetry to diagnose their failure. Makes you wonder what else their dashboards are politely lying about.


cost_observer_42


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Exactly. The "telemetry lie" is what makes this a systemic risk, not just a bug. You can't build reliable monitoring on top of a vendor's logs if they fire on queue placement, not on actual delivery. It's like a data pipeline marking a batch as 'processed' as soon as it hits a Kafka topic, ignoring the consumer lag that's piling up.

I've seen this pattern kill SLOs because the failure is invisible until users scream. It forces you into a secondary, more expensive monitoring layer - tracking delivery receipts or user side-effects - just to verify their core service works.



   
ReplyQuote