Good to see a clear root cause and a path for manual verification. The spam folder check is a necessary first step, but it's often a dead end when the real issue is sender reputation decay you can't see. That test trigger better use the same sending infrastructure, or it's just theater.
Postmortem numbers would be useful. What was the cap, and what's the new one? A 3-attempt cap is a rookie move for anything beyond password resets.
Yeah, the point about needing the test to use the same infrastructure is spot on. If it's just a mock, you won't catch weird stuff like a specific port being blocked by a new firewall rule.
> A 3-attempt cap is a rookie move
I'm just starting to set up our own mail service. What's a more reasonable number for a general notification system? Is 5 tries a better starting point?
Containers are magic, but I want to know how the magic works.
Your manual test bypasses user rate limits. Good. That's a valid diagnostic.
But if you're serious about catching this, publish the retry stats for that test path. Time in queue, number of attempts, final disposition. Otherwise you're just checking if the pipe is open, not if it's efficient.
A 3-attempt cap is absurdly low. You'll lose notifications. The cost of a few extra retries is zero compared to the support load for missed emails.
cost per transaction is the only metric
It was the max number of tries, capped at three. As others have said, that's way too low for general notifications. We've bumped ours to seven with exponential backoff.
Your spam folder check story about the subject line changing is a perfect example of how fragile delivery can be. We once had a vendor's name in our "From" field that got flagged by a new filter overnight. The content was identical, but the domain reputation shifted.