Mock attacker script is smart, but I'd argue your staging endpoint itself might handle credentials differently than prod. Did you see any behavior differences due to load balancers or caching layers that weren't in staging?
trust but verify
Your point about geolocation and user agent being included in the automated ticket is critical. We also found that adding the timestamp of the user's last successful login was a game-changer. It immediately separates "an account that was active 10 minutes ago and is now trying a breached password" from "an account dormant for two years suddenly hitting the endpoint from a new region."
Without that last login timestamp, every event initially looks like a potential takeover attempt, even when it's just a returning user with bad password hygiene.
It absolutely works with custom auth! Your A/B test plan is the perfect approach. I'd go even further and suggest running it in log-only mode on 100% of your traffic for a week or two first, just to see what the match rate looks like without any user-facing impact at all.
Mapping the fields is the crucial part. Like others said, you'll be building a custom rule. The trickiest bit for us wasn't the email field, it was making sure the rule only triggered on a genuine login attempt and ignored other POST requests to the same endpoint, like password resets. We had to get clever with the conditions, checking for the presence of both username and password fields in the specific payload structure we use.
Seeing a few legitimate blocks was actually encouraging, it meant the system was working. It's always a user with a password they haven't changed in a decade. We handle it by logging a distinct security event and presenting our normal "invalid login" message, but we also flag the account to gently prompt for a password update on their next successful login.
hugo
So the trigger is only for genuine login attempts, not resets. That's a good catch. How do you differentiate them in your condition, just by checking if both fields are present? What about login attempts that are missing one field due to a client-side error, could that cause a false positive?
Yep, that's exactly what we do - check for the presence of both fields. Our condition looks for `$.password` AND `$.username` in the JSON payload. A client-side error that omits one would be ignored by the rule, which is actually what you want. It's not a genuine attempt, so there's nothing for the breach check to validate.
The tricky part, as you hinted, is making sure your condition is airtight. For us, that meant also checking the HTTP path and method, and making sure the request body isn't empty. We saw a few stray POSTs from health checks hitting the same endpoint that we had to filter out.
Have you looked at your recent access logs to see what the actual payloads look like for failed client-side validations? That'll tell you if you need to worry about those false positives.
Keep deploying!