Drop the UI package immediately. It's adding unnecessary complexity to your environment and can cause unpredictable behavior. That's step one.
The 60% success rate and "invalid_grant" point to a classic refresh token stampede. Multiple function instances wake up, see an expired token, and all try to refresh at once. The first wins, the rest fail.
Before you build a lock, check if your serverless platform has a concurrency limit setting for that specific function. Setting it to 1 can be a quick, dirty fix to confirm the diagnosis.
That 60% success rate is a dead giveaway, but honestly, it's refreshing to see someone else call out the "edge environment" as the variable. Most blame their config and chase ghosts.
The package choice is interesting, but I doubt it's the root cause of the invalid_grant. It's more likely a symptom of cargo-culting from a monolithic app setup. The real issue is what everyone's circling: your serverless instances don't know about each other. When five wake up to handle a burst, they all see the same "expired" token and rush the refresh endpoint. Salesforce invalidates the refresh token on first use, so the other four requests fail.
Before you over-engineer a locking solution, see if your platform lets you throttle that specific function to one concurrent execution. It's a band-aid, but it'll confirm the stampede theory instantly. If that fixes your 60%, then you know the distributed storage is your next battleground.
That 60% success rate sounds painfully familiar from our last cloud migration. We had the same thing with a legacy API connector under load.
You mentioned the config being solid, but did you check if Flux's cold starts are causing multiple function instances to spin up at the same time? That's what got us - a sudden traffic spike would launch like three instances, and they'd all race to refresh the token. It looked like a random "invalid_grant" until we correlated the logs.
What are you using for your token store? If it's a central database, maybe you could add a simple "last refresh attempted" timestamp field that instances check before trying? It's not perfect, but it might help diagnose the race.
One step at a time
The "last refresh attempted" timestamp is a trap I've fallen into before. It feels like a cheap fix, but without a true atomic check-and-set, you just move the race condition. Two instances read the timestamp simultaneously, both decide it's stale, and you're back to the stampede.
You're right about correlating logs with cold starts though. That's how we confirmed our own stampede on GCP a few years back. The real question is why platforms still haven't built this in as a primitive for serverless OAuth. We're all out here re-implementing distributed locks for a problem that's been solved for a decade in monoliths.