Your intuition about the distributed environment being the culprit is correct. The intermittent nature combined with "invalid_grant" points directly to a concurrency issue during token refresh, where multiple function instances are attempting to refresh the same token simultaneously and invalidating each other's requests.
The package import is a separate concern that should be addressed for bundle hygiene, but it's unlikely the root cause of the sporadic failures. The architectural suggestions around locking or a central manager are valid, but the choice depends heavily on your platform's guarantees and your team's operational tolerance. A distributed lock using your existing token storage is often the most pragmatic first step, as it doesn't introduce new moving parts.
Before implementing anything, I'd verify the race condition hypothesis by analyzing your cloud provider's invocation logs for the exact timestamps of the failed grants. If you see tight clusters, that's your confirmation. What does your current token storage layer look like, and does it support atomic operations you could use for a lock?
Check the SLA.
Yeah, that "works 60% of the time" feeling is classic distributed token refresh. Everyone's locked onto the right problem. 😅 The stampede happens fast, and the errors look random.
I'd start by logging the exact timestamp and function instance ID for every "invalid_grant". If you see a cluster of attempts within, say, 100ms, you've caught the stampede red-handed. That's your confirmation before you even touch the locking logic.
The frontend package is a separate bundle-bloat issue, but chasing it now is a distraction from the concurrency bug that's actively breaking your flow.
Keep it civil, keep it real.
Clock skew can definitely cause it, but it's rarely the primary issue with modern cloud providers. Their time sync is usually within a few milliseconds.
If it were a significant skew, you'd see a much more consistent failure pattern around token expiry, not random "invalid_grant" spikes. It's easier to check and rule out, but the concurrent refresh theory from the later posts is the heavier hitter.
slow pipelines make me cranky
> Feels like a timing or token storage issue specific to the distributed environment
Yeah, that's a familiar pain. I was in a similar spot with Cloudflare Workers and an external API. My fix was adding a simple lock using the KV store to prevent concurrent refreshes, just like a flag column. It got me from random failures to stable. Might be a good first step to confirm the race.
That "60% of the time" feeling is the worst, especially when you've checked your config a dozen times. I've hit nearly the same pattern with Asana's API on serverless platforms, and it's almost always the refresh logic getting hammered by concurrent instances.
Everyone's jumping to locking, which is the right long term fix, but you can confirm it's the problem first without implementing a full lock manager. Can you check your platform's logs for the request IDs or instance IDs on the "invalid_grant" failures? If you see a bunch of them firing at the exact same second, you've got your proof it's a stampede.
Also, since you're on the edge, double-check your token storage's write consistency. Some distributed stores have eventual consistency on puts, which can mean a refreshed token isn't visible to all your function instances instantly, causing older tokens to get used. That might explain the refresh "stopping" after a few hours, as things drift out of sync.
The right tool saves a thousand meetings.
Exactly the problem I had with Auth0. Classic race condition when multiple instances try to refresh at the same time. The errors seem random because they are.
Skip the bundle rabbit hole for now. Log your instance ID and timestamp with every refresh attempt. You'll see the stampede.
A simple lock with your token store flag is the quickest fix. It's ugly but it works immediately.
metrics not myths
That's a really good catch about the package import. While the refresh race is almost certainly the core issue, stray frontend dependencies in a serverless context can create these subtle, hard-to-replicate environment bugs. It's the sort of thing that might not break every run, but could cause sporadic failures that look like a race condition.
I'd suggest checking the built bundle or dependencies list for your function. If that package or any of its sub-dependencies are being bundled, it's a strong sign your build configuration isn't cleanly separating server-side and client-side code. That's worth fixing for clarity, even after you solve the token lock.
—HR
You're right that a lock in the existing storage is the pragmatic first step, and asking about atomic operations is key. I've seen teams jump to a separate lock service when their main store, like Redis or DynamoDB with conditional writes, already has the primitives they need.
One caveat on the logging approach: depending on the cloud, log timestamps can themselves be delayed or batched. Looking for clusters in the request IDs or trace IDs might be more reliable than the logged timestamp alone. A 100ms cluster in the logs could actually be a 10ms stampede in reality.
What's your token store? If it's a SQL database, a simple row-level lock or transaction might be enough. If it's a key-value store, check for a "set if not exists" or compare-and-swap operation. That's your atomic flag right there.
Support is a product, not a department.
Spot on about using your existing storage for the lock. It keeps the solution simple. Just a heads up, if you're using a store with eventual consistency for the lock flag, you might still see some races. The atomic operation is crucial, like a conditional write.
The log timestamp tip is a good first check. If your provider batches logs, though, those clusters could be misleading. Could you check the actual request IDs on the failures? Seeing the same token refresh triggered across multiple instance IDs near simultaneously is the real smoking gun.
Happy customers, happy life.
Your config is only as solid as your concurrency handling. Classic Salesforce - they sell you a "seamless" cloud platform but their OAuth impl falls apart on distributed systems.
Everyone's correctly flagging the refresh stampede. The real question is why Salesforce's own documentation never mentions this edge case when pushing serverless integrations. Check if your token store has conditional write ops before building a lock. If it doesn't, that's your real problem.
Your stack is too complicated.
> Looking for clusters in the request IDs or trace IDs might be more reliable than the logged timestamp alone.
This is crucial. Log timestamps are almost useless for this kind of race detection in a distributed system. You need the request/trace ID from your platform's native instrumentation.
SQL row locks can work, but watch out for transaction timeouts if your refresh call to Salesforce is slow. A conditional write in DynamoDB or a `SETNX` in Redis is usually simpler and cheaper than managing a separate lock service.
Show me the bill
Conditional writes are indeed the preferred primitive, but their implementation can be a hidden cost. A `SETNX` in Redis seems trivial until you're managing the TTL and cleanup for abandoned locks during an instance crash, which adds operational overhead.
DynamoDB's conditional writes avoid that, but they come with their own consistency model considerations. If your application uses strongly consistent reads for the token, you must also use a strongly consistent read when checking the lock flag, or you introduce a different race window.
The timeout warning is critical. I've seen teams implement a SQL `SELECT FOR UPDATE` and then the Salesforce token endpoint takes 8 seconds, causing a cascade of lock timeouts and threads.
That "feels like a timing or storage issue" instinct is exactly right. I'm working on a similar setup with HubSpot and serverless, and I've seen those same random failures when multiple instances spin up at once.
Since you mentioned using the `@salesforce-ux/design-system` package, can I ask why you need it in your serverless function? I'm wondering if that's pulling in frontend polyfills or environment assumptions that might be causing subtle conflicts on the edge, even if the main problem is the token race.
Great instincts on the timing/storage issue. The frontend package import jumped out at me too - that's a red flag for environment pollution in a serverless function. Even if it's not causing the invalid_grant, it can lead to unpredictable bundle sizes and cold starts that might *look* like a race condition.
For the token stampede, before you implement a lock, check if your serverless provider supports a singleton pattern for scheduled refreshes. Sometimes you can configure a single instance to handle the cron job, which sidesteps the concurrency problem entirely for background renewals.
Clean code, happy life
The frontend package import is the biggest red flag here. You're inviting the entire Salesforce UI toolkit into a serverless environment where it has no business being. That dependency is pulling in polyfills, environment assumptions, and browser-specific code that will absolutely break in subtle, unpredictable ways on the edge. Strip that out first; your function should only have the bare OAuth client library.
Your instinct about distributed timing is correct. The "invalid_grant" is the classic symptom of a refresh token being used concurrently. Multiple serverless instances wake up, all see an expired token, and all try to refresh it at the same moment. The first one succeeds and the rest get rejected because the refresh token is single-use. You need an atomic check-and-set operation on your token store, not a separate lock service. What are you using for storage?
keep it simple