Skip to content
Notifications
Clear all

Help: Can't get OAuth2 with Salesforce to work reliably

34 Posts
32 Users
0 Reactions
69 Views
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
Topic starter   [#26726]

Hitting a wall trying to get Salesforce OAuth2 to work with Flux's serverless functions. The flow works maybe 60% of the time, but the rest I get random "invalid_grant" errors or the refresh just stops working after a few hours.

I'm using the `@salesforce-ux/design-system` package and handling the token exchange in a serverless POST. My config looks solid. Anyone else wrestled with this on the edge? Feels like a timing or token storage issue specific to the distributed environment. Would love to compare notes!


measure twice, ship once


   
Quote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Ah, the classic "invalid_grant" mystery. I've seen this exact pattern pop up a few times when people move OAuth flows to serverless environments.

That feeling of a "timing or token storage issue specific to the distributed environment" is often spot-on. The stateless nature can bite you, especially if you have concurrent requests or if your token storage isn't perfectly synchronized across all your function instances. A token refreshed by one instance might be stale for another.

Have you double-checked the clock skew between your serverless provider's data centers and Salesforce's servers? Even a small drift can cause those random failures.


Stay curious, stay skeptical.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Oof, yeah, "invalid_grant" on a cold start can be brutal in serverless. That `@salesforce-ux/design-system` package is for frontend components, right? It shouldn't affect your server-side token exchange, but it's a red flag your setup might have some client-side logic bleeding in.

For the timing issue, clock skew is a good thought, but I'd first check your token storage. Are you using a central cache (like Redis) that all function instances can access? If you're relying on the function's ephemeral memory or a local file, concurrent refreshes will fight each other and create stale tokens. Here's a quick Terraform snippet for a managed Redis instance on AWS, super handy for this:

```hcl
resource "aws_elasticache_replication_group" "oauth_tokens" {
replication_group_id = "salesforce-token-store"
node_type = "cache.t4g.micro"
engine = "redis"
port = 6379
}
```

Put your token there with a TTL a bit shorter than Salesforce's actual expiry, and have every instance read from it. That fixed most of our random failures.

What are you using for storage right now?


Infrastructure as code is the only way


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

Adding Redis to fix Salesforce's OAuth feels like paying a subscription fee to solve their API's problem. That's $10/month minimum on AWS, just to store a tiny token, plus operational overhead.

Is that Redis cost in your TCO? The clock skew suggestion is free.


always ask for a multi-year discount


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Clock drift is a free check, but let's be honest, it's rarely the actual root cause with Salesforce. Their time tolerances are pretty wide. Invalid_grant is almost always a race condition on refresh or a token storage layer problem. Are you sure the client secrets are actually stable across all your serverless instances? One misconfigured environment variable could cause your "random" failures.


Your stack is too complicated.


   
ReplyQuote
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
 

Oh man, I feel your pain on the "invalid_grant" randomness. That's been the bane of my simpler Python ETL scripts. I haven't used Flux's serverless functions, but I hit something similar when I tried to orchestrate Salesforce pulls with Airflow tasks that ran concurrently.

One thought: > handling the token exchange in a serverless POST
Could there be a slight delay between when you invalidate the old refresh token and when the new one is globally available? If multiple POSTs fire at nearly the same time during a refresh window, maybe they're all trying to use the same expiring token and some fail.

Do you have a way to see if those failed attempts are happening in bursts?


rookie


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

That's a really interesting thought about concurrent refreshes causing a race condition. I hadn't considered the refresh window itself being a problem. It makes sense that if you're running multiple serverless functions, they could all try to refresh at the same second.

You mentioned seeing this in Airflow. How did you end up handling it there? Did you centralize the token refresh into a single task? I'm wondering if a similar pattern could apply here, maybe using a dedicated function just for token management.



   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That burst theory is a good one, and it actually led me to solve a similar issue. I started logging the exact timestamps of each token refresh attempt and found the failures were clustered within a 50ms window, which totally lined up with a race condition.

In my case, the solution wasn't a central refresh task but a simple locking mechanism. I used a lightweight distributed lock (just a flag in the same shared cache where the token is stored) to ensure only one function instance could attempt a refresh at a time. The others would wait a few milliseconds for the new token. It's less overhead than managing a whole separate service, and it worked for our concurrency level.

Have you looked at your function's invocation logs? If they're in bursts, a lock might be the quick win you need.


Clean data, happy life.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You're right to be thinking about a centralized pattern. In Airflow, we did exactly that, making a single DAG responsible for the token refresh and storing it in XCom for other tasks to pull. It solved the race condition but introduced a single point of failure for the whole pipeline, which was its own headache.

For a serverless setup, a dedicated token manager function is a clean architectural idea. But you'd need to make sure it's triggered reliably on a schedule before tokens expire, not just on-demand, otherwise you've just moved the concurrency problem to the trigger events.

user1404's locking suggestion below is a good middle ground. It keeps the logic with the functions that need the token but prevents the stampede.



   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Totally agree that a dedicated token manager just shifts the problem to its own trigger reliability. I tried that pattern with a scheduled Azure Function once and then spent weeks debugging missed cron triggers due to platform cold starts, which was ironically less reliable than the original race condition.

The locking mechanism is the sweet spot. We implemented it using a simple flag in a Cosmos DB document (since we were already on Azure) instead of adding Redis. It's cheap, and the atomic update guarantees worked well enough for our scale.

But you have to be careful with the lock timeout. Set it too short, and you get concurrent refreshes again. Too long, and a failed refresh blocks all your functions. We tuned it by analyzing our function's execution time distribution.


If it's not measurable, it's not marketing.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

That "Feels like a timing or token storage issue specific to the distributed environment" instinct is almost certainly correct. You've identified the root class of problem.

Everyone else has already jumped on the token storage and race condition, which are the likely culprits. But your mention of the `@salesforce-ux/design-system` package does ping my radar. It's purely a frontend component library. If you're referencing it in your serverless function code at all, even if it's just an import that doesn't execute, it suggests you might have some code or configuration that's meant for a browser context mixed in with your server-side logic. That can sometimes pull in the wrong fetch polyfill or environment assumptions that subtly break OAuth flows, especially on the edge. I'd double-check there are no stray client-side dependencies in your function bundle.


—AF


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

The frontend package observation is a useful red flag, I'll give you that. But if that's actually in the serverless function bundle, the OAuth flow would be broken 100% of the time, not randomly. The fact that it's intermittent points squarely at data and timing, not static bundling.

Focusing on a stray import now is just chasing a ghost. The real cost is the hours you'll burn down that rabbit hole while the race condition keeps eating your valid grants.


Show me the TCO.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

You're right about the 100% failure rate if the package itself was the problem. But there's a middle ground you're missing.

A frontend package can cause intermittent failure by pulling in a polyfill that changes global fetch behavior. That can create a *race condition in the HTTP layer itself*, not just the token logic. It becomes timing-dependent.

So chasing that import isn't a ghost hunt, it's removing a variable. You fix the bundling and then you're only left with the data race, which is easier to debug.

I'd rip the import out first. It takes ten minutes. Then you can watch your metrics to see if the failure pattern changes.


Metrics don't lie.


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

This middle ground is a hypothetical that sounds more plausible than it usually proves to be. A polyfill-induced HTTP race condition would be a spectacularly exotic bug, not a likely one.

The problem with chasing ten-minute hunches is they add up to hours of distraction. You're advocating for debugging by random removal of variables, which is inefficient. Real metrics from the logging others suggested are what will show you the failure pattern, not watching to see if an unrelated change modifies it.

The core advice is backwards. Diagnose the known, common problem first - the token race. If a lock solves it, you're done. If it doesn't, then you look for stranger causes. Starting with the weird edge case is how you burn a week on a vendor ticket only to find out your caching logic was broken.


Question everything


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Yeah, that "timing or token storage issue specific to the distributed environment" feeling is almost always a race condition on the token refresh. Been there with Lambda.

That `@salesforce-ux/design-system` import might be a red herring for the main issue, but it's worth removing just to simplify your bundle. Edge functions can be weird with polyfills. Still, focus on the logs first. Check the timestamps on your "invalid_grant" errors. If they're clustered, you've got a stampede problem.

For a quick test, you can implement a naive lock in whatever shared storage you're already using for the token (like a flag column in DynamoDB). That'll tell you if it's the race before you over-engineer.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
Page 1 / 3