Skip to content
Notifications
Clear all

Help: My function's environment variables are not updating on deploy.

26 Posts
24 Users
0 Reactions
48 Views
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
Topic starter   [#26086]

Alright, let's see if we can cut through the usual "just click harder" advice. I've got a Lambda function deployed via the Serverless Framework. The environment variables are defined in the `serverless.yml` under `provider.environment`. I update a value, run `sls deploy`, the deployment completes, but the function still uses the old values. The CloudFormation stack shows the new values. The function's configuration in the AWS console *sometimes* shows them, but the runtime does not.

I've ruled out the obvious:
* No, I'm not confusing stages or service names.
* No, I'm not hitting a cached version in my testing (tried multiple invocations, cold starts).
* Yes, I've checked that the IAM role has the necessary permissions (it does, because it worked initially).

This feels like a platform quirk masquerading as a feature. I'm on the standard provisioned concurrency setup. The only thing that seems to force a refresh is toggling the function's state or waiting an absurd amount of time (20+ minutes). So much for "serverless agility."

Anyone else run into this and found a reliable, non-manual workaround? Or is this just another case where the managed service abstraction starts leaking and we're left poking it with a stick until it behaves?


Show me the unit economics.


   
Quote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You're hitting a known, albeit poorly documented, behavior with Lambda's environment variable propagation. When you're using provisioned concurrency, you're dealing with pre-warmed execution environments that are initialized once and reused. The environment variables are baked into that initialization context at the time the provisioned concurrency environment is spun up.

The CloudFormation stack updates the function's *definition* correctly, which is what you see in the console, but the already-running provisioned environments retain the old variables from their init phase. The platform only guarantees the new variables for new environments, which only get created when the old ones scale down or are recycled.

The non-manual workaround is to include a version or timestamp in your environment variable names, or as a dummy variable, forcing a true configuration change that triggers a fresh provision. For example, add a `CONFIG_VERSION: ${sls:stage}-${timestamp()}` variable. This changes the function's configuration hash, prompting a full replacement of the provisioned environments. It's a hack, but it works within the platform's constraints.

You could also script a step after deploy to update the provisioned concurrency configuration, which effectively does a reset. But the dummy variable method is less brittle.


—BJ


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

That's exactly the behavior, and user1008 has the correct cause. The provisioned concurrency environments are the culprit. The console shows the *definition*, but the *running sandboxes* have a baked-in copy.

The timestamp-in-environment trick works, but it's a hack. The clean method is to trigger a version update that forces container recreation. You can do this by publishing a new alias pointing to the updated $LATEST, or, if you're using SAM or pure CloudFormation, you can set `AutoPublishAlias` and update the alias version. The Serverless Framework handles aliases differently, but you can manually bump the function's description property with a timestamp or hash. It forces a new version, which forces new provisioned containers.

Waiting 20 minutes is just waiting for AWS's internal garbage collection to cycle the old sandboxes. The alias update is immediate.


Your fancy demo doesn't scale.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

The alias update works, but it's not always immediate in practice. I've seen delays up to 90 seconds before the new provisioned concurrency containers are fully ready, even after the alias points to the new version. You get a brief window of mixed variable states during the cutover.

If you need deterministic behavior, you have to manage the cutover yourself. Use weighted aliases or deploy a new alias pointing to the new version, then update your traffic router after you verify the new containers are active.


Metrics don't lie.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

The delay you mention is real, but calling it a "brief window" is optimistic. I've seen mixed states persist for several minutes in production, which is an eternity for anything remotely transactional. Weighted aliases sound great until you realize you're now managing a traffic routing layer just to push a config change, which is exactly the kind of vendor-induced complexity I'm allergic to.

The real problem is that AWS sells this as a seamless deployment model while quietly offloading all the state management headaches to you. Provisioned concurrency turns a simple config update into a multi-step orchestration problem, and the documentation never quite catches up.


— skeptical but fair


   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Totally feel your pain on the vendor-induced complexity. It's why I started treating provisioned concurrency like immutable deployments. I bake the env var values into a small config file inside the Lambda package itself for anything that needs to be deterministic on deploy. It's an extra build step, but it bypasses the whole propagation race condition.

The real kicker? This all disappears if you're using something like App Config or a secrets manager, because you're pulling config at runtime anyway. But then you're paying latency for every cold start. No winning sometimes!


Always optimizing.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

That config file approach is a valid workaround, but it does lock you into a specific deployment artifact for each variable change. It's effectively a manual versioning system that bypasses the platform's configuration management, which can be brittle if you forget to update the file.

Your point about runtime config introduces a useful distinction. The trade-off isn't just latency versus determinism, it's also about the operational model. Using a config service like App Config decouples deployment from configuration, but it shifts the consistency problem to your cache strategy and secret rotation. You now have to manage TTLs and ensure your runtime client handles staleness gracefully.

A hybrid method I've used is to embed a configuration version hash as an environment variable itself, which the Lambda uses to decide whether to fetch fresh values from a central store on a cold start. It adds some logic but keeps the deploy artifact generic while avoiding the propagation delay.



   
ReplyQuote
 amyt
(@amyt)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Oh man, you've perfectly described the provisioned concurrency wall. That "platform quirk masquerading as a feature" feeling is spot on 😅

The reliable workaround for me has been adding a dummy environment variable with a timestamp on each deploy. Something like `ENV_BUILD_TIME: ${env:BUILD_TIMESTAMP}`. It forces a new version, which eventually spins down the old containers. It's still a hack, but at least it's automated in the pipeline.

It's wild that we have to game the system like this just to get a config refresh.



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

Oh wow, I didn't even know that could happen! I've just been getting started with a simple Lambda and I'm definitely saving this thread for later. The "platform quirk masquerading as a feature" line is too real 😅

So, to make sure I understand the workarounds right - if you're just using the Serverless Framework and not doing any extra version alias stuff, would adding a dummy timestamp variable in the YAML be the simplest first thing to try? I'm a bit scared of the whole alias/routing solutions the others mentioned. Seems like a lot just to update a config!



   
ReplyQuote
(@clarak2)
Estimable Member
Joined: 2 months ago
Posts: 143
 

Yep, the dummy variable trick is absolutely the simplest place to start if you're using Serverless Framework and don't want to touch aliases.

One small caveat: make sure the timestamp value actually changes on every deploy. I've seen pipelines reuse the same build timestamp across runs, which defeats the whole purpose. A commit hash or a real timestamp from the deploy step works best.

It feels silly, but it genuinely unblocks you. Once you start needing zero-downtime config flips, that's when the alias/routing complexity becomes necessary.


Docs save time


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Ah, the ol' "my config is right but the runtime is wrong" classic. You've nailed the culprit in your last line - it's absolutely provisioned concurrency. The CF stack updates, the console eventually shows the new env vars, but the pre-warmed containers are happily humming along with the old snapshot.

The timestamp hack everyone's suggesting works because it tricks the system into creating a new $LATEST version, which eventually replaces the provisioned containers. The catch is you have to make sure your alias is actually pointing to $LATEST. If you're using a fixed version alias, you're just shuffling papers.

If you really need agility, consider if you can drop provisioned concurrency for that function. Most of my "agile" config updates happen on functions where the cold start penalty is worth the sanity.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Yep, you've hit the exact wall. It's 100% the provisioned concurrency containers holding the old environment snapshot.

The dummy variable trick (like a timestamp) forces a new function version, which helps. But there's a subtlety if you're using semantic versioning in your pipeline. If your `serverless.yml` uses a fixed `version:` attribute, the new version won't become `$LATEST`, and your alias might not update. Make sure you're either letting Serverless auto-version or updating that manually.

It's a frustrating leak in the abstraction. Sometimes the simplest fix is to just temporarily scale your provisioned concurrency down to zero, deploy, then scale it back up. Not elegant, but it clears the stale containers immediately.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

That's a really good point about the fixed `version:` attribute. I ran into that exact issue last month! The deployment pipeline showed "success," but the alias was still pointing to the old version because we'd pinned it.

If you're stuck with fixed versions, you could add a small script to update the alias automatically. Something like this in your `serverless.yml` hooks:

```yaml
custom:
version: 1.2.3

resources:
Resources:
MyLambdaAlias:
Type: AWS::Lambda::Alias
Properties:
FunctionName: !Ref MyLambdaFunction
FunctionVersion: !Ref custom.version
Name: live
```

But then you have to manage that version bump yourself. It's a trade-off for sure. The scale-to-zero trick you mentioned is often the fastest path to sanity during a fire drill.


Clean code, happy life


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

Yeah, you've hit the main culprit. That "platform quirk masquerading as a feature" feeling is often provisioned concurrency containers holding onto the old environment snapshot.

While the timestamp/dummy variable trick forces a new version, there's another subtlety: if you're using any fixed function aliases for traffic shifting or canary deployments, you need to make sure that alias updates to point to the new version after the deploy. The Serverless Framework doesn't always handle that automatically depending on your config. You might see the new version created but your traffic still routed to the old, stale one.

The scale-to-zero-then-back trick is the most direct, if manual, way to clear the stale containers immediately. For a more automated pipeline workaround, some teams bake the env var values into a small JSON file within the Lambda package itself, but that's a whole different operational model.


Keep it civil, keep it real


   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You've perfectly described the classic provisioned concurrency problem! The dummy variable trick is the standard go-to, but let me add one more nuance based on my testing.

If you're using any sort of parameter store or secrets manager integration that fetches values at deploy time, make sure your dummy variable isn't being optimized away by the framework. I once used a commit hash that was identical in two back-to-back deploys because of a pipeline hiccup, and it didn't force the new version. So now I always append the actual deployment timestamp from within the CI job itself.

Have you checked if your provisioned concurrency is set on an alias that actually points to $LATEST? That's another spot where the abstraction leaks and can leave you with stale containers even after a successful deploy.


test everything twice


   
ReplyQuote
Page 1 / 2