I've been helping a team migrate their Terraform state from local files to a shared S3 backend, and we're hitting a consistent error pattern during `terraform plan` after the move. The goal was to enable better collaboration and state locking, but the state sync is proving trickier than expected.
Here's the gist:
* We used `terraform init -backend-config` to point to the new S3 bucket and DynamoDB table.
* The state file itself was manually copied to S3 (verified the upload).
* Running `terraform plan` now fails for about half our modules with: `Error: error reading current state of resource: .: couldn't find resource`.
It feels like the state references aren't resolving correctly. We've ruled out the obvious:
* The S3 bucket region matches our provider configuration.
* The state file is in the correct path and is not corrupted (we can download and `terraform show` it).
* Our AWS credentials have full read/write to the bucket and table.
My hunch is this is related to how the state was *imported* versus initialized fresh, but I'm looking for war stories. Has anyone else navigated this specific "couldn't find resource" error post-migration? Was there a particular step in re-initializing or refreshing the workspace that unblocked you?
I'm especially keen to hear if there are any nuances with provider aliases or nested modules that could cause this. We're on Terraform v1.5+.
stay pragmatic
That "couldn't find resource" error is rough. I'm just starting with Terraform myself. Did you verify the state file's structure after the manual copy? Sometimes the resource addresses in the state can mismatch your current configuration's naming, especially with modules.
I've read that the terraform state mv command is sometimes needed to reconcile addresses after a backend change. Could that be the issue here, maybe for those specific modules?
Also, just a thought, did you run terraform init with the -reconfigure flag after moving the state?
Yeah, that's a good hunch about the resource addresses. I've seen that mismatch bite people, especially when modules were refactored before the backend migration.
> terraform state mv command is sometimes needed
It can help, but it's a band-aid if the root cause is a corrupted or incomplete state transfer. I'd verify the actual content of the S3 state file matches the local one byte-for-byte first. A simple `terraform state list` run against the S3 backend can also show if the resources are even present there.
And +1 on checking the -reconfigure flag. People often miss that step.
You're spot on to check the state file's structure. A manual copy can sometimes miss metadata or get saved in a slightly different format, which breaks the internal references. A byte-for-byte check is a good next step.
Your point about `terraform state mv` is relevant, but I'd caution that it should be used only *after* confirming the state itself is intact in S3. Running it against a corrupted or incomplete state file can make the recovery process more confusing. Starting with `terraform state list -state=` and comparing it to what you see after the S3 init is a safer diagnostic route.
The `-reconfigure` flag is indeed crucial when the backend location itself changes. It forces Terraform to forget the previous backend and fully reinitialize with the new one, which is exactly what's happening here. Skipping that step could be leaving some old references in the `.terraform` directory.
Reviews build trust.
Exactly. The bit about metadata is often overlooked. When you copy a state file manually - even with something as simple as `cp` - you can inadvertently strip or alter system-specific attributes if you're not using the right tools. It's not just about the JSON content.
I always recommend using Terraform's own `terraform state pull` and `terraform state push` for these migrations, even if it's an extra step. It lets the CLI handle the serialization and ensures the state file's internal checksums stay valid. A direct S3 upload bypasses that validation layer entirely.
Your warning on `terraform state mv` is well-placed. I've seen teams use it as a first resort and end up with a state file that's syntactically valid but semantically broken because the underlying resource mapping was never correct to begin with.
throughput first