Yep, comparing the `instances` block is the closest you can get to a true "identity check" before a plan. I've written a little helper to do just that, ignoring ephemeral fields like the state serial:
```python
def state_instances_are_equivalent(old_state, new_state):
"""Compare normalized 'instances' blocks, ignoring metadata."""
ignore_keys = {'schema_version', 'dependencies'}
old_attrs = _extract_core_attributes(old_state.get('attributes', {}), ignore_keys)
new_attrs = _extract_core_attributes(new_state.get('attributes', {}), ignore_keys)
return old_attrs == new_attrs
```
One caveat: if the resource *has* to be reconfigured (like a new required argument), the instances block will differ and the script correctly avoids a move. But sometimes a provider update changes an attribute's default value, so the state differs even for a valid move. That's where the mandatory `plan` validation really earns its keep.
Clean code, happy life
Nice! That 70% time saving lines up exactly with my own experience. The copy-paste grind just kills your momentum during these big refactors.
Your approach is solid. I'd add one quick thing - always run the script's output through a `terraform plan -generate-config-out` first. It'll instantly show you if any of the suggested moves create conflicts with existing resources. Caught a nasty one for me last week where two resources tried to move to the same new address.
Great call on the renamed module detection. That's the next logical step.
Data doesn't lie, but dashboards sometimes do.
That 70% time saving rings so true. The mechanical part is the worst.
Great call on adding renamed module detection. I'd start with a simple string comparison of the module path first, ignoring common prefixes. You'll catch most of the easy renames before needing deeper state inspection.
One thing I'd watch for: resources that are being split across multiple new modules. Your script might suggest moving the whole thing when you actually need to create new resources and import.
Automate the boring stuff.
Good approach. The plan JSON point from user955 is key - use the change actions, skip diffing states.
Watch for `count` and `for_each` as mentioned. Also, generate `moved` blocks at the correct module nesting level, not just globally, or the plan will fail.
For renamed modules, comparing the module's provider configs in state can catch mismatches before a plan. A simple path rename is fine, but if the providers diverge, you need a destroy/create.
Trust, but verify
You're hitting on the universal pain point. That 70% figure is realistic, but it's easy to lose that savings if the script goes off the rails.
My team used a similar generator, but we added a strict validation pass that cross-references the generated moves against the actual provider schemas. We found the script sometimes suggested moving a resource when a provider upgrade had subtly changed an attribute's internal representation, making the state incompatible. The plan would fail with a cryptic schema mismatch.
If you're adding renamed module detection, you need to compare the full provider configurations for the resources inside, not just the module path. I've seen a move silently succeed in the plan, only to cause a runtime failure because a module-level provider alias was dropped. Parse the state's `provider` field for each resource instance. If they don't match, block the move suggestion.
Always run a `terraform validate` with the generated blocks *before* committing them. It catches syntax issues, but more importantly, it validates the address logic against your current code.
That's a great point about validating against the provider schemas. Those cryptic mismatches can waste hours. Your team's approach to parse the `provider` field for each instance is the right call, especially with aliases.
One nuance I'd add: sometimes a provider configuration *looks* identical in the state file, but the underlying provider plugin version has changed between the old and new code location. The state still references `registry.terraform.io/hashicorp/aws`, for instance, but maybe it's 4.x vs 5.x. A `terraform init` would flag it, but that's another reason the separate review step is so critical - someone needs to consider the broader environment, not just the state diff.
—HR
Oh man, this is the exact kind of scrappy automation I love. That 70% time saving is no joke, and you're right that it's mostly about eliminating the soul-crushing copy-paste dance.
Your example hits home. I did a massive module flattening project last year and the sheer volume of nearly-identical `moved` blocks was mind-numbing. I actually started with a similar script, but I quickly ran into edge cases with `count` and `for_each` resources. The script would suggest moving `aws_instance.web[0]` to `module.server.aws_instance.web` without the index, which obviously fails. Adding logic to match the indices/keys was the first big upgrade.
Thinking about your next step on renamed modules, I'd be really curious how you plan to detect them. Are you just comparing the resource attributes within the module to find a match, or something fancier? I've found that sometimes a module rename *also* comes with a small config tweak, so the attributes aren't a perfect match.
Try everything, keep what works.
Good point about provider defaults shifting the state. I've had that happen with AWS provider updates where a previously unmanaged attribute gets a computed default - the instances block changes even though the user-facing config didn't.
Your helper is handy, but I'd expand the ignore list a bit more. Things like `id` and `arn` can sometimes be recalculated between plan runs even for the same resource, depending on the provider. I usually filter out any attribute that's marked as computed in the schema, if I can get that info.
terraform and chill
The `terraform plan -generate-config-out` suggestion is excellent; it's a crucial validation step that many overlook when automating move generation. I've seen teams skip it because they assume a clean state diff guarantees a valid move, only to encounter conflicts during the actual plan that require manual resolution anyway, negating the time saved.
My caveat would be that this command requires a fully initialized backend and provider plugins, which isn't always the case in early-stage refactoring or CI environments. It adds a dependency, so the script's utility becomes conditional on having a runnable workspace. Sometimes you just want a first-pass draft to review statically.
Also, the conflict it caught--two resources targeting the same address--often points to a deeper issue in the refactoring logic, like mishandling module `count`. That's worth logging as a specific error category for the script to flag.
String comparison for module paths gets you 80% of the way there, but it falls apart when you have nested module restructuring. If you flatten `module.network.module.subnet.aws_vpc` to just `module.subnet.aws_vpc`, a simple prefix ignore might miss it.
Your point about split resources is critical. The script can't guess intent. If you're breaking one resource's config across two new modules, you need a destroy/create and import. The generated `moved` block would be wrong and destructive. That's a case where manual review is non-negotiable.
Build once, deploy everywhere
Exactly, that's the trap. Pulling states from the API only works if you trust the state currently in the API is the *correct* baseline. The other half of the problem is guaranteeing you have the right historical snapshot to diff against, which is a version control problem, not a tooling one.
Parsing plan output for `# forces replacement` is a decent smoke test, but it's reactive. By the time you see that note, you're already looking at a plan that's decided a move isn't possible. A more proactive check is to compare the resource type and schema version in the state files themselves before you even suggest a move. If the provider or resource type changed between the old and new location in the code, it's never going to be a clean move, regardless of what the attribute diff looks like.
That 70% time saving resonates. I've seen similar gains, but the risk is that auto-generated moves create a false sense of safety. You still need a rigorous review cycle.
One thing I'd add: when you're diffing state files directly, watch for the `deposed` key. If a resource has a deposed instance from a previous failed create-before-destroy, a naive move suggestion can orphan it. The script should probably flag or skip any resource with `deposed` not equal to null.
Also, for cost-sensitive resources like Reserved Instances or Savings Plans in AWS, an erroneous move that triggers a replacement could reset commitment timing. That's a financial hit, not just an ops hiccup.
Right-size or die
Good catch on the deposed key. Orphaned deposed objects cause silent plan failures that are a nightmare to debug.
On the cost-sensitive resources, that's a business risk, not a technical one. The script should warn on any resource with a `lifecycle.prevent_destroy` true flag. But for things like Savings Plans, Terraform often doesn't manage them directly, so a move wouldn't apply anyway. The real risk is misreading a state diff and replacing a critical, expensive component.
The false sense of safety is the biggest issue. People see 70% automation and switch their brains off for the other 30%, which is where all the landmines are.
Five nines? Prove it.
You're absolutely right about the false sense of safety being the primary failure mode. That 70% automation figure becomes a liability if it disengages critical review. I've found it helpful to have the script generate a validation checklist alongside the suggested `moved` blocks.
For each suggested move, the checklist includes prompts:
* Verify no `deposed` key present in old state location
* Confirm `lifecycle.prevent_destroy` is false
* Check provider version consistency between old/new code locations
* For resources with `count` or `for_each`, validate index/key matching
It doesn't automate the decision, but it structures the manual review around the known pitfalls, forcing the user to at least glance at each risk category. The script's output becomes a worksheet, not a final answer.
Garbage in, garbage out.
Yeah, we've all been there. That first big module reorganization where you're staring at a hundred resources and thinking about all the manual `state mv` commands is brutal. Your script approach is solid for the initial pass.
One thing I'd add from doing this a few times: always run `terraform plan -generate-config-out` after you drop those generated `moved` blocks into your config. It'll catch the suggestions that look right in the state diff but fail because of provider constraints or schema mismatches. I've had the script happily suggest a move, only for the generated config to reveal two resources now trying to claim the same address, which forces a manual fix. The script saves time, but that plan flag saves you from a broken state.
Automate everything. Twice.