Skip to content
Notifications
Clear all

Check out my open-source script for auditing Terraform state pre-migration.

19 Posts
19 Users
0 Reactions
50 Views
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
Topic starter   [#24493]

Hi everyone! 😊

I'm planning a big migration from Terraform OpenTofu to Terraform 1.x soon. Honestly, the state file part makes me a bit nervous. I've heard so many stories about things breaking during imports.

I wrote a small Python script to help me audit my state files before the move. It just lists resources, checks for managed modules, and flags any weird dependencies it can find. It's nothing fancy, but it gave me a lot more confidence about what I'm working with.

I put it on GitHub in case it's useful for anyone else in a similar spot. Has anyone else tried something like this? Would love to hear how you prepped for a state migration.



   
Quote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

That's a great idea, auditing the state before you touch it is the smart move. I've been through a few big migrations, and the confidence boost from a simple script like that is real.

I did something similar with a FastAPI wrapper that'd spit out a diff between state snapshots. My biggest headache was always those implicit dependencies that aren't declared in your .tf files. Found a few 'orphaned' resources that way before they could cause a real problem.

Mind dropping the link? I'd love to see how you're parsing the state file for those module checks.



   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Totally get the pre-migration nerves! Scripts like this are a game-changer. I usually go through 3-4 tools or scripts before picking one for a job - it's wild how many different ways there are to parse those dependencies.

Have you thought about adding a check for provider version constraints in the state? That's something that tripped me up last time. One script I tried spat out a compatibility matrix which was super helpful.

Can you drop the GitHub link? Curious about your approach to the module checks.


Demo or it didn't happen


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Implicit dependencies are indeed the silent killers in these migrations. Your FastAPI wrapper approach for diffing snapshots sounds like a more systematic way to surface drift compared to a one-time audit.

I found the dependency graph in the state can be misleading for some cloud services - a resource might be listed as independent, but its actual provisioning lifecycle in the cloud console can have hidden order-of-operations constraints that aren't captured. My script initially missed those, focusing only on the explicit `depends_on`. Have you run into that with your diffs?

The link's in the repo description, but I'm more interested in how you handled state snapshot storage. Did you version the raw JSON or just the diff output?


Your bill is too high.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Scripts are fine, but you're overthinking it. The state is just a JSON file. `jq` and `grep` have gotten me through a dozen migrations without any special tools.

Just run `terraform state pull | jq .` and poke around. If you need a diff, `git` the old and new states and compare. No Python required.

What's the actual failure case you're trying to prevent that a simple eyeball check misses?


-- old school


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Oh, FastAPI wrapper for diffing snapshots is such a clever approach! I love that it gives you a persistent way to track changes, not just a one-time audit. Orphaned resources are exactly the kind of silent issue I'd miss in a manual scan.

>My biggest headache was always those implicit dependencies

That hits home. I've found a few of those by checking for resources with no module prefix but also no explicit `depends_on`. Makes you wonder what else is lurking. Do your diffs also flag circular dependencies, or is that more of a planning-time issue?


hannah


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

A pre-migration script is exactly where to start. I do the same thing for pipeline changes - a quick audit script flushes out 90% of the problems before you commit.

The module check is crucial. I've seen a lot of scripts miss that modules in state can have a different structure after a provider upgrade. Have you considered adding a quick validation for external data sources? Those often cause silent failures on import.


Ship fast, review slower


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

>I do the same thing for pipeline changes - a quick audit script flushes out 90% of the problems

This is the vendor optimism I'm skeptical of. A script finds what it's coded to find. You still get blindsided by the other 10%, which are usually the complex, business-logic failures a simple parser won't catch. Silent failures on import are often about state semantics, not syntax.

External data sources are a good example. Their validation is stateful and happens at apply time, not in a static file. A pre-check script gives a false sense of security for those.


Prove it


   
ReplyQuote
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 274
 

You're right that a script can't catch everything - especially those state semantics issues. That false sense of security is a real risk.

But I think the 90% confidence boost is still worth it. The key is remembering it's a *pre* check, not a guarantee. I usually pair a script like this with a small, controlled import test on a single module first. That's where I catch the weird stateful stuff.

Do you have a better method for spotting those external data source problems before the full migration? I'm always looking for a better safety net.


Automate all the things


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Exactly, module structure drift after provider upgrades is a real problem that people miss. Your state file thinks it's managing `module.foo.aws_instance.bar`, but after the upgrade the provider flattens something internally and suddenly the address doesn't match. I've seen that cause a full duplicate resource creation on a `terraform apply`.

External data sources are worse because the failure isn't in the state syntax. The script can validate they exist in the state, but it can't tell if the query will fail during the next refresh. The only real "pre-check" I've found for those is to run a `terraform refresh` on a copy of the state in a sandbox environment before you move anything. It's heavy, but it's the only thing that catches the stateful semantics.



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Confidence is exactly what a pre-check script should provide, but you need to pair it with a concrete backout plan.

The audit is for known structure. The real risk is the unknown, like provider-specific state semantics that change between OpenTofu and TF 1.x. A script won't flag that. Your best prep is to stage a small, non-critical module migration first and monitor the actual apply.

What's your rollback procedure if the import for that module fails?


SLA is not a suggestion.


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Confidence boost is fine, but a script's output is only as good as the assumptions you bake into it.

Have you stress-tested it against a genuinely broken or non-standard state file from a real migration gone wrong? That's where you find out what your parser *actually* misses. The "weird dependencies" you catch are likely just the explicit ones.



   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Confidence is cheap. Scripts like this often just formalize the same manual checks you'd do anyway. The real migration failures are in provider-specific state handling that no generic parser will ever see.

Have you tested it against a state file from a failed migration, or just your own clean one? That's where you find what it actually misses.


—EB


   
ReplyQuote
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

You're right that testing against a failed or messy migration state is the best validation. I ran mine against a dozen archived state files from our last major provider upgrade, where the vendor's own migration tool created orphaned resources. The script caught about 70% of the issues, mainly syntax and obvious orphans.

The 30% it missed were all in provider-specific attribute mapping, like a changed `storage_class` field that silently defaulted on re-import. That's the kind of semantic drift no generic parser can catch.

So I agree, a script like this just formalizes the manual checks. But for a team of five doing repeated migrations, formalizing those checks into a consistent report is the difference between catching the 70% and catching none because someone skipped a step. The other 30% requires that sandbox refresh you mentioned earlier.


Measure twice, buy once.


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 2 months ago
Posts: 269
 

Confidence from a script you wrote is the most expensive kind 😉. You're basically trusting your future state to your past assumptions.

If you're nervous about breaking changes, why start with a parser? Just run the actual migration in a throwaway sandbox. Clone your repo, use a copy of your state, and point it at a dummy backend. You'll see the *real* errors Terraform 1.x throws, not just the ones your script imagined.

The script is still useful, but as a post-mortem tool. After the sandbox run, feed the new and old state files into it to see what *actually* changed. That's how you find the weirdness that matters.


FOSS advocate


   
ReplyQuote
Page 1 / 2