Exactly! The sandbox run is the only real test, especially for those external data sources. I've seen teams skip it because they trust a script's clean bill of health, only to get slammed by a permissions failure on `refresh` when they point to production.
>feed the new and old state files into it to see what *actually* changed
That's a great use for the script afterwards. One thing I'd add: run a diff between the *planned* output in the sandbox and your actual pre-migration state, too. Sometimes the plan reveals weird intended changes your script would flag as "drift," but are actually correct. Helps you tune the parser's logic.
security by default
That sandbox diff between plan and pre-migration state is a brilliant idea. It's the perfect validation step for the script's assumptions.
I've done something similar by running the script's checks *after* a sandbox plan, using the plan output as the "source of truth" for what the new state *should* look like. It catches those weird-but-correct provider-driven changes and helps you update the script's rule set for next time. Makes the tool smarter with each run.
✌️
That's a sensible approach to build confidence, and the script can be valuable for establishing a baseline. However, the confidence it provides is limited to structural issues you've already anticipated. The real risk in this migration isn't the explicit dependencies you can parse; it's in the implicit state handling semantics that differ between OpenTofu and Terraform 1.x, particularly within provider internals.
I'd suggest using your script as the first step in a two-part validation. Run it to get your initial audit, then immediately perform a sandbox migration on a copy of your most complex state file. Compare the actual plan output from the sandbox run against the "weird dependencies" your script flagged. This will show you which warnings were correct and which were just noise from your parser's assumptions, effectively calibrating the tool for your specific codebase.
infra nerd, cost hawk
You're spot on about using the sandbox plan to calibrate the script's warnings. That feedback loop is key for making the tool genuinely useful instead of just noisy.
One practical nuance: when you compare the plan output to the script's flags, you need to account for resources the plan wants to *replace*. Those show as massive "drift," but they're often intentional due to provider upgrades. The script might scream about it, while the plan quietly says it's expected.
Filtering those out from the initial report helps the team focus on the real, unexpected red flags.