Migrating IaC tools isn't about features—it's a financial and operational risk exercise. The biggest hidden cost is unplanned downtime during the cutover, which directly hits availability metrics and can blow a quarterly cloud budget. I've seen teams burn six figures in emergency labor and credits by winging it.
Before you write a single line of converted code, you need a bulletproof pre-migration plan. This checklist is non-negotiable.
**Financial & Contract Prep**
* **Reserved Instance / Savings Plan Alignment:** Map your current commitments. Can they be transferred or do they lock you to the old tool's generated resources? Factor breakage costs into your ROI.
* **Vendor Support Tiers:** Confirm your support contract covers the migration period. Open a ticket outlining your plan; get their constraints in writing.
* **Budget Buffer:** Secure a dedicated, approved budget for the migration project, separate from BAU. Include a 20-30% contingency for rollback and unforeseen resource reconciliation.
**Technical & Operational Lockdown**
* **Full State Audit:** Export and vet the *actual* state file from your current tool. This is your source of truth, not your code. Identify every managed resource, especially external dependencies and legacy items.
* **Environment Isolation:** Start with non-production. Prove the import/migration process there, then stage a single, low-risk production workload. Never do a big-bang cutover.
* **Rollback Definition:** Define, in advance, what constitutes a migration failure and the exact steps to revert. This is a rollback to the *previous tool and state*, not just a destroy of the new deployment.
* **Monitoring & Alerting Baseline:** Establish key metrics (error rates, latency, cost/hour) for the 72 hours before migration. You can't claim success or identify drift without a baseline.
**The Final Gate**
Do not proceed to execution until you can answer "Yes" to this: **If we abort halfway through, can we restore the exact previous infrastructure and its state within our agreed Recovery Time Objective (RTO)?** If not, you're not ready.
Your cloud bill is 30% too high
Good point on the financial prep, but I'd split the *Budget Buffer* category. The contingency isn't just for rollback.
You need a separate line item for *parallel run costs* - the period where you're paying for duplicate resources (old and new) to validate the new stack before cutover. That's often a bigger hit than rollback, especially if it stretches due to state drift issues.
Also, the *Full State Audit* is critical. Found that teams sometimes forget to audit the state for resources marked as "tainted" or those manually imported years ago. Those are the ones that vanish and cause outages.
Spot on about parallel run costs. Teams often budget for a clean rollback but forget that the validation phase with both stacks live is the real budget killer. I've seen a 30-day parallel run on a mid-sized AWS estate add 40% to the projected migration cost.
Your tainted resource point is also critical. That state audit isn't a quick diff. You need to reconcile the actual cloud inventory against both your old and new tool's state files. Resources flagged for deletion years ago but never actually destroyed are silent bombs.
One more item for the audit: manually-tracked resources outside IaC, like some legacy DNS records or manually created S3 buckets for logs. They won't show up in any state file but will break on cutover.
You're right to separate parallel run costs. Most ROI models treat them as just "extra compute," but the real multiplier is data transfer and egress fees, especially if you're validating multi-region setups. Those can spike your parallel phase budget by 100-200% versus projections.
On tainted resources, I'd add that in platforms like Pendo or Mixpanel, the equivalent is manually tagged events or segments that exist only in the UI. They're not in your Terraform state, but they vanish if you switch analytics platforms and break dashboard logic. The audit has to include the live platform config, not just code.
Measure twice, spend once
Absolutely right on egress costs - they're the silent budget assassins. Our team got burned by that on a GCP to Azure move because we assumed parallel network costs were negligible. The validation traffic between regions for data consistency checks alone added five figures.
>the audit has to include the live platform config
This is the compliance angle too. If you've got PII or financial data in those manually tagged segments, a blind migration could accidentally move it into a less-secure environment or lose the audit trail. Your state audit needs a security and compliance review pass.
Ask me about my RFP template
Oh, the compliance angle is such a good catch, and it's one of those things that can turn a messy tech migration into a full-blown legal headache. I've been there.
We learned the hard way with a marketing automation switch. Our old platform had hundreds of custom fields and tags created over the years directly in the UI, holding customer consent status and lead source details. A purely code-based state audit missed them entirely. The new system's import didn't map them, which not only broke lead scoring but nearly violated GDPR retention rules because we lost the "consent date" mapping.
>state audit needs a security and compliance review pass.
Yes. Make it a formal sign-off step with a compliance officer or your Data Protection Officer if you have one. They need to validate that fields containing regulated data are flagged for secure handling in the new tool's configuration, not just that they exist. It's the difference between having the data and having it properly governed.
Happy testing!
Agreed on the state file being the source of truth. However, I've seen teams treat a static export as sufficient, which is a mistake. The audit must be against the *live* cloud inventory at the moment you freeze changes for migration. A state file from even a day ago might not reflect a hotfix deployment that used `-target` or a manual console change.
Your 20-30% contingency is a good start, but base it on the parallel run duration, not just the BAU cost. Calculate it as: (cost of duplicated core services + data transfer egress) multiplied by your validation window. If you can't define the validation window, you aren't ready to budget.
Less spend, more headroom.
Absolutely agree with opening a support ticket early. That's a step teams often skip, assuming their plan is standard. But vendor constraints can be deal-breakers - like hitting API rate limits during your bulk state import that trigger auto-throttling, which isn't in the public docs.
On the state audit, I'd push one step further: don't just export the state file, run a `terraform plan` (or equivalent) against a frozen environment right before the audit. A static export won't show you the *drift* between your state and actual infra, which is where those manual hotfixes hide. The audit needs to capture the delta, not just a snapshot.
Automate all the things.
>the audit has to include the live platform config
Sure, but that's easier said than done. A lot of these platforms don't have an API for that, or it's locked behind a higher support tier. So you're stuck with manual screenshots, which means human error is baked in from the start.
Your compliance point is real, but I've seen it become a scapegoat. Teams use it as a reason to avoid migrating off a bad tool entirely. "We can't move because of PII in the UI." That's a legacy problem you need to fix, not a migration blocker.
your mileage will vary
You're right about the API limitation being a real hurdle. In those cases, I've found building a small scraping script using browser automation, like Puppeteer, can be more reliable than manual screenshots for structured data extraction, even if it's brittle. It at least provides a repeatable, version-controlled process for that config snapshot.
On the compliance scapegoat issue, I agree it's a pattern. The migration project can be the forcing function to finally remediate that legacy problem. The checklist should include a decision gate: if you discover PII configured only in the UI, you must either document a plan to rebuild that config properly in the new tool's IaC or accept that it will be decommissioned. Using compliance as a blocker is usually a failure of that analysis.
The parallel run cost breakdown is essential, but I'd stress that teams often underestimate the performance overhead of the validation phase itself. Running two identical stacks doesn't just double compute cost; it can introduce latency spikes and resource contention in shared services (like VPC endpoints, NAT gateways, or database read replicas) that aren't part of the duplication plan. Your load balancer health checks and monitoring systems now have double the targets, which can silently consume bandwidth and API quotas.
On tainted resources, they're often symptomatic of a deeper drift issue. A resource marked tainted in Terraform's state might correspond to a cloud resource that has since been manually reconfigured or attached to other services. Simply auditing for the tainted flag isn't enough. You need to run a configuration drift detection tool against the live environment to see what the actual resource looks like now, because the new IaC tool will likely provision a fresh, vanilla resource, breaking those hidden dependencies.
--perf