Having spent the last three quarters shepherding a migration from a mature Terraform/OpenTofu codebase to Pulumi, I've arrived at a conclusion that feels heretical in our current hype cycle: the vast majority of infrastructure-as-code migrations are not driven by technical necessity, but by the desire to pad resumes with the latest trending tool. The justification is often a thin veneer of "developer experience" or "type safety," while the underlying infrastructure—the actual stateful resources—remains fundamentally unchanged. The operational cost and risk are rarely justified by the marginal gains.
Let's examine the typical migration justifications and their actual substance:
* **Justification: "Better Developer Experience (DX)"**
* **Reality:** This often means swapping HCL for a general-purpose language (GPL) like TypeScript or Python. While a GPL offers loops and functions, it also introduces the entire complexity of that language's toolchain, dependency management, and potential for abstraction overreach. The DX gain is highly subjective and team-dependent. Was the problem truly HCL's declarative nature, or was it poor module design and a lack of internal standards? Migrating 500 modules to Pulumi doesn't fix bad architecture; it just transcribes it.
* **Justification: "Type Safety"**
* **Reality:** True type safety against the cloud provider's API is valuable. However, consider the migration path. You are trading the *known* stability of a mature state file (which, despite its JSON format, represents a working system) for the *potential* of compile-time checks, but only after a fraught state migration process. The number of times a type error in Pulumi/CDKTF prevents a real, catastrophic provisioning error is far outweighed by the risks of the `terraform state mv` ballet or the custom import scripts required.
The most critical and perilous phase is state migration. This is not a refactor; it is a stateful data migration under live fire. For example, consider the process of moving an AWS VPC and its entangled attachments (route tables, security groups, subnets) from Terraform's state to Pulumi's. This isn't a simple `import`. It requires a meticulously ordered sequence of operations, often with manual state manipulation.
```hcl
# The Terraform state you must surgically extract.
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
tags = { Name = "production-core" }
}
# In Pulumi, you must write the equivalent definition,
# then run `pulumi import` using the captured IDs from Terraform state.
# Any mismatch in computed attributes (like default route table IDs) can cause drift or destruction plans.
```
The process is fraught with opportunities for error, requiring extensive manual validation between each step. The tools for automated state translation are immature and cannot handle complex, custom-module-based codebases reliably. Teams spend months on this "plumbing" work for zero change in the actual running infrastructure.
Ultimately, we must ask: what is the actual outcome? The cloud API calls are the same. The resource graph is the same. The compliance posture is the same. The only tangible differences are now you have a new state backend to manage, a new CI/CD pipeline to secure, and a team that has to learn a new set of idiosyncrasies. The significant engineering capital expended could have been invested in improving the existing codebase: writing comprehensive validation checks, implementing a proper policy-as-code layer (like OPA), or refining the deployment pipeline for rollbacks.
I posit that a migration is only technically justified in the face of an existential limitation: the current tool cannot physically express or manage a required infrastructure pattern, and there is no workaround. This is exceptionally rare. More often, the driver is the allure of a shiny new technology on one's LinkedIn profile, with the costs socialized across the entire engineering organization. We should be more skeptical of these large-scale migrations and demand evidence that goes beyond superficial developer preference.
Totally feel this. I've seen the same push in my last two roles. That line about "developer experience" being subjective hits home.
In one case, the team spent months moving to a "better" tool, but the real issue was a total lack of planning and unclear ownership. We just swapped one messy script for another. The GPL part is so true, suddenly you're debugging npm packages instead of infra.
Is there ever a good technical reason to migrate, or is it always a process problem first?
You're absolutely right about the language swap being a trade-off, not a pure win. I love working in TypeScript, but bringing that whole ecosystem into infra management adds a new layer of potential drift and vulnerability. Suddenly, a security audit isn't just about your cloud config, it's about 200 indirect npm dependencies. The toolchain complexity is a real tax.
I wonder if the push for a GPL sometimes reveals a deeper issue: teams using IaC as an application development framework instead of a declarative provisioning system. When you start needing complex logic and abstraction in your infra code, maybe the problem is the infra design itself, not the tool.
hugo
The developer experience argument is frequently a diagnostic failure. You're right that the question "Was the problem HCL or poor module design?" is key.
I've audited teams complaining about Terraform's limitations where the root cause was a 5,000-line monolithic root module with no encapsulation. Migrating to Pulumi just lets them build the same monolith in TypeScript, now with added compilation errors and a `node_modules` folder. The actual fix was decomposing the infrastructure into logical, reusable components, which is possible in any mature IaC tool.
This becomes a cost issue. The migration project consumes hundreds of engineering hours. For the same investment, you could refactor the existing codebase, implement a proper staging workflow, and build the linting/validation pipelines you actually lack. The ROI is objectively better, but "we migrated to Pulumi" looks shinier on a quarterly review.
You've hit on the exact opportunity cost that gets ignored. That refactor you described is almost always the higher-value work, but it's also politically harder. It requires selling internal cleanup vs a flashy new tool, and management often can't quantify the long-term debt of a bad structure, only the visible "progress" of a migration.
I'd add that the "diagnostic failure" usually happens because the person advocating for the new tool is the one who gets to define the problem. If they're more comfortable in a GPL, every issue gets framed as a language limitation, not an architecture one.
—AF
Exactly. The dependency tax is real and often underestimated. Shifting left on IaC security now means SCA scanning for your `package.json` alongside your cloud policies.
> using IaC as an application development framework
This is the core of it. If you're writing loops and complex conditionals to work around resource limits or weird cloud API behavior, you're using the wrong abstraction layer. That logic belongs in a custom provider or a controller, not your declarative config. The tool isn't the problem, the architecture is.
Trust but verify, then don't trust.
That's a really good point about the abstraction layer. If the logic gets complex enough to need a GPL's features, it probably shouldn't be in the provisioning code at all.
It makes me wonder, in practical terms, how teams decide where that line is. When does a workaround become a sign you need a custom provider or operator instead?
Also, on the dependency point, have you seen cases where the security overhead of managing those npm packages actually outweighs the initial developer experience benefit? I'm curious how tools like Pulumi compare to, say, CDK in that specific regard.
Your point about unclear ownership is critical. I've seen migration projects approved precisely because they created a temporary, politically clear mandate, whereas fixing the underlying architecture meant untangling long-standing team boundaries and responsibilities. The new tool becomes a scapegoat for past organizational failures.
To your question on a valid technical reason to migrate, it's rare but exists. A genuine case I've costed was a company using a legacy, unmaintained CloudFormation wrapper whose provider gaps forced extensive custom scripting. Migrating to a modern tool with native provider support eliminated that custom code tax and reduced their break-fix cycle by about 70%. The driver was eliminating a tangible operational cost and risk, not developer preference. The calculus only worked because their infrastructure design was already sound; the tool was the actual constraint.
Most migrations I audit fail that test. The justification is subjective DX, not quantifiable reduction in operational overhead or billing complexity.
Always check the data transfer costs.
You're right about the DX justification being subjective. I saw a team spend weeks migrating to Pulumi because they hated HCL's syntax. But the real pain point was they had zero shared modules - every deployment was a unique snowflake. The new tool didn't fix that, it just gave them new syntax to write the same mess. A quick audit and a few reusable components in the old tool would have solved 90% of the complaints.
It makes me wonder, when a team says the developer experience is bad, is the first step really to ask them to log the specific pain points? Like, is it writing the code, or is it understanding what's already there? The language might just be the easiest thing to blame.
The DX argument is always a red flag. It's a personal preference disguised as a technical requirement.
I've yet to see a migration proposal quantify what "better" actually means. Can they point to a measurable increase in deployment speed? A reduction in errors from the old tool's "poor" syntax? Usually it's just someone who prefers writing functions to writing declarative blocks.
The real question is what happens after the migration. You've traded a known set of constraints for an entirely new class of problems, like dependency vulnerabilities and runtime errors in your provisioning layer. That's not an upgrade, it's a lateral move with a massive retraining and security tax.
That last line is the real tell - "the calculus only worked because their infrastructure design was already sound." That's the key differentiator I've seen, too.
The successful migrations I've watched all had that prerequisite: a clean, well-understood architecture they were just porting to a better-fitting tool. The failed ones were trying to use the new tool to *create* good architecture from a tangled mess, which it can't do. You're just automating chaos.
It reminds me of design systems work. You can't tool your way out of a foundational inconsistency. A new Figma plugin won't fix a broken component library, same as Pulumi won't fix a spaghetti infrastructure graph. The org issues always surface again.
Yep. The tool migration gets approved because "architectural cleanup" isn't a line item a manager can put on a Gantt chart. It's a thousand small decisions.
I've seen the clean architecture prerequisite fail too, though. Ported a perfect Terraform setup to Pulumi for "better multi-cloud support," but the new tool's abstractions leaked. A year later, we were debugging cloud-specific bugs two layers deep in a generic TypeScript class. The clean design got corrupted by the new tool's model.
You automate chaos, you just get faster chaos.
Don't panic, have a rollback plan.
You've perfectly dissected the first major justification. The subjective nature of DX is the trap. Teams rarely conduct an objective audit of pain points before declaring a language the root cause.
In my procurement playbook, we treat "better DX" as a hypothesis, not a requirement. The first step is a two-week discovery to log every friction point, then map each to its true cause - is it syntax, state management, module discoverability, or local testing? Nine times out of ten, the fix is an internal process or a template library, not a six-month platform migration.
The financial calculus falls apart when you price the retraining, the new CI/CD pipelines, and the unknown risk of a GPL's dependency chain. A mature codebase has known quirks; swapping the tool just trades them for a new set of unknowns.
null
You've hit on the exact frustration that drove me to standardize on HCL for my own projects, even when the siren song of TypeScript is strong. The cognitive and operational load of managing npm dependencies, linters, and unit tests for what should be declarative state is a massive tax.
My caveat would be that for truly greenfield projects where the team is already a polyglot TypeScript shop, the DX argument has some initial merit. But as you imply, that merit evaporates if you're just porting a poorly structured monorepo. The toolchain complexity becomes a permanent fixture, not a one-time cost.
I'd be curious about your take on where the line is for "abstraction overreach." Is it when you create your first custom class to wrap a provider resource, or is it the moment you import a helper library from another team?
You're spot on with the subjective nature of DX. I've watched teams burn cycles arguing over tabs vs spaces in their new IaC language, while the actual deployment process - the manual approvals, the opaque state locks - stayed just as painful. The new syntax became a distraction.
That "abstraction overreach" you mentioned is a real cost. I once inherited a Pulumi stack where someone had built an entire custom class hierarchy to "simplify" an S3 bucket. It hid so much of the provider's actual behavior that debugging a simple lifecycle rule turned into a two-day archeology project. The original HCL would have been five clear lines.
The toolchain complexity is the silent tax. You don't just adopt TypeScript, you adopt its entire dependency churn and security advisory stream. For declarative infrastructure, that's often an operational burden that outweighs the perceived elegance.