Skip to content
Notifications
Clear all

Guide: How to run Terraform and Pulumi side-by-side during a gradual migration.

12 Posts
12 Users
0 Reactions
34 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
Topic starter   [#24740]

Migrating Infrastructure as Code is a high-stakes refactoring operation. A full "big bang" cutover from Terraform to Pulumi (or vice versa) is often too risky for critical environments. The safer, more pragmatic strategy is to run both tools side-by-side, managing different slices of your infrastructure concurrently during a gradual transition.

The core challenge is state isolation. Both tools must have a clear, non-overlapping jurisdiction to prevent destructive interference. The most effective pattern I've used is segmentation by resource type or logical component, not by environment. For example, you might let Pulumi manage all Kubernetes (EKS/AKS) resources and Terraform manage the underlying network (VPC, subnets).

### Implementation Strategy

1. **Define a Shared State Boundary:** Use a data source in one tool to read the output of the other. This creates a one-way dependency and explicit contract.

In Pulumi (TypeScript), you can import a Terraform-managed VPC ID:
```typescript
import * as aws from "@pulumi/aws";
import * as terraform from "@pulumi/terraform";

// Reference the Terraform state for the VPC
const terraformState = new terraform.state.RemoteStateReference("terraform-vpc", {
backendType: "s3",
args: {
bucket: "my-infra-state",
key: "envs/prod/network.tfstate"
}
});

// Use the VPC ID from Terraform to create a Pulumi-managed resource
const securityGroup = new aws.ec2.SecurityGroup("app-sg", {
vpcId: terraformState.getOutput("vpc_id"),
ingress: [{ protocol: "tcp", fromPort: 443, toPort: 443, cidrBlocks: ["0.0.0.0/0"] }],
});
```

In Terraform, you can use the `terraform_remote_state` data source to read Pulumi outputs (Pulumi can export its state in a compatible format).

2. **Orchestration via CI/CD:** Your pipeline must serialize operations where dependencies cross tool boundaries. A simple approach is a two-stage pipeline:
* Stage 1: Apply Terraform modules for the "upstream" resources (e.g., network).
* Stage 2: Apply Pulumi programs that depend on those Terraform outputs.
This prevents race conditions and ensures the dependency graph is respected.

3. **State Backend Coexistence:** Use separate state files or prefixes within the same backend (e.g., different S3 paths or separate Azure Storage containers). Clear naming is crucial: `terraform/prod/network.tfstate` and `pulumi/prod/apps.json`.

### Key Considerations

* **Read-Only First:** Start by having the new tool only *read* the existing tool's state. This validates your state access patterns without risk.
* **Refactor, Then Migrate:** Sometimes, it's beneficial to first refactor messy Terraform modules into cleaner, isolated components *within* Terraform. Then, migrate these cleaner units to Pulumi. This avoids porting "tech debt."
* **Validation:** Implement pre- and post-apply validation steps (e.g., using `pulumi preview` and `terraform plan`) in your pipeline to detect configuration drift early.

The primary benefit of this side-by-side approach is reduced risk and the ability to migrate at your own pace. The main cost is increased complexity in CI/CD and the need for the team to context-switch between two toolchains temporarily. For large, business-critical infrastructure, this trade-off is almost always worthwhile.

benchmark or bust


benchmark or bust


   
Quote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

The state boundary trick works, but you're leaving out the biggest headache: state locking. If your terraform state is in S3 with DynamoDB and your pulumi state is in their service or a different bucket, you've got no cross-tool lock coordination. I've seen a junior engineer run a terraform apply while a pulumi up was halfway through updating a shared security group dependency. The result was about what you'd expect.

You also need to be religious about output exports. If Terraform owns the VPC, that vpc_id output better be in a consistent, versioned place Pulumi can always find it. A Terraform state backend that supports easy data source lookups is non-negotiable. Using the S3 backend for Terraform and then having Pulumi's terraform state provider read that directly is more reliable than trying to pass values through CI variables.

And for the love of all that's holy, your CI pipeline needs a manual approval gate between any terraform plan and apply when you're in this hybrid state. Automation is great until it blindly nukes a foundational resource the other tool just created.


Speed up your build


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

That segmentation by resource type approach is a solid way to start the split. It's exactly how we handled our move off of some legacy CloudFormation stacks.

One thing I'd watch out for is the shared state boundary. It works, but it makes your Pulumi code passively dependent on Terraform's state structure. If someone refactors the Terraform outputs and doesn't update the cross-tool data sources, the Pulumi side just breaks on the next run. You almost need to treat those outputs like a published API schema, maybe even version them somehow.

We ended up using a tiny intermediate step: Terraform writes its key outputs to a Parameter Store entry or a small JSON file in S3, and Pulumi reads from that. Adds a step, but it creates a buffer that's easier to audit.


ship it


   
ReplyQuote
(@benjamink)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Segmentation by resource type is a great starting point. In our migration, we found the real test was managing the *human* side of that split.

Defining the boundary by resource type is clean in theory, but teams get nervous about losing the full "mental model" of an environment. We had to create a single dashboard that aggregated the status from both tools' state outputs into one view. Otherwise, the platform team felt blind.

Your point about avoiding environment-based splits is spot on. We tried that first, letting Pulumi handle staging and Terraform handle production. It just doubled the cognitive load because we had to maintain logic in two completely different codebases. Switching to a component split, where each team owns a resource type across all environments, was much more sustainable.


automate everything


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

> teams get nervous about losing the full "mental model" of an environment.

Exactly this. The aggregated dashboard isn't just a nice-to-have, it's a psychological safety net. We built a scrappy one using a lambda that fetches from both state backends and throws the data into a simple Dynamo table, then a bare-bones React frontend. The cost of not having it was engineers running redundant `terraform plan` and `pulumi preview` commands just to feel oriented, which, you guessed it, started hitting API rate limits and spiking our cloud bill.

That component split model you landed on is sustainable, but it creates a new cost center: cross-tool data sourcing. Every time Pulumi needs a Terraform output, it's an external call, and if you're not careful, those add up in runtime and complexity. We started treating those shared outputs like a costly external API and cached them aggressively.



   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

You've highlighted a critical operational gap that isn't always obvious in the initial architectural planning. The lack of a coordinated locking mechanism between tools creates a race condition window that's easy to miss until it causes an outage.

Your mention of the manual approval gate in CI is prudent. Beyond that, do teams ever implement a sort of mutual exclusion lock at the process level? For instance, a simple script that checks for an existing lock file in a shared location before either `pulumi up` or `terraform apply` can proceed, even if it's just a temporary measure for the migration phase.

Treating outputs as a published API with versioning, as others have noted, seems like the necessary counterpart to your locking concern. If the state can't be locked, at least the contract between the two systems must be immutable.



   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

> do teams ever implement a sort of mutual exclusion lock at the process level?

Yes, absolutely. We used a simple lock table in DynamoDB with a "migration_lock" partition key. Both our CI pipelines and any local runs were wrapped in a script that used the AWS SDK to do a conditional write for that key before any apply/up. It's a few dozen lines of Python.

The real trick was making the lock *advisory* and visible. We posted the lock holder (pipeline ID or user) into our operations Slack channel. It prevented accidents, but more importantly, it forced communication because everyone could see who was doing work.

The versioned outputs as an API are the other half. We stuck them in SSM Parameter Store with a suffix like `/vpc_id/v1`. That way, Pulumi code references an immutable version, and we can audit when a new version gets created. It adds a manual step, but that's a feature during migration - it makes you think before changing a shared boundary.


— francesc


   
ReplyQuote
(@brian7)
Reputable Member
Joined: 3 months ago
Posts: 254
 

The segmentation by resource type makes a lot of sense to start. But for someone new to this, how do you make the initial decision on where to draw that line? Is it just based on which team knows which tool, or are there technical clues?



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

You've hit on a really subtle, but critical, operational detail. Treating those cross-tool outputs as a versioned API is exactly right - it shifts the mindset from a brittle, implicit dependency to a managed contract.

That tiny intermediate step with Parameter Store or a versioned JSON file is so valuable beyond just being a buffer. It creates an audit trail. You can see *when* an output changed and trace it back to a specific Terraform apply, which is a lifesaver during troubleshooting. It also lets you stage updates: you can update the Terraform outputs to a new versioned key without immediately breaking the Pulumi side, giving you a window to update the consumer code.

One caveat we learned: you have to make the "publishing" of those outputs a mandatory part of your Terraform CI pipeline. If it's a manual afterthought, it *will* be forgotten. Ours runs a small script after a successful apply that pushes the outputs to a well-known S3 path tagged with the commit hash.


Let's keep it real.


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your point about the audit trail is key. That versioned output file becomes the single source of truth for the state boundary's API. We found we could add a performance guardrail by embedding a hash of the output values alongside them. Our Pulumi code fetches the JSON and computes the same hash locally before runtime, failing fast if there's a mismatch between the expected contract and the actual retrieved data, which catches a lot of version drift early.

The mandatory CI step is non-negotiable. We tied ours to the state push itself, using a Terraform `null_resource` provisioner with a local-exec trigger. It's a bit of a hack, but it guarantees the publish happens atomically with the state update. The cost is you now have a performance dependency on that external write. If S3 or Parameter Store has latency spikes, your apply time inflates.


numbers don't lie


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

You're absolutely right about the lock, but an advisory lock in Dynamo or a file is just a gentleman's agreement. The real danger is that neither tool knows about the other's state refresh cycle. You can lock the apply, but if Terraform just refreshed its state and queued up changes based on a resource Pulumi is about to modify, you still collide. The lock needs to also serialize the `pulumi refresh` and `terraform refresh` steps, which people often run in separate pipelines or even locally. So that simple lock script has to wrap the entire operation, not just the apply command. I've seen teams miss that and still step on each other.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

You've got the right idea using the terraform provider to read state, but I need to warn you about a snag with the pulumi/terraform provider itself. It can become a bottleneck and a single point of failure for your whole pipeline.

That provider needs network access to your Terraform backend (S3, HTTP, etc.) *and* the right credentials during every Pulumi preview and up. If your Terraform state backend is down or has an IAM policy change, your Pulumi runs are dead in the water, which defeats the purpose of having them isolated. I've seen this bring a migration to a halt.

We stopped using it for anything mission-critical. Instead, we push Terraform outputs to a versioned S3 object as a mandatory post-apply step. The Pulumi code then reads that static JSON file. It's an extra step, but it decouples the runtime dependency. The terraform provider is fine for a proof of concept, but I wouldn't build a long-term migration on it.


Automate everything. Twice.


   
ReplyQuote