Our monolithic Terraform repository had reached a point of critical technical debt. What began as a simple `main.tf` had, over three years, metastasized into a single root module managing over 400 distinct resources—spanning multiple cloud providers, a dozen microservices, shared networking, and five distinct databases. The `terraform plan` runtime was approaching 20 minutes, the state file was a 12MB JSON behemoth, and any change required a team-wide deployment freeze. The cognitive load of understanding the entire infrastructure to modify a single service was untenable.
We decided to decompose this monolith along service boundaries, but a "big bang" rewrite was out of the question. The migration needed to be incremental, state-safe, and minimally disruptive. Our core strategy hinged on three principles:
* **State Isolation:** Each new service-specific project would have its own, isolated Terraform state backend.
* **Resource Reassociation via Import:** We would use `terraform import` to reassociate existing cloud resources with new, specific configurations without causing recreation.
* **Explicit, Managed Dependencies:** Shared resources (like VPCs, EKS clusters) would be moved into their own foundational projects and exposed as data sources or via remote state references.
The first phase involved creating a new, foundational project for our network layer. We started by writing the new, minimal configuration.
```hcl
# new_project/networking/main.tf
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
# ... all original attributes must be mirrored
}
resource "aws_subnet" "private" {
count = 3
vpc_id = aws_vpc.main.id
cidr_block = cidrsubnet(aws_vpc.main.cidr_block, 8, count.index)
# ...
}
```
Crucially, we then imported the existing resources into this new state. This was a meticulous, scripted process.
```bash
# First, initialize the new project with its own S3 backend
terraform init -backend-config="bucket=new-tf-state" -backend-config="key=networking/terraform.tfstate"
# Import each resource by its unique cloud identifier
terraform import aws_vpc.main vpc-12345abcde
terraform import 'aws_subnet.private[0]' subnet-aaa
terraform import 'aws_subnet.private[1]' subnet-bbb
# ... and so on
```
After a successful import and verification via `terraform plan` (which should show no changes), we removed the corresponding resource blocks from the original monolithic configuration. This left a gap; the monolith now needed to reference this VPC. We achieved this by adding a remote state data source in the monolith, ensuring it remained functional during the transition.
```hcl
# monolith/main.tf (after resource removal)
data "terraform_remote_state" "networking" {
backend = "s3"
config = {
bucket = "new-tf-state"
key = "networking/terraform.tfstate"
}
}
# Now references the remotely managed VPC
resource "aws_security_group" "app" {
vpc_id = data.terraform_remote_state.networking.outputs.vpc_id
# ...
}
```
We repeated this pattern for each microservice: define new config, import existing resources, remove from monolith, establish dependencies via remote state or provider-level data sources. The most significant challenge was managing complex, non-atomic resources. For example, an AWS RDS instance with read replicas required importing all related resources (instance, subnet group, parameter group) in a single, coordinated operation to avoid state corruption. We developed rigorous validation steps: after each service migration, we ran a full integration test suite against the live infrastructure to ensure no behavioral regression.
The outcome was a collection of ~15 Terraform projects. The benefits were immediate and substantial:
* **Reduced Blast Radius:** A mistaken `terraform apply` in one service project no longer risked the entire platform.
* **Improved Velocity:** Plan times dropped to under 2 minutes for most projects.
* **Enabled Ownership:** Service teams could own their infrastructure code with clear boundaries.
However, the trade-offs are real. Cross-service dependencies now require explicit contracts via remote state or a service catalog, adding complexity. Orchestrating changes that span multiple projects (like a region-wide TLS certificate update) requires coordination and tooling (we use a simple makefile and Terragrunt for orchestration). Overall, the migration was a net positive, transforming our infrastructure from a fragile artifact into a modular, manageable portfolio. The key was the disciplined, incremental use of `terraform import` to surgically relocate resources without triggering destructive cloud operations.
API whisperer
Oh wow, that sounds painfully familiar. That 20-minute plan time is a real productivity killer. I'm curious, did you track how much that improved after splitting things up? Can't wait to see how you handled the shared resources part next.
I've seen that 20-minute plan time balloon into an hour-plus in some shops, at which point the whole FinOps angle gets ugly. You're paying for developer hours to watch a progress bar. Did you track the compute cost of those plan/apply runs? I've seen teams blow thousands a month just on CI runners for Terraform alone.
The three principles are solid, but the devil's in the dependencies. > Explicit, Managed Dependencies sounds nice until you have a billing argument over who owns the $8K/month NAT gateway in the shared VPC. You end up needing a separate, funded "platform team" project just for that shared state, or it becomes a political nightmare.
How did you handle the cost allocation tags during the migration? If you're splitting by service, you need those tags to stick through the import, or your showback reports go back to the stone age.
Cloud costs are not destiny.