We had a "platform modernization" initiative handed down from leadership. The mandate: replace our "legacy" CI/CD (Jenkins), IaC (a mix of CloudFormation and old Terraform), and monitoring (Nagios + some hacky scripts) with a "cloud-native" stack in one quarter.
The plan was to go from:
- Jenkins -> GitLab CI (on K8s runners)
- CloudFormation/Terraform 0.11 -> OpenTofu + Terragrunt
- Nagios -> Prometheus stack (Grafana, Alertmanager, Prometheus)
The forcing function was a security audit that flagged our Jenkins instance and Terraform versions. Instead of incremental upgrades, the architect decided on a "big bang" cutover. The sequencing was decided by dependency: IaC first, then CI/CD, then monitoring.
Where it slipped immediately: the team revolted in week two. Here's why.
* **Context loss:** The new Terraform module structure and Terragrunt wrapper introduced a 300-line `terragrunt.hcl` that nobody understood. The original author of the plan left mid-project.
* **Toolchain fatigue:** Debugging a GitLab pipeline failure meant understanding K8s pod scheduling, GitLab runner config, *and* the new Terragrunt flow, all at once. Productivity dropped to near zero.
* **No parallel run:** The old Jenkins jobs were decommissioned before the new pipelines were stable. We had a 48-hour period where no deployments could happen because of a gitlab-runner nodeSelector mismatch.
```
# Example of the "simplified" Terragrunt config we were handed:
include "root" {
path = find_in_parent_folders()
}
dependency "vpc" {
config_path = "../../../network/vpc"
mock_outputs = {
private_subnets = ["mock"]
}
}
terraform {
source = "git:: https://example.com/modules.git//app?ref=v3. 2"
}
inputs = {
subnets = dependency.vpc.outputs.private_subnets
}
```
This looks clean, but when it failed, the error trace was four layers deep in remote modules.
We forced a halt and regrouped. The compromise:
1. Keep Jenkins for existing services, but freeze changes.
2. New services only use the new stack.
3. Migrate monitoring last, after the other pieces are stable.
The lesson wasn't about the tools themselves—they're fine. It was about changing the entire feedback loop for engineers simultaneously. You can't ask people to re-learn their entire development and deployment workflow overnight and expect anything but revolt.
-shift
shift left or go home