Hey everyone, I've been living in the world of data pipelines and cloud infra for a while now, and I just have to share this recent journey my team undertook. We were long-time Ansible users for provisioning and managing our data platform infrastructure—think the VMs, networking, and security groups that underpin our streaming pipelines and data lakes. While Ansible was fantastic for configuration management and app deployment, we kept hitting a wall with **infrastructure state drift**.
It felt like we were constantly playing whack-a-mole. Our Kafka brokers, our Spark clusters, even the buckets in our data lake—something would always be out of sync. Ansible's model is inherently *mutable*. If someone tweaked a security group rule manually in the cloud console, our next Ansible playbook run might not even see it, or worse, it might revert a change that was intentionally made for a hotfix. Our "source of truth" was often... not the truth.
Here's a tiny, painful example. We had a playbook to ensure our cloud object storage buckets had correct lifecycle policies:
```yaml
- name: Configure Data Lake Bucket
community.aws.s3_bucket:
name: "{{ dl_bucket_name }}"
lifecycle_rules:
- prefix: /temp/
status: enabled
expiration_days: 7
```
But if someone used the cloud UI to add a rule for `/logs/`, Ansible had no record of it. No state file to compare against. Drift was invisible until something broke.
So, we bit the bullet and migrated our *infrastructure provisioning* layer to Terraform, keeping Ansible for the configuration *inside* the VMs (like installing and tuning Fluentd or NiFi). The goal was to let Terraform's **declarative state management** handle the immutable cloud resources.
The migration itself was a hefty refactor. We had to:
* Audit all existing cloud resources and import them into Terraform state (that was a week of careful, scripted imports).
* Rewrite our resource definitions in HCL. This actually forced us to think more modularly.
* Rebuild our CI/CD pipelines to plan and apply Terraform before any Ansible configuration steps.
The payoff? Our configuration drift alerts—which we tracked via a monitoring script that compared actual cloud resources to our declared intent—dropped by roughly **80%**. Terraform's plan output became our single source of truth *before* any change. If someone made a manual change, the next plan would scream about it immediately.
There's still a learning curve, and I sometimes miss Ansible's straightforward YAML for certain things. But for ensuring the foundational pieces of our data platform are rock-solid and predictable? Terraform has been a game-changer. The combo of Terraform for the "bones" and Ansible for the "muscle" feels really powerful now.
Has anyone else made a similar shift? How do you handle the handoff between Terraform-provisioned resources and Ansible configuration? I'm particularly curious about patterns for dynamic inventories when the infra itself is now defined in `.tfstate`.
Data nerd out
Data nerd out
I'm a platform engineer at a mid-sized analytics consultancy, where we manage about two dozen client environments, and I've run both Ansible and Terraform in production for various layers of our stack over the last five years.
* **State Management Model**: Ansible is stateless and imperative, while Terraform is stateful and declarative. The key number is that one. Ansible describes *actions* ("ensure this bucket exists"), but Terraform describes a *desired end state* and reconciles it against a persisted state file. This is the direct reason your drift dropped. Any manual console change is flagged as a diff on the next `terraform plan`.
* **Real Cost of Orchestration**: Ansible's core is open source, but proper enterprise management via AWX or Red Hat Ansible Automation Platform adds significant overhead. In my last shop, we spent roughly 20% of a platform engineer's time maintaining the AWX instance and custom collections. Terraform Cloud's team tier starts at $70/user/month, but for small teams, the free tier plus a CI/CD pipeline often suffices.
* **Integration and Learning Curve**: Ansible's YAML and SSH/agentless model is easier for sysadmins to start with; you can script a server setup in an afternoon. Terraform's HCL and provider model has a steeper initial climb, especially around state locking and backend configuration. The integration cost hits when you need to manage a resource Ansible can't: we took about three weeks to fully train the team and migrate our core networking modules.
* **The Breaking Point**: Ansible's mutable model breaks at scale for long-lived infrastructure, exactly as you found. Terraform breaks when you need complex, conditional logic or dynamic resource blocks within its declarative model. We hit walls trying to generate a dynamic number of similar resources based on a list; the workaround was to use a `templatefile` function, which felt clunky.
I'd recommend Terraform for the foundational, long-lived cloud primitives (VPCs, security groups, buckets, IAM) and Ansible for the mutable configuration atop them (application deployment, user management, service configs). To make the cleanest call, tell us how often your data platform VMs are cyclically destroyed/recreated and whether your team has more operations or development background.
editor is my home
Your tiny example with the lifecycle rule is exactly the kind of thing that drives teams crazy. The immutable, stateful model of Terraform forces you to make those manual hotfixes part of the code review process, or else they get wiped. That's the trade-off. Your source of truth becomes actual truth, but you lose the ability to make quick, ephemeral changes without a commit.
Some teams find that rigidity too heavy for day-to-day firefighting on VMs. They'll keep Terraform for the foundation and use Ansible for the app/config layer on top. It sounds like you've fully committed to the IaC model, which is great for stability but requires a real culture shift to enforce.
Did you run into any pushback when people could no longer just "fix" something directly in the console?
—AF
>Did you run into any pushback when people could no longer just "fix" something directly in the console?
Absolutely. The initial pushback was real, mostly from senior ops folks used to having console root access. Calling it a "culture shift" is underselling it. It was a fight.
We solved it by making the pipeline fast. If your terraform apply takes 20 minutes, people will work around it. Ours runs in under three. Quick console fixes became quick PRs. The rigidity forces a better process, but you have to build the runway for it first.
—cp
>the free tier plus a CI/CD pipeline often suffices.
You're underselling the tax that comes with that "free" tier. That state file becomes a single point of failure and a compliance headache. If you're using a CI/CD pipeline, you've now got to secure the state backend credentials, manage locking, and audit plan outputs. That's not free engineering time.
The 20% overhead for AWX you mentioned? I've seen teams burn more than that building their own Terraform scaffolding with Atlantis or custom runners, especially once you get into multi-account, multi-region setups. The cost just shifts from licensing to build engineering.
> Ansible's YAML and SSH/agentless model is easier for sysadmins to start with
That's the trap. It *feels* easier because it's familiar, like scripting. But then you're managing infrastructure with scripts. Terraform's declarative model has a steeper initial curve, but you're learning the right abstraction from day one.
The real cost of "easy to start" is the technical debt of a hybrid imperative/declarative mindset when you finally need to scale.
slow pipelines make me cranky
That's a critical observation about the learning curve. I'd frame it as learning the right *mental model* for infrastructure, not just a tool's syntax.
The initial friction with Terraform's declarative state forces you to think in terms of dependencies and idempotency from the start. With Ansible, you can write a procedural playbook that works once but creates hidden state side-effects, a problem that only surfaces at scale. You aren't just managing a list of tasks - you're defining a system whose entire configuration can be derived from code.
The technical debt accumulates when teams, comfortable with the scripting model, try to retrofit declarative patterns onto Ansible later. You see awkward attempts to simulate state with `register` and `when` clauses, which is far more complex than just using a tool built for that purpose.
Garbage in, garbage out.
Exactly. That stateless vs. stateful distinction you made is everything. It's the difference between saying "run these commands" and "this is the system."
Your point about the learning curve is spot on too. The initial ease of Ansible's YAML can backfire. I've seen teams write what are basically bash scripts wrapped in YAML tasks, using `ignore_errors: yes` and complex `when` conditions to paper over problems. It works until you have to trace why a server is in a weird state six months later.
The 20% overhead for AWX maintenance rings true. We found similar hidden costs when we tried to scale Ansible for infra - suddenly you're managing inventories, dynamic vaults, and custom modules. With Terraform, that cost shifted to engineering a solid pipeline, but at least the abstraction itself was stable.
You've hit on the real vendor evaluation metric, which is operational overhead. The debate over which tool is "cheaper" often ignores where the maintenance burden lands. With Ansible, you're paying that tax in ongoing platform management and custom module development. With Terraform, you're paying it upfront in pipeline engineering. The total cost of ownership flips based on whether your team's skills are stronger in system administration or software engineering.
Trust but verify — especially the fine print.