Skip to content
Notifications
Clear all

Terraform vs Pulumi for a 20-person dev team - which one caused fewer daily headaches?

2 Posts
2 Users
0 Reactions
2 Views
(@finnj)
Estimable Member
Joined: 1 week ago
Posts: 57
Topic starter   [#20383]

Alright, let's wade into the trenches. Everyone's screaming about "declarative vs imperative" like it's the holy war, but after shepherding a team through both, I've found the real pain points are far more mundane. It's not about which one is more "powerful"—it's about which one causes fewer engineers to contemplate a career in goat farming on a Tuesday afternoon.

We ran Terraform for two years. The state file became this eldritch horror everyone was afraid to touch. The sheer ceremony around `terraform import` for any existing resource felt like performing open-heart surgery with a spoon. And don't get me started on the module dance—writing a re-usable module felt like crafting a treaty, and then you'd get version drift because someone forgot a `required_providers` block. The abstraction was leaky, and the leaks were always filled with glue and prayers.

Switched to Pulumi six months ago. The immediate win? We could finally *reason* about our infra. It's just code. Need a loop? Write a loop. Need to transform some data? It's your familiar language (TypeScript, in our case). The cognitive load of context-switching between HCL and, well, *everything else* vanished. But—and there's always a but—you trade one set of headaches for another. Now you have to worry about dependency management, SDK versions, and the occasional "fun" of the Pulumi engine deciding your explicit deletion order is merely a suggestion. The state is conceptually similar, but at least the CLI feels less like you're diffusing a bomb.

So, which caused fewer daily headaches? For us, Pulumi. But only because our team was already neck-deep in TypeScript. The headaches shifted from "how do I express this in HCL" to "why is my node_modules this large," which was a trade we were willing to make. If your team lives in Go or Python, maybe it's different. If they're not coders, stick with Terraform and its particular brand of safe, frustrating sanity.

The real free alternative, though? Stop over-engineering. Half the "migration" was realizing we'd defined 500 lines of code for what could be a 20-line shell script and a well-documented README. Sometimes the best IaC is the one you don't write.

― Finn


FOSS advocate


   
Quote
(@avab)
Trusted Member
Joined: 6 days ago
Posts: 50
 

I'm a platform engineering lead at a 250-person fintech. We run over 600 services across AWS and GCP, and we've had Pulumi in production for two years after a prior three-year stint with Terraform.

1. **Cost Visibility and Sprawl**
Terraform Cloud's pricing felt opaque at scale, landing us in the $5-7/user/month band before we even accounted for the time spent managing runners. Pulumi's per-seat model was clearer, but the real cost was engineer hours. Our weekly Terraform plan review meetings were easily 8-10 collective hours for senior staff, mostly parsing HCL diff noise. With Pulumi, that dropped to 2-3 hours because the diffs are programmatically filtered.

2. **The State Horror vs. The Secret State Horror**
Yes, Terraform state files are a known nightmare, requiring strict locking protocols and still causing paralyzing merge conflicts. Pulumi's secret weapon is that its state backend is just as critical, but the console makes it feel managed. The catch? You are now irrevocably tied to Pulumi's backend or your own S3/Blob storage with a custom layer. The vendor lock-in is subtler but just as real. Migrating out of Pulumi's state format is a multi-week project.

3. **Module Abstraction vs. Package Management**
Writing a truly reusable Terraform module is a treaty negotiation, as you said. We had a "core networking" module that required 47 input variables. Pulumi's win is using actual package managers (npm, NuGet, PyPI). But this introduces a new headache: dependency management and security scans. We now have to audit our `package.json` for infra changes and have had two incidents where a transitive library update broke deployments.

4. **The Provider Gap and Maturity**
For the core AWS/GCP/Azure resources, both are fine. For niche SaaS providers (think Datadog, Okta, Fastly), Terraform providers are almost always more mature and updated faster. We hit a wall with Pulumi's Snowflake provider last year that lacked support for a critical resource; the workaround involved writing raw SQL in our Pulumi program, which defeated the purpose. The Terraform provider had it for six months prior.

I'd recommend Pulumi for a greenfield stack where your team is already strong in a supported language (TypeScript, Python) and you accept the backend lock-in. If you have significant legacy infra or rely heavily on niche third-party services, Terraform's provider ecosystem will cause fewer acute headaches.

Tell us what percentage of your infra is in the big three clouds versus other services, and whether your team's primary language is already Go or Python. That decides it.


Question everything


   
ReplyQuote