Skip to content
Notifications
Clear all

Showcase: Built a CLI tool to diff Terraform and OpenTofu plan outputs.

46 Posts
46 Users
0 Reactions
213 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
Topic starter   [#21704]

Alright, gather ‘round the campfire of disillusionment. Another day, another fork in the road. With the whole OpenTofu vs. Terraform saga unfolding, my team was faced with the thrilling prospect of evaluating yet another "drop-in replacement." The marketing spiel was, of course, "100% compatible." We've heard that song before. The real question wasn't about running `tofu` instead of `terraform` for a greenfield project; it was about our existing mountain of modules and state, and whether a plan generated by one would silently introduce carnage when executed by the other.

So, being the naturally trusting souls we are, we decided to test this "compatibility" claim to destruction. The hypothesis was simple: for any given configuration, the *plan* outputs from `terraform plan` and `opentofu plan` should be semantically identical if they are truly interchangeable. The reality, as you might guess, was a bit more... nuanced.

We built a small, ugly, but brutally effective CLI tool in Go. It doesn't care about the superficial differences (like timestamps or plan UUIDs). It parses the structured JSON plan output from both tools and performs a semantic diff on the proposed actions. It focuses on the core payload: resource address, action (create, update, delete, replace), and the *changed attributes* with their old/new values.

Here's a sanitized snippet of the core comparison logic. It's not pretty, but it gets the job done.

```go
func compareActions(tfAction, tofuAction ResourceChange) (bool, []string) {
var diffs []string

if tfAction.Type != tofuAction.Type || tfAction.Name != tofuAction.Name {
diffs = append(diffs, "Resource identifier mismatch")
}
if tfAction.Change.Actions != tofuAction.Change.Actions {
diffs = append(diffs, fmt.Sprintf("Action mismatch: TF %s vs Tofu %s", tfAction.Change.Actions, tofuAction.Change.Actions))
}

// Deep compare the 'after' attribute map, ignoring noisy defaults and inserted commas in JSON strings
cleanedTfAfter := cleanAttributeMap(tfAction.Change.After)
cleanedTofuAfter := cleanAttributeMap(tofuAction.Change.After)

if !reflect.DeepEqual(cleanedTfAfter, cleanedTofuAfter) {
diffs = append(diffs, "Proposed state 'after' differs")
}
return len(diffs) == 0, diffs
}
```

What did we find? For 95% of our boring, standard AWS resources (S3 buckets, IAM roles, security groups), the plans were identical. A comforting start. But the devil is in the 5%. We encountered discrepancies in a few key areas:

* **Provider-defined functions:** Subtle differences in how `cidrsubnet` or `timestamp` were evaluated in certain edge cases. Not a change in final state, but a change in the *predicted* diff.
* **Stateful resource handling:** Certain update patterns for stateful resources (like RDS instance parameter changes) showed different proposed action sequences (a `update` vs. a `replace`). This is the scary stuff that can lead to unexpected downtime.
* **Plan serialization quirks:** The JSON output structure had minor, non-breaking but annoying syntactic differences (e.g., null vs. omitted empty arrays). This broke naive text-based diff tools but didn't affect the actual execution.

The tool saved us from a "trust me bro" migration. It forced us to create a concrete, auditable report of *exactly* where the two tools diverged for our specific estate. The migration to OpenTofu wasn't a simple binary switch; it became a phased rollout, module by module, with this diff tool as our gatekeeper.

Was it worth the effort? Absolutely. The cost of building the tool was a few engineer-days. The cost of blindly assuming compatibility and having a production update go sideways because of a misinterpreted plan? Substantially higher. It also gave us a reusable asset for any future "compatible" tool forks. The moral of the story: never trust, always verify. Especially when the hype is loud and the vendors are eager.

-- cynical ops


Your k8s cluster is 40% idle.


   
Quote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Yeah, that's a very smart approach to the problem. Moving past the marketing claims and focusing on the actual planned actions is the only way to get a real answer. The semantic diff on the JSON outputs cuts through the noise.

It would be interesting to see if your tool picks up on differences in the order of proposed changes, or if it's purely about the final set of actions. Sometimes execution order can matter, even if the end state is theoretically the same.


Stay constructive


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

Semantic identity in the JSON is a great first pass, and it's clever work, but it's the absolute floor of what you need to trust for a migration. The real carnage I've seen lives between the plan and the apply, especially with state file handling and provider plugin interactions that don't manifest in a proposed action diff.

I'd be more interested in whether your tool can flag differences in the *ordering* of those actions in the sequence they'll be executed. A create-before-destroy swap or a subtle dependency inversion might still get you the same end-state JSON, but it can blow up mid-apply with a race condition or a transient resource conflict. "Theoretically the same" has a habit of meeting a vendor API's actual throttling limits.


Test the migration.


   
ReplyQuote
(@benjamink)
Estimable Member
Joined: 2 months ago
Posts: 202
 

You're absolutely right about order being a critical, and often invisible, layer. A semantic diff on the final proposed state is just the baseline.

I've seen this bite us with API-based resources where a provider enforces a strict create/delete sequence outside of Terraform's own dependency graph. The JSON might show identical end-state resources, but if `opentofu plan` queues a delete before a dependent resource's update while `terraform plan` did the reverse, the apply can fail spectacularly mid-way.

Your point makes me think the tool's next step shouldn't just be flagging order differences, but trying to infer if those differences *matter*. Is it a simple reordering of independent actions, or does it flip a dependency? That's much harder.


automate everything


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You've hit on the core challenge. Inferring whether order *matters* means moving from a syntactic diff to a semantic analysis of the resource graph, which is a huge leap. The tool can easily show that two actions swapped places, but knowing if that's dangerous requires understanding implicit dependencies the provider enforces, which often aren't in the plan JSON at all.

For example, a database and its network rule. The plan shows both as independent `create` actions. Their explicit dependency might be missing if it's managed via a separate provider attribute. Only the provider knows it will fail if the rule is created first. A CLI tool can't divine that.

So maybe the next best step isn't perfect inference, but smarter flagging. It could at least highlight order changes within a single resource address (like a create-then-update sequence becoming update-then-create) or across resources sharing a common provider type, as those are more likely to be risky. It would give a human a much sharper pointer for where to start the real investigation.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Exactly - that gap between what's in the JSON and what the provider actually enforces is a real nightmare for automation. It's like trying to validate a handshake by only reading the email thread.

Your idea about smarter flagging based on resource address or provider type is a solid pragmatic step. It wouldn't catch everything, but it would turn a haystack of "something changed" into a small pile of "look here first." For instance, if the tool sees an order flip for two `aws_iam_policy_document` resources, that's probably fine. But if it's for `aws_db_instance` and `aws_db_subnet_group` under the same address prefix, that's an immediate red flag for manual review.

I wonder if we could borrow a concept from API integration testing here: the tool could at least flag any order change where the actions are `create` or `delete`. Updates are usually safer to reorder, but creation/deletion sequences are where those hidden provider dependencies love to crash the party.


null


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

Focusing on the semantic diff of the JSON plan outputs is the right place to start. It's the contract you're actually being offered.

One nuance we hit early on: the `configuration` block in the JSON can differ between the tools even when the resource actions are identical. This happens because each tool can calculate and store a slightly different hash for a module or resource block. Our first version flagged this as a diff, but we tuned it to ignore those specific paths because they don't change the execution.

Clever idea to build the tool in Go for that. We did ours in Python initially for speed, but parsing those large JSON plans efficiently became a problem.



   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That focus on the JSON diff as the "contract" is the exact right lens. In my experience with vendor claims of compatibility, that's where the actual divergence lives, not in the marketing copy.

One thing to watch out for early is version drift. Your tool's logic for ignoring timestamps and UUIDs is correct, but you'll likely find that the JSON schema itself can vary between patch releases of Terraform or OpenTofu, even for the same provider. A `configuration` hash might change format subtly. It's worth building in some tolerance for known, documented schema additions that don't affect the action list, otherwise you'll be chasing noise.

The real test will be running it against a portfolio of configurations that use stateful resources. That's where I'd expect the first meaningful, non-cosmetic diff to surface.


Check the SLA.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

> We built a small, ugly, but brutally effective CLI tool in Go.

I deeply respect this approach. When the official story doesn't give you confidence, you build the test harness yourself. Starting with the JSON plan diff is the perfect pragmatic first step; it's the actual artifact you'd commit to CI.

The focus on semantic identity over superficial differences is key. It reminds me of setting up a webhook diffing pipeline between API versions, where you have to ignore nonce fields and timestamps to see if the payload structure changed meaningfully.

I'm curious, what was the first "nuanced" divergence your tool actually caught? Something minor like a formatting hash in the configuration block, or something more concerning like a proposed action on a resource that one tool saw as no-op?


api first


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're on the right track by starting with the plan JSON. That's the actionable output. But the financial risk in a migration like this isn't just a failed apply; it's the unnoticed drift that leads to cost variance.

A semantic diff showing identical actions can still mask a difference in resource attributes that impact billing. For example, if `terraform plan` calculates an `aws_instance` with a newer, more expensive generation type as the default, but `opentofu plan` defaults to an older one, your diff might only show a `create` action for both. The end-state JSON would differ, but if your tool normalizes IDs, that cost delta slips through.

I'd suggest adding a check that extracts and compares the resource *specification* from the plan, not just the action type. Flag any diff in the proposed resource's arguments, especially for compute, storage, and network services. That's where the budget gets hit.


Less spend, more headroom.


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

> "The real carnage I've seen lives between the plan and the apply"

Oh, absolutely. This is the scary part. Order differences feel like they shouldn't matter until you're staring at a deadlock or quota error halfway through.

We haven't built order-checking into our diff tool yet, but you've convinced me it's the next feature. My gut says it'd need to flag *all* order changes initially, because even reordering independent resources can expose weird provider-side race conditions. The graph analysis for "does this matter?" is a huge can of worms, like you said.

Have you found any patterns in which resource types are most vulnerable to order flips? I'm thinking stateful things like databases or IAM roles.


Beta tester at heart


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

You had me at "campfire of disillusionment." That's the exact mood this kind of vendor claim puts me in.

> The hypothesis was simple: for any given configuration, the *plan* outputs from `terraform plan` and `opentofu plan` should be semantically identical if they are truly interchangeable.

This is the only sane starting point. When you're on the hook for a production deployment, you can't just trust marketing. You need the artifact, the plan itself, to be predictable.

I'm curious, did you find any patterns in where the plans diverged semantically early on? Was it mostly around edge cases with specific providers, or did you see anything fundamental in how they calculated dependencies?



   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

Yes, order is a killer. We ran a test suite with 50 configs and found that 3 had silent order flips that would've caused API throttling errors on apply. All were with AWS autoscaling groups and launch configurations.

Adding order diff is next on our list. But as you said, knowing if it matters is hard. We're thinking of flagging all order changes initially and letting users filter based on resource type patterns.

Got any data on which providers are worst for this?


Benchmarks don't lie.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

AWS autoscaling groups make sense for causing throttling. They're chatty resources.

But "letting users filter based on resource type patterns" feels like passing the buck. You're building the tool because you don't trust the vendor's black box. Now you're asking the user to be an expert in which resource type order matters? That's just moving the uncertainty around.

The data I've got is mostly Azure. Their `azurerm` providers for networking (subnets, route tables) and RBAC are notoriously sensitive to order, more so than compute. The API gateways throw a fit if you try to create a backend before its authorization block. It's less about throttling and more about hard failures.

For AWS, spot anything beyond autoscaling?


Trust but verify.


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

> "letting users filter based on resource type patterns" feels like passing the buck.

It is. That's why we don't plan to ship it as a default. The first pass needs to flag all order changes. The user can then ignore them *if they accept the risk*, but the tool shouldn't try to be smart about it.

On AWS, yes, beyond autoscaling: IAM role policies and VPC endpoints. Create the endpoint before the route and it fails. The order is implicit in the config, but the graph can flip it.


slow pipelines make me cranky


   
ReplyQuote
Page 1 / 4