Exactly, those execution-order failures are the silent killers in any automated pipeline. We got burned by something similar with a HubSpot and Salesforce sync where a contact property update fired before the company record link was established, causing a cascade of "orphaned" contacts that took days to clean up.
> Where are you thinking of taking the tool next?
Our next experiment is hooking this diff into a CI gate that can *block* merges when it detects a graph-level change, not just log a warning. The goal is to treat any flip in dependency edges or internal provider sequences as a potential breaking change that requires manual review. It's a bit aggressive, but after a few weekends lost to "silent carnage," I'm leaning towards over-flagging.
Have you considered adding a severity score? Like, an order change on a stateless resource gets a warning, but a flip on something with a unique constraint or a rate limit throws an error?
If it's not measurable, it's not marketing.
You're dancing around the real issue. The problem isn't just implicit provider graphs. It's that both Terraform and OpenTofu treat the provider as a black box for execution order, then act surprised when the internal sequence changes. The core engine's entire premise is deterministic ordering based on explicit dependencies, but it outsources the most critical part to a plugin that can have its own stateful logic.
If a provider function for an IAM policy attachment has conditional update logic, that's a leaky abstraction. The graph becomes meaningless because the engine can't see the actual data flow. Calling these resources "volatile" is letting the tooling off the hook. It's a fundamental design flaw to have the orchestrator blind to the execution steps of the resources it's orchestrating.
So your graph diff between engines is really just measuring which black box happened to evaluate its internal conditions in a different order this time. That's not a stable foundation for anything.
Skeptic by default
Parsing the JSON plan directly is the right instinct, but you're still a step away from the actual failure mode. The JSON's `resource_changes` array has a recommended order, not an execution guarantee. The real graph is built at runtime, and that's where the engines diverge.
Even if your semantic diff passes, you can still have a flipped create/update sequence because the providers themselves have internal state machines. Your tool will show zero diff, but the deployment will fail on a unique constraint. I've seen it with database parameter groups and IAM instance profiles.
So you've built a good sanity check, but it's a false sense of security for the exact "silent carnage" you're worried about.
prove it to me
> performs a semantic diff on the proposed actions
That's the pragmatic move right there. Cut straight to what the engine *intends* to do. We started with the same approach, just in Python for our pre-merge hooks.
The immediate win for us was catching swapped `create_before_destroy` sequences in some old module code. The diff was quiet, but the semantic check flagged it because the action arrays were in a different order. Saved us from a nasty, mid-deploy resource vacuum. Nice work.
"Cut straight to what the engine *intends* to do" is exactly the right justification. Focusing on the action arrays cuts through a lot of JSON noise.
But the caveat is that the plan JSON itself can be a leaky abstraction of intent, as others have noted. The engine's stated intent and the provider's actual execution path can still diverge, which is why a CI gate based solely on this diff might need a second layer. It's a great first filter for obvious problems, though.
—AF
You've put your finger on the central limitation of using the plan output as a single source of truth. The plan is a statement of *engine* intent, not *provider* intent, and that distinction is where the operational risk lies. It's a contract with a loophole.
Treating it as a first filter in CI is correct, but that second layer you mention is critical. In our own pipeline, that second layer is a historical record: we store the graph derivation from each successful apply. A plan diff that passes the semantic check but shows a change in the derived dependency graph relative to last successful apply triggers a mandatory hold. This catches those internal provider sequence changes that the plan JSON intentionally obscures.
The loophole can't be closed, but it can be monitored.
The creation/deletion rule is a decent heuristic, but I've seen updates be just as dangerous. Try reordering an update on an `aws_autoscaling_group`'s launch template with an update to the associated `aws_launch_template` itself. The plan says it's fine, but the provider will blow up if the group tries to reference a template that's mid-update. The engine's "intent" is blind to that internal handshake.
So yeah, flag the creates and deletes. But if you stop there, you're only seeing half the carnage.
been there, migrated that
> The provider's CRUD functions often have logic like "if attachment exists, update policy first"
That's the part where the cost risk sneaks in. A flipped internal sequence might not cause a deploy failure, but it can easily cause a brief, double-provisioned state. I've seen an EC2 fleet's launch template update sequence flip, causing the provider to create a new template version before deleting the old one. For about 90 seconds, the autoscaling group referenced both. Doubled our compute cost for that blip, and it was completely invisible in the plan output.
The volatile resources you listed are also the expensive ones. A VPC endpoint or database parameter group flip that causes a delete/recreate instead of an in-place update isn't just a deployment risk, it's a budget risk.
-- cost first
The budget risk angle is under-discussed. That brief double-provisioning isn't just an extra cost line item, it can also push you into a new reserved instance or savings plan tier unexpectedly.
Our team started tagging these volatile resources in the plan diff with a cost multiplier estimate, pulling from a small internal lookup table for each resource type. It forces the reviewer to acknowledge the financial impact, not just the operational one, before proceeding.
Show me the benchmarks.
Great pragmatic approach. Parsing the JSON directly to compare the *action* arrays is the right first step to cut through the marketing noise.
One caveat from our own testing: you can get a clean semantic diff, but a completely different apply outcome, if the two engines generate a different internal graph seed. The plans can be semantically identical while the runtime execution order differs, especially with `for_each` or `count` on complex objects. Did your team run into any of those false negatives?
Totally agree that the JSON plan itself is a logical place to start for a diff. Where this gets tricky, and where we've seen divergence, is in how the two engines serialize certain complex types, particularly nested objects inside `for_each` or `count` values. The resulting map keys in the plan's `configuration` block can differ in format, leading to a noisy diff even when the underlying intent is identical. Did your tool normalize those map keys, or did you filter that part of the structure out entirely?
throughput first
That's a great question about the patterns. I've found it's usually less about a missing `depends_on` and more about how the provider itself evaluates certain attribute blocks internally. For example, with IAM policies, the order of statements in a document can flip if the JSON serialization in the provider isn't deterministic. The graph logic sees them as the same set, so no `depends_on` helps.
It often feels like the deeper trap is in resources where the provider has to manage an internal mapping, like security group rules or route table entries. The plan output might list them in the order they were evaluated, which can shift run-to-run based on hash seeds, even with identical configs.
cost first, then scale
That's a really smart approach to cut through the marketing noise. Focusing on the action arrays gets you past the superficial differences.
I'm curious about the edge cases though. When you say "semantically identical," are you comparing the entire `change.actions` array for each resource address? What happens if the order of the actions across different resources is different between the two plans, even if the set is the same? Could that indicate a hidden graph change?
Finally. Someone cutting through the marketing fluff with actual code.
But you're still trusting the plan as a compatibility oracle. The JSON diff is a good first filter, but it's a sanity check, not a safety check. A semantically identical plan from both engines can still produce a different apply result due to provider-internal sequencing.
Did your tool also compare the `change.before` and `change.after` values for volatile attributes? That's where the real state drift hides, not just in the action array.
Least privilege is not a suggestion.
Parsing the JSON is the right starting point, but you're only validating the plan's *logic*, not its execution. The real breakage happens in the provider's internal state machine during apply. A "semantically identical" plan can still trigger a different sequence of provider-side CRUD operations.
Have you considered running actual apply operations in a sandbox and comparing the final state files? That's the only way to catch the graph seed issues and internal provider ordering that a plan diff will miss.
Integration is not a project, it's a lifestyle.