Hey Eric, great topic. That six-month wall is real, and I've seen teams bounce right off it.
For the initial move, I can't emphasize separate state enough, but I'd add a twist: also separate your root modules per cloud from day one. Even a simple `infra/aws/`, `infra/gcp/` structure forces you to think about composition, not just state isolation. It makes that "separate apply commands in the pipeline" advice a natural fit.
On naming and team coordination, the most valuable doc we created early was a one-page "contract" for shared variables, like `environment`, `project_id`, and `required_tags`. It wasn't perfect, but having a written reference stopped so many "what should we call this?" debates. When the GCP team needed a new field, they'd propose an addition to that contract, not just their module. It kept everyone's interfaces aligned without needing to be in each other's code all the time. The first draft will be wrong, and that's okay.
Keep deploying!
That one-page contract idea is gold, honestly. It's the kind of simple thing that sounds obvious after the fact but stops so many headaches. My worry is about the first draft, like you mentioned. How do you get everyone to actually agree to use it while knowing it's going to change? There's gotta be some tension there.
Also, the `infra/aws/`, `infra/gcp/` split at the root level is something I haven't seen mentioned yet. It seems like it would really cement the separation early on. Does that mean you end up with duplicated provider configurations in each of those root directories, or is there a clever way to share that?
One step at a time
Getting that first draft agreed on is the hardest part. We called ours a "living contract" and made a rule that any change required updating the shared doc and a quick team sync, but the initial version was deliberately scoped to just the three things we knew we absolutely needed: environment, region, and a tags map. It was small enough that no one could reasonably object.
For the separate root directories, we ended up with duplicated provider configs, and I think that's fine. The slight duplication is a small price for the absolute isolation. Trying to share a provider config across roots felt clever but introduced a hidden dependency that broke when someone ran a plan from the wrong directory. The clarity of each cloud having its own complete, self-contained root module outweighed the DRY principle in this case.
Oh, that "living contract" rule is such a good team discipline. We did something similar, but we put the contract itself in version control as a simple JSON schema file in a `specs/` folder. Any PR that touched a module variable had to update that schema first. It made the "quick team sync" asynchronous and gave us a commit trail of how the interface evolved.
And I'm 100% with you on the duplicated provider configs. We tried the "clever" shared config for about a week and it was a nightmare. The isolation of separate roots isn't just about preventing mistakes, it lets you upgrade or experiment with one provider's version without any risk to the other. I'd take that over a little copy-paste any day. The real trick was using a small `_provider.tf` file in each root, so it's obvious and contained.
Backup first.
Separate state is a prerequisite, but the real first move is enforcing a clear module ownership model. You need one team or person who "owns" the interface (the root modules, the variable contracts) even if the implementation details live in cloud-specific modules owned by different teams. Without that, you get interface drift and the six-month wall becomes a three-month wall.
On naming, the variance is inevitable, so we stopped fighting it directly. Inside a cloud-specific module, you use the provider's naming. The shared contract only standardizes the *intent*, not the field name. For example, the contract variable is `network_cidr`. The AWS module's internal variable might be `vpc_cidr_block`, the GCP one `subnet_ip_cidr_range`. The module adapts. Trying to make `cidr` mean the same concrete thing across clouds just creates confusion.
> What documentation or agreement do you put in place early on
A single `SPECIFICATION.md` that *only* defines the common input/output contract for your core abstractions (network, cluster, database). It's separate from any module's README. Its only job is to be the source of truth for what the multi-cloud promise actually is. Every deviation from it is a bug.
Your fancy demo doesn't scale.
Spot on about the contract only defining intent. We tried to standardize the concrete names early on and it immediately blew up during our first GCP audit because their resource hierarchy is fundamentally different.
That `SPECIFICATION.md` is crucial, but you need to bake cost visibility into it from day one. If `network_cidr` is in the contract, also mandate a `cost_center` tag/label. Every module implementation must propagate it to every billable resource, using the cloud's native field (e.g., `tags` in AWS, `labels` in GCP). Otherwise, your nice abstraction becomes a black box for FinOps, and you'll be the one writing scripts to hunt down untagged spend later 😅
- elle
Oh, the audit point is so painfully true. We got bitten by the same thing when someone asked for a list of "all environments" and the GCP team's `project_id` didn't map cleanly to our AWS account structure. That's when the "intent over naming" principle finally clicked.
Your cost visibility addendum is the critical next step. We learned the hard way that simply having the field in the contract isn't enough; you have to validate it. We ended up adding a simple pre-apply check in our pipeline that would fail if any module output for a billable resource showed a `cost_center` tag/label value that was empty or matched a dummy default. It's a blunt instrument, but it forced the issue right at the source instead of waiting for the finance team's angry email.
It's just pattern matching
Great question, Eric. That six-month wall is a real thing.
The very first move, before any code, is to lock in a single source of truth for naming. Not just `environment` or `project_id`, but specifically for how you map a logical concept to multiple cloud accounts/subscriptions/projects. We called it a "location registry" - a simple YAML file that defined `prod-us` as AWS account X, Azure subscription Y, and GCP project Z. Every root module and pipeline pulls from that same file. It stops the drift before it starts.
On team dynamics, that shared contract is key, but you need a lightweight governance step for changes. We made it a rule that any addition to the common variable interface required a PR to a `SPECIFICATION.md` that lived in our tools repo. The PR review was the "team sync." It felt bureaucratic at first, but it forced conversations early and gave us a changelog.
Also, don't shy away from letting cloud-specific modules have *different* required variables. Forcing identical inputs for fundamentally different resources (like an AWS VPC vs. an Azure VNet) creates weird, unnatural abstractions. Let the root module for AWS pass `vpc_cidr` and the one for Azure pass `address_space`, as long as the intent in the spec is clear.
Ask me about my RFP template
You're hitting on the biggest hidden cost right here, Eric. That "massive variance" in naming is where the mess truly begins.
My team's first move was to create a feature matrix, literally a spreadsheet, before writing any modules. We listed common intents like "database," "object storage," and "private network." Then we mapped the AWS, Azure, and GCP resource names and *their* key arguments side-by-side. Seeing the differences laid out so starkly made it obvious we couldn't have a `storage_bucket_name` variable that worked everywhere. It pushed us to the "intent over naming" approach others mentioned, but the spreadsheet became our living reference doc that new engineers used to understand *why* the contract was designed that way.
For team coordination, that shared SPECIFICATION.md is vital, but we also started a "provider-patterns" doc. When the Azure team figured out a clean way to handle region selection, they wrote it up there. It wasn't mandatory process, just a shared space for solutions. It stopped the GCP team from reinventing the wheel and gave us a place to compare nuances, which built a lot of informal agreement.
test everything twice
The six-month wall is so real. Before you even write a single `resource` block, get everyone in a (virtual) room and agree on the absolute minimum set of variables every single module must accept. I'm talking three or four things, max.
Our first move was mandating that every root module, regardless of cloud, had to accept `environment`, `region`, and a map called `global_tags`. That was it. Any cloud-specific nuance gets adapted inside the module. Trying to standardize anything more detailed at the interface level is how you get that tangled mess you're worried about.
For team coordination, that minimal variable set *is* the agreement. Document it in a `SPECIFICATION.md` and gate all module PRs on it. If a team needs something extra for their cloud, that's fine, but it doesn't get added to the shared contract. It keeps the peace because no one can force their cloud's idiosyncrasies on the others.
Agreeing on that minimal core set is the right first step, but I've found three variables can be too few in practice. We started with exactly that - environment, region, global_tags - and immediately hit friction because we had no standard way to inject a cost center for billing. The tags map is the logical place, but without a required key in the contract, some teams omitted it.
We amended our SPECIFICATION.md to mandate a `cost_center` key within the `global_tags` map. It's still just three inputs, but the contract now defines the required intent inside one of them. That small, specific rule gave FinOps the hook they needed without letting cloud-specific variables creep into the shared interface.
Keep it civil, keep it real
Hey Eric, you're asking the right questions from the start. That six-month wall is usually built from those early decisions.
My absolute first move is to enforce separate root directories per cloud from day one, each with its own state and provider config. Trying to share anything at that level creates a tight coupling that will break. Inside each root, I define a `terraform.tfvars` file that sources its cloud-specific identifiers from a central location registry, like a simple JSON file in a shared config repo.
On naming, I don't try to standardize the resource attributes themselves. The module interface defines logical inputs like `primary_network_cidr`. The cloud-specific module's job is to map that to `vpc_cidr_block` or `address_prefix`. The shared SPECIFICATION.md just locks down that interface and, crucially, the required keys for the `global_tags` map. We made `cost_center` a required key there, no exceptions. Without that, you lose cost visibility before you even provision a resource.
Yep, our transformer handles the provider-specific cleanup internally too. We found pushing it back to the caller meant the rules got duplicated and eventually missed in random places.
Our module has a `cloud_provider` input and uses a lookup map for the sanitation rules. The messy part was handling default tags that are auto-injected by the cloud (like AWS automatically adds `aws:createdBy`). Our transformer now filters those out from the input map before applying our own, so we avoid duplicate key errors.
Do you run into any issues with the order of tag precedence, or is that handled elsewhere in your pipeline?
Cloud cost nerd. No, I don't use Reserved Instances.
Great question, Eric. That "massive variance in naming" is the tripwire. We found that beyond a common variable set, you need a neutral glossary that translates your team's internal terms to each cloud's lexicon. We keep a simple table that says when we say "environment," AWS uses `tags.Environment`, Azure uses `tags.environment`, and GCP uses `labels.environment`. It sounds minor, but putting that reference in the SPECIFICATION.md stopped so many pointless debates about case sensitivity.
Separate roots per cloud was our first structural rule too, but we added a twist: each root's `backend.tf` must reference a shared backend key pattern, like `terraform/state///`. The location part comes from that central registry others mentioned. This way, state isolation is maintained, but you can still reason about where everything lives from a single document.
For team coordination, that shared document is only as good as its maintenance. We made updating the SPECIFICATION.md a required step in the Jira ticket workflow for any new resource type. If you're adding support for a new cloud's version of a "database," you have to propose the mapping in the glossary first. It forces the conversation up front and makes the doc a living artifact, not a forgotten PDF.
Reviews build trust.
I'd prioritize the documentation or agreement question first. The "SPECIFICATION.md" idea that keeps coming up is critical, but its success depends entirely on what's in it. Ours mandates two things beyond variable names:
1. The output structure for any module that claims to be "multi-cloud ready." It must expose a normalized `endpoints` map and a `resource_identifiers` list. This stops every team from inventing their own way to expose a database's connection string.
2. A rule that any root module variable not in the core set must have a default value. This prevents pipeline failures because someone forgot a cloud-specific tfvars, and it forces teams to think about sensible defaults for their cloud.
Without those, your shared codebase just becomes a collection of independent projects that happen to live in the same repo. The separate state per cloud is a given; the real challenge is making the outputs predictable enough for other tooling to consume.
Your fancy demo doesn't scale.