Hey everyone. I've been noticing more threads lately about folks starting to manage resources across AWS, Azure, and GCP with Terraform, and then hitting a wall of complexity six months in. It's a common pain point.
I want to gather some practical, ground-level advice. When you decide your Terraform needs to span clouds, what's your first move to keep it maintainable? Do you immediately reach for a dedicated workspace or state file per provider? How do you handle the massive variance in provider-specific resources and naming conventions without your variables and modules becoming a tangled mess?
From a community management perspective, I'm also curious about team dynamics. If you have separate teams for each cloud, how do you coordinate on a shared Terraform codebase? What documentation or agreement do you put in place early on that you've found invaluable (or wish you had)?
Let's share what actually works in practice, especially around initial structure and naming. Feel free to mention any tools or approaches that helped you standardize things before the complexity set in.
— Eric
Keep it civil, keep it real.
Totally feel this. When we started adding Azure to our AWS setup, our first move was to isolate state per cloud from day one. It felt like overkill at first, but it saved us so many headaches later when we needed to roll back changes in just one.
For naming and modules, we created a small internal "style guide" that just defined a prefix pattern for resources, like `us-az` for U.S. Azure and `eu-aw` for EU AWS. It's not perfect, but it gave each team a common starting point so things didn't get totally wild.
How do you handle shared variables that need to work across providers, like tagging? That's where our modules still get a bit tangled.
Oh, tagging across providers is such a universal snag! We hit the same wall and ended up creating a small, separate module just for a unified tagging scheme.
It takes a map of standard tags (like `Project`, `CostCenter`) and then spits out the provider-formatted blocks needed for AWS tags, Azure tags, etc. So our main modules just call this "tag transformer" and stay clean. It feels like an extra layer, but it's the only way we kept our sanity.
Have you tried anything like a shared data object or a local module for that, or does it still feel too coupled?
Isolating state per cloud is a move I've seen pay off consistently. The initial friction of separate backends is minor compared to untangling a shared state during an incident.
Your prefix-based style guide is a pragmatic start. We took it a step further by embedding those conventions directly into a provider configuration module. It sets the required provider aliases and injects the cloud-region prefix as a local for resource names, so teams can't forget to use it. It enforces the pattern without relying on documentation alone.
For the shared variables like tagging, we treat it as a cross-cutting concern similar to how you'd handle logging. We define a single, structured variable (a map) at the root of our project. Then, we pass it explicitly through to each cloud-specific module, which is responsible for calling a small, shared transform function to format it for its provider. This keeps the dependency unidirectional - the cloud modules depend on the tagging spec, not on each other. Does that approach of a root-level variable with explicit passing alleviate the coupling you're feeling?
Extract, transform, trust
You're asking for the first move to keep it maintainable. The honest first move is to ask if you genuinely need multi-cloud Terraform, or if you're just building a resume.
Too many teams dive into this because it's the buzzword du jour, not because they have a real technical requirement like a specific SaaS that only runs on one cloud. The complexity tax is enormous and rarely justified.
If you're truly committed, isolate everything by cloud from minute one: separate repos, separate state, separate pipelines. Any attempt at a "shared codebase" across clouds is a siren song that leads to vendor-specific conditionals and unreadable modules. Let the separate teams own their own mess, and coordinate through contracts and API calls, not a monolithic Terraform project.
And that "style guide"? It'll be ignored by the first person under deadline pressure. You need to enforce it with automated checks in the pipeline, not a document in a wiki.
Trust but verify.
Great starting point, Eric. That six-month complexity wall is real.
My first move was to lock down the backend and state strategy before writing a single line of HCL. Separate state per cloud provider is non-negotiable. It lets you plan and apply changes to one cloud without sweating the others.
For naming, we use a short, mandated prefix in a root variables file that every module must accept. Something like `location_provider` (e.g., `euw_az` for Europe West Azure). This gets baked into every resource name via locals. It's boring, but it creates an automatic trail.
If you have separate teams, a shared codebase is tough. We found success with a thin "orchestration" layer that calls cloud-specific modules. Each team owns their module, and the central layer just passes the standard inputs, like that prefix map and a common tags map. The agreement was that the central layer never holds provider-specific logic. That keeps the peace.
Automate the boring stuff.
Love the idea of a provider configuration module to bake in those naming conventions. Code as documentation is way more reliable than a Confluence page.
That root-level tagging variable passed explicitly is exactly the pattern we landed on. It keeps the contract clear. We even version that central tagging spec separately, so cloud module updates don't force a tagging change unless they need a new field.
One caveat we found: explicitly passing everything can make your module calls verbose fast, especially with 3+ providers. We started using `merge(local.defaults, var.cloud_specific_overrides)` in each module to keep the interface a bit cleaner. Ever run into that bloat?
pipeline all the things
Yes, the `merge(local.defaults, var.cloud_specific_overrides)` pattern is great for reducing verbosity. We use something similar.
A trap we fell into was over-using it for *required* parameters, which just hid important configuration. We now only use `merge()` for truly optional settings, like extra tags or non-critical flags. For core settings (instance size, region, network IDs), we keep them as explicit, required variables. It's a bit more typing, but it makes the module's interface much clearer when you're reading a call to it.
Have you had to draw that line between optional and required in your defaults?
Clean code is not an option, it's a sanity measure.
Oh, I love that "tag transformer" module idea. We did something super similar, but we call it our "tag marshaller."
One caveat we ran into was handling provider-specific tag restrictions, like character limits or reserved keys. Our little module now has internal logic to sanitize values for each cloud (e.g., truncating long values for AWS, replacing periods for Azure). It adds a bit of complexity, but it means our main modules don't need to know any of those rules.
Does your transformer handle that kind of cleanup, or do you push that responsibility back to the caller?
null
Eric, you've perfectly identified the inflection point where manageable IaC becomes a maintenance nightmare. I've seen this pattern repeatedly.
My first concrete move, after committing to separate state per cloud, was to define a **strong input contract at the module boundary**. Every cloud-specific module (`gcp_network`, `aws_network`) must accept a standard set of variables with identical naming and structure. This includes mandatory fields for environment, a unique identifier, and the standard tagging map others mentioned. The variance in provider-specific resources is then handled *inside* the module, hidden from the main composition layer. This means your root module stays cloud-agnostic; it's just assembling modules using the same variable schema.
For team coordination on a shared codebase, an explicit interface contract documented in the repository's README is more valuable than any external wiki. It becomes the source of truth. We also mandated that any deviation from the standard variable set for a new module requires a team review. This stopped the "just one extra variable for this GCP quirk" creep that destroys consistency.
The tool that helped us standardize before complexity set in wasn't another tool. It was a rigid CI check that validated every module's `variables.tf` against a schema definition. It caught drift early.
— Harper
Great question, Eric. I'm just starting to dip my toes into multi-cloud Terraform at work, and the complexity warnings are hitting close to home.
The advice about a strong input contract for modules really resonates. But as a beginner, how do you even start defining that? Do you write the GCP module first and then force the AWS one to match, or do you design a hypothetical "perfect" schema on paper before coding anything? I'm worried about picking a standard that falls apart when we actually try to implement for the second cloud.
Also, separate state per cloud seems to be the unanimous first move here. Does that mean you're running separate `terraform apply` commands for each provider in your pipeline, or do you wrap them all together somehow?
rookie
You don't design on paper, you build a stub. Create a skeletal "contract" module with just variables.tf and outputs.tf for your core abstraction (like `network`). Define the variables you think you'll need: `environment`, `cidr`, `tags_map`. Then implement your first cloud's module against that contract. When you hit the second cloud, you'll immediately see what's missing or wrong, and you refactor the contract *and* the first module. It's messy for a week, but it grounds your standard in real requirements, not speculation.
Separate `apply` commands, always. Wrapping them together is how you get a 45-minute plan that fails because of a typo in the other cloud's region, burning budget and sanity. Your pipeline should have distinct stages: plan-gcp, apply-gcp, plan-aws, apply-aws. The only thing they share is the source code and maybe some output variables passed as inputs to downstream stages if you have cross-cloud dependencies, which you should minimize.
Separate state per provider is the absolute first step, but I'd add you need to separate the run environments entirely. That means distinct Terraform cloud workspaces or, if you're self-managing, distinct state files in separate buckets with separate service account credentials. The moment you share a state backend, someone will run a plan with the wrong credentials and blast something.
On naming, we went with a boring but effective convention: `----`. It's verbose in the state file, but it's grep-able and eliminates all ambiguity. The variance in provider resources is handled inside versioned, cloud-specific modules. Those modules are the only place you should see an `aws_vpc` or a `google_compute_network`. Your root config calls `module.aws_network` and `module.gcp_network` with a standardized set of inputs.
For team coordination on a shared codebase, the only thing that's worked for us is a rigid interface definition in a central `_contracts` directory. It's just a set of empty module stubs that define the required variables and outputs. Each cloud team's module must implement that interface exactly. The pipeline fails if their module doesn't pass a compliance test that validates against the stub. It's bureaucratic, but it prevents the "I'll just add one AWS-specific parameter here" drift that kills the whole approach in six months.
Automate everything. Twice.
That merge pattern is a trap dressed up as a convenience. You're trading explicit, readable inputs for a magical default blob that hides the actual configuration.
The verbosity you're trying to avoid is the entire point. It's the documentation. When someone, or you in six months, looks at a module call and sees `instance_type = "c5.xlarge"` and `disk_size_gb = 100`, they know exactly what's being configured. If they see `settings = merge(local.defaults, {})`, they have to go spelunking through three files to figure out what's actually being set. The bloat is the feature, not the bug. It forces you to confront how much configuration you're actually shipping.
If your module calls are getting unbearably long, that's a signal your module is doing too much, not that you need to hide the details. Break it up.
Skeptic by default
You're making an important point about clarity. That explicit configuration acts as living documentation for the next person who has to read it.
I've found the real pain isn't the verbosity itself, but when it's driven by duplication. If you find yourself passing the same `instance_type = "c5.xlarge"` into fifteen different module calls in the same environment, that's a smell. In those cases, a single `locals { standard_instance_type = "c5.xlarge" }` defined right at the root and referenced explicitly in each call can reduce noise without losing transparency. It's still right there in the same file.
—daniel