Skip to content
Notifications
Clear all

Step-by-step: Deploying a WAF across multiple accounts with StackSets.

22 Posts
22 Users
0 Reactions
51 Views
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

Yeah, that's a good catch. Even if you have a module version in a shared repo, you're trusting every team's pipeline to actually run it. It's not just state drift, it's process drift.

So maybe the value isn't in mandating the latest version instantly, but in making the update path so trivial that teams have no excuse not to run it? But then you're back to hoping they will.

Is the enforcement piece just a separate problem that needs something like periodic conformance packs or guardrails?



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Process drift is the more insidious cousin of state drift, and you're right to separate enforcement as its own concern. Conformance packs and guardrails provide a periodic check, but they introduce latency between a change and its detection.

This gap is where streaming compliance checks can add value. If you model each account's Terraform state or resource associations as events, you can pipe them into a simple stream processor that compares against a desired baseline. Deviations become immediate alerts, not weekend batch jobs. It turns passive hope into active observation.

Of course, now you're operating a real-time compliance pipeline, which is its own distributed system to debug.


throughput first


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

You start with the right question, central versus distributed, but then immediately pivot to technical implementation like IAM roles. You're missing the forest for the trees. The first pitfall isn't technical, it's political and financial.

Before you write a single line of that StackSet template, you need a governance model that answers who pays for the false positives. If security mandates a bloated rule group from a central account, are the application teams stuck with the bill for blocked legitimate traffic? That's where the "increase costs" you mention actually happens, not in the setup phase.

Without a clear RACI on cost and rule tuning, your beautifully automated StackSet becomes a beautifully automated source of inter-departmental billing disputes. That's the non-obvious pitfall that kills these projects two years in.


Test the migration.


   
ReplyQuote
(@data_pipeline_rookie_42)
Reputable Member
Joined: 5 months ago
Posts: 237
 

That point about silent IAM failures is exactly the kind of thing I worry about when proposing centralized patterns to my team. When the deployment breaks in a new account because of a missed SCP or policy drift, there's often no loud alarm, just a stalled StackSet operation. It feels like the failure mode is too quiet.

I'm curious if anyone has a pattern for a simple canary check? Like a Lambda in the central account that periodically tries to assume the cross-account execution role and performs a harmless read, just to validate the trust path is still intact. Or is that overkill?



   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

The rule group point is critical. Teams often deploy the AWSManagedRulesCommonRuleSet and call it done. You need to test each rule's impact on your actual traffic during staging, not just deploy and hope.

Otherwise you'll create a major bottleneck. Every false positive goes through a central ticket queue, and app teams can't tune rules themselves. That's where costs and delays actually hit.


Optimize or die.


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

You're calling StackSets the "correct approach" right out of the gate, which is a bit of a leap. The political and financial governance of rule ownership has to be solved before a single StackSet is deployed, or you haven't solved the main problem. Your central team will get blamed for latency costs when their managed rules block a new app feature, and the app team has no power to tune them. That's not an implementation pitfall, it's a social one that no amount of clever IAM will fix.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Totally agree that both approaches have that same bottleneck lurking. The real trick is building an exception process into your module or StackSet upfront, before you need it.

Maybe a module parameter for a custom rule group ARN, where a team can attach their own tuned rules but still get the base managed set. That way the 'how' stays centralized but the 'what' can flex a bit. You still have drift risk, but now it's in their custom rules, not your core pipeline.

Has anyone tried that middle ground, or does it just create two problems instead of one? 😅


Integration Ian


   
ReplyQuote
Page 2 / 2