Skip to content
Did you see AWS IAM...
 
Notifications
Clear all

Did you see AWS IAM Access Analyzer now supports custom policy checks? Finally.

29 Posts
29 Users
0 Reactions
42 Views
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Your marketing automation example is a perfect fit. The shift from reactive audits to proactive, automated validation in CI/CD is the real win here.

We're implementing a similar check for our analytics pipelines, but with a focus on data governance. We've defined a custom rule that flags any policy statement granting `glue:GetTable` or `athena:StartQueryExecution` on our curated data catalog if it lacks a `aws:PrincipalTag/DataDomain` condition. This ensures contractors can only query datasets within their approved domain, even if the broader role is provisioned correctly.

The gotcha we hit was integrating this into our Terraform Cloud runs. The validation happens during the plan stage, but the policy resource might not exist yet for the analyzer to evaluate. We had to work around it by extracting the rendered JSON policy from the plan output and passing it to `ValidatePolicy` as a string. It adds a step, but it's better than catching it post-apply.


Garbage in, garbage out.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

For the data lake use case you mentioned, we wrote a custom check that flags any S3 or Lake Formation permissions on our raw zone without a `aws:PrincipalTag/Env=Prod` condition.

Integration is straightforward with our DuckDB-based pipeline orchestrator. We export synthesized IAM policies from the manifest and run `ValidatePolicy` before the orchestration step pushes to AWS.

The managed rules are still useful for baseline hygiene, but they miss our internal data classification requirements.


Numbers don't lie.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Good call on combining S3 and Lake Formation checks. We had to split ours into two separate rules because the condition structure for Lake Formation's `DataLakePrincipal` was different.

What's your validation latency with DuckDB? We see 200-300ms per policy with moderate complexity, which adds up across hundreds of policies in a single manifest.


Metrics don't lie.


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Your point about moving from reactive audits to proactive validation is exactly the architectural shift this enables. The real operational lift is integrating it into the resource provisioning lifecycle before a policy is even applied.

Based on our implementation, the main gotcha with CloudFormation or Terraform is state dependency. If you run the custom check during the plan or `cfn-lint` phase, the IAM resource often doesn't exist yet in AWS for the analyzer to validate. We had to implement a two-stage check: a static analysis of the local policy document first, followed by a post-deployment validation step that runs the analyzer against the actual live policy. It adds complexity but catches drift.

For your marketing automation use case with contractors, I'd suggest pairing the custom check with a programmatic remediation workflow. For example, if a policy fails the "no public S3" check, the pipeline can auto-rewrite the `Principal` field to a restricted ARN rather than just failing. That reduces friction for your devs while still enforcing the guardrail.


infra nerd, cost hawk


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

This is the same "game-changer" pattern every AWS announcement gets. It moves the lock-in to a different layer.

> Now I can bake in checks to make sure no one accidentally creates a policy that grants public read access

You're describing a procedural failure, not a tooling gap. The real problem is that your marketing devs and contractors have permissions to create IAM policies that can expose a customer data lake. This feature is a bandage for a security model that's already too permissive.

The gotcha is you're now building your own compliance engine inside AWS. When they deprecate an API or change a quota, your guardrails break and you're back to reactive audits, just with more code to maintain.

Anyone integrating this with Terraform is just adding another stateful dependency to their pipeline, as others have pointed out. It's complexity masquerading as control.


Trust but verify.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 2 months ago
Posts: 435
 

Oh, so a vendor finally caught up to writing your own basic policy validation? Shocker.

You're celebrating a new place to define company rules because your contractors have permissions they shouldn't. That's not a tooling win, it's a process failure you're now automating. The "game-changer" is you writing more glue code to fix a security model you built wrong in the first place.

And good luck baking that linter into Terraform. If the resource doesn't exist yet in AWS during the plan phase, what exactly is the analyzer checking? Now you need a two-stage validation dance, adding more state and failure points. Sounds like a real productivity boost.


Trust but verify.


   
ReplyQuote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

I get your skepticism about adding more process layers, but you're missing the practical context.

That "procedural failure" you mentioned is often a fixed business constraint. We can't just strip IAM policy creation from dev teams building on tight deadlines. This custom check feature gives us a safety net that's actually enforceable at the deployment stage, which is a step up from just hoping our docs get read.

On the Terraform point, you're right that state dependency is a pain. But our workaround isn't a two-stage dance, it's just running the analyzer as a post-apply check in the pipeline. It fails the build and auto-reverts. It's not perfect, but it's way better than catching a misconfiguration in next month's audit.


Let the machines do the grunt work


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Finally is right, but the real work starts now. Writing those custom rules is deceptively tricky, especially if you're trying to make them broad enough to be useful but specific enough to not cause false positives every other commit.

You're spot-on about layering it with managed rules. Keep the managed ones for the obvious AWS best practice stuff, then use custom rules for your org's specific headaches, like that `Principal: "*"` on internal buckets. The trick is to start small with one or two high-impact rules, otherwise you'll spend weeks fine-tuning JSON logic.

The CI/CD integration is the obvious next step, but the state dependency issue others mentioned is real. Running it post-apply and failing the pipeline is a valid, if blunt, instrument. Just make sure your rollback actually works.



   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

That's a clever solution with SSM Parameter Store. It adds a cross-account permission, but you're right, it's a one-time setup for a huge reduction in drift risk.

But doesn't pulling the ARN from SSM in every pipeline stage introduce a new failure mode? If that parameter gets deleted or the org root account is having issues, your entire CI/CD grinds to a halt. We've seen that with centralized configs before. It forces you to treat that parameter like a critical piece of infra, which is maybe the point.

Your `sts:GetCallerIdentity` pattern for assumed roles is gold. So many tag-based conditions break silently when you don't account for that final hop. That's going straight into our rulebook.



   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

The fintech DynamoDB rule is a perfect example of a concrete business constraint that the managed rules could never cover. Your point about the evaluation logic is crucial, though.

You're matching document structure, not simulating a request. That means a rule to enforce VPC-only access, as you noted, is brittle. You'd have to check for the existence of a `aws:SourceVpc` condition and the absence of a `Principal: "*"`, but you can't validate the VPC ID is correct. It's a structural lint, not an authorization test.

We've found this limitation pushes the complexity into the rule's JSON logic, making them harder to maintain than we initially expected. A "deny *" check is simple; a "allow only from these specific IPs" check becomes a convoluted pattern match.


Measure twice, spend once


   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Wait, so you can actually build rules that block specific AWS services for certain accounts? That's exactly what we need for our contractor workflows.

I'm new to setting up these guardrails though. When you say you're baking checks into your pipeline, are you running this validation before a policy gets created, or after? I keep seeing people mention a timing issue with Terraform.



   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Yes, you can block services for specific accounts. The rule's `forAnyResource: true` condition with a `StringNotLike` on the ARN pattern works for that. For example, blocking `s3:*` for a contractor account's DynamoDB resources.

On the timing issue, you're right to focus on that. The validation happens against a live policy in AWS, not a local file. So if you run it in the Terraform plan phase, the resource doesn't exist yet for the analyzer to check. Most teams run it post-apply as a pipeline step, then fail and roll back if a custom rule is violated. It's a trade-off between catching it early and having a valid target to analyze.


prove it with data


   
ReplyQuote
(@ethanw9)
Trusted Member
Joined: 2 months ago
Posts: 85
 

The post-apply check and rollback pattern makes sense for this. But doesn't that create a window where the bad policy exists live, even if briefly? How are you handling the potential for any automated process to pick up those new, non-compliant permissions before the rollback finishes?



   
ReplyQuote
(@billyp)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Exactly, that proactive angle is the real win! We run a lot of campaign data through S3 and our mail service SNS topics. Being able to stop a policy with `"Principal": "*"` before it goes live is huge for compliance. No more scrambling after a quarterly audit flags something.

We're definitely layering it. The managed rules catch the AWS no-brainers, then we add one or two custom rules for our marketing ops quirks - like preventing SQS policies in the production segment of our account. The gotcha? Start with those one or two high-impact rules. If you try to boil the ocean on day one, you'll spend all your time tuning JSON and fighting false positives instead of actually securing things.

For CI/CD, we run it post-apply in the pipeline, just like others mentioned. It fails the build if a custom rule trips. The window where a bad policy exists is brief, but we mitigate by having our automation scripts pull IAM from a separate, locked-down deployment account that doesn't use this pipeline. For campaign resource deployments, the risk is low enough that the trade-off is worth it.


Always A/B test.


   
ReplyQuote
Page 2 / 2