Having recently completed the implementation of a quarterly access review process for our cloud infrastructure, I found myself dissatisfied with the manual, spreadsheet-heavy approach. The process was not only tedious but also prone to human error, particularly when scaling across multiple projects and service accounts. To systematize and de-risk this, I've developed an automated audit script that enumerates and evaluates user permissions across our core GCP and AWS environments, executing on a monthly cadence.
The script's primary function is to move beyond simple role listing and perform a contextual analysis. It cross-references IAM policies, Identity Center (AWS SSO) assignments, and resource-level permissions against a pre-defined, organization-specific "least privilege baseline" defined in a YAML configuration. The output is a detailed report highlighting deviations, such as standing permissions that violate the principle of least privilege, unused service accounts older than 90 days, and direct user assignments to production resources that should be mediated through groups.
```python
# Example core function for GCP IAM analysis
def analyze_gcp_iam(project_id, baseline_roles):
"""Fetches IAM policy and flags non-baseline bindings."""
from google.cloud import resourcemanager_v3
client = resourcemanager_v3.ProjectsClient()
policy = client.get_iam_policy(request={"resource": f"projects/{project_id}"})
findings = []
for binding in policy.bindings:
if binding.role not in baseline_roles:
for member in binding.members:
findings.append({
"resource": project_id,
"member": member,
"excessive_role": binding.role,
"recommendation": f"Evaluate for reduction to {baseline_roles.get(binding.role, 'viewer')}"
})
return findings
```
The implementation leverages a combination of cloud provider SDKs (primarily for enumeration) and a custom rules engine. Key components include:
* **Enumeration Module:** Programmatically lists users, service accounts, groups, and their associated roles/policies across defined projects and accounts.
* **Rules Engine:** Applies logic from the baseline YAML, checking for patterns like admin roles on non-admin users, broad storage permissions, or cross-project access.
* **Reporting Module:** Generates a markdown report for human review and a JSON-line file for ingestion into our SIEM (Security Information and Event Management) for trend analysis.
* **Orchestration:** The entire pipeline is containerized and scheduled via Cloud Scheduler, invoking a Cloud Run job. This design eliminates the need for persistent compute and simplifies dependency management.
Initial results were, predictably, alarming. The first run identified 17 instances of overly permissive roles attached to service accounts meant for specific, narrow functions, and 3 user accounts with direct project-owner roles that had not been accessed in over six months. The script has now been integrated into our compliance workflow, with findings creating tickets in our incident management system automatically.
I'm curious if others in the community have tackled similar problems, particularly around:
* Managing false positives for legitimate but broad permissions (e.g., break-glass accounts).
* Incorporating SaaS application permissions (like GitHub teams or Salesforce profiles) into a unified view.
* The trade-offs between building such a tool in-house versus using a commercial Cloud Security Posture Management (CSPM) platform, especially concerning long-term maintenance cost.
- alex
Measure twice, cut once.
That move from spreadsheet to script sounds like it was desperately needed. I'm curious how you got buy in to shift from a quarterly to a monthly cadence. Was that an easy sell on the basis of risk reduction, or did you have to show specific, near-miss examples to make the case for the increased frequency?
Great question about the cadence shift. In my experience, selling a monthly review came down to reframing it not as extra work, but as less risk per unit of work. Quarterly reviews often became huge, overwhelming projects - people dreaded them. A monthly automated script is actually lighter touch; you're catching drifts when the context is still fresh, and the fixes are smaller. The "near-miss" examples helped, but the bigger sell was showing that monthly, automated checks would take *less* total human time than the quarterly manual fire drill. It turned the conversation from "more audits" to "better, cheaper audits."
Raise the signal, lower the noise.
Exactly, that reframing from "more work" to "less drama" is the only way these operational changes ever get traction. The quarterly monster review is a perfect example of a process that's so painful it incentivizes shortcuts and blind spots, which defeats the whole purpose.
My cynical addition is that you often have to prove the automation is actually cheaper. I've seen teams spend more engineer-months building and maintaining the "automated audit" than the manual process ever cost. The script becomes a pet project, then a legacy system, then a migration headache.
So the real sell isn't just monthly vs quarterly, it's that your script is a one-hit wonder with minimal ongoing fiddling. If you need a dedicated FTE to babysit it, you've just built a more expensive, faster-spinning hamster wheel.
keep it simple
You've nailed the cynicism perfectly. The pet project to legacy system pipeline is a well documented failure mode in ops engineering.
I'd add that the real cost often hides in the integration surfaces, not the core logic. If your script relies on a specific library version for the AWS SDK, or a particular GCP service API that gets deprecated, the maintenance burden spikes silently. You can build a beautifully clean script for today's APIs, but you're signing up for a subscription to their future change logs.
The "one hit wonder" ideal requires a design that minimizes dependencies on volatile interfaces. Sometimes that means accepting a slightly less detailed audit in favor of using stable, long term CLI tools instead of the latest SDK. The tradeoff is between completeness and longevity.
Measure twice, cut once.