Skip to content
Notifications
Clear all

Just built a script that tags untagged resources. Saved us $3k/month.

1 Posts
1 Users
0 Reactions
20 Views
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
Topic starter   [#6142]

Hey folks, I've been deep in the weeds of our multi-cloud cost optimization lately, and I just had to share a win that felt too good not to post. As you know, we run workloads across AWS, Azure, and GCP, and untagged resources have been a silent budget killer for us. We finally bit the bullet and built a script to automatically find and tag them, and the initial results are saving us a staggering **$3k/month** on average. It turns out a huge chunk of our "mystery spend" was just untagged development and test resources that nobody wanted to claim ownership of.

The core idea is simple: iterate through all major services (compute, storage, databases) in each provider, check for missing mandatory tags (like `Owner`, `Environment`, `Project`), and apply tags based on some fallback logic. The real magic, though, is in the error handling and the multi-provider abstraction. We used the respective SDKs (Boto3, Azure CLI, Google Cloud Python Client) and wrapped them in a consistent function.

Here's a simplified snippet of our core tagging logic for AWS EC2, which was our biggest offender:

```python
def tag_aws_instances(session):
ec2 = session.client('ec2')
instances = ec2.describe_instances(Filters=[{'Name': 'tag:Owner', 'Values': ['']}])
untagged = []

for reservation in instances['Reservations']:
for instance in reservation['Instances']:
if not any(tag['Key'] == 'Owner' for tag in instance.get('Tags', [])):
untagged.append(instance['InstanceId'])

if untagged:
# Fallback logic: derive owner from IAM user or last started-by event
derived_owner = derive_owner_from_cloudtrail(instance['InstanceId'])
ec2.create_tags(
Resources=untagged,
Tags=[
{'Key': 'Owner', 'Value': derived_owner},
{'Key': 'Environment', 'Value': 'unclassified'},
{'Key': 'CleanupDate', 'Value': (datetime.now() + timedelta(days=14)).isoformat()}
]
)
print(f"Tagged {len(untagged)} instances.")
```

We built similar modules for Azure VMs and GCP Compute Engine, and then for object storage (S3, Blob, Cloud Storage) and managed databases. The key gotchas we had to navigate:

* **Rate Limiting:** Each cloud provider has different throttling limits. We had to implement exponential backoff with jitter, especially in Azure.
* **Permission Scoping:** The service principal/role needed *list*, *read*, and *tagging* permissions across a dizzying array of services. Took a few iterations to get the IAM policies just right without being overly permissive.
* **Fallback Logic:** Not all resources have clean metadata. For some, we had to cross-reference CloudTrail/Azure Activity Log/Google Audit Logs to infer an "owner," which added complexity.
* **Idempotency:** The script runs daily via a scheduled Lambda/Cloud Function, so it had to be safe to re-run without duplicating tags or causing errors.

We're now running this as a serverless function in each cloud, and the reports it generates have become indispensable. It's not just about cost savings; it's about accountability and cleaning up old resources. The `CleanupDate` tag has triggered our existing automation to email owners and then shut things down, which compounded the savings.

Has anyone else tackled cross-cloud untagged resources? I'm particularly curious about your strategies for handling container services (like EKS, AKS, GKE) or serverless resources (Lambda, Functions, CloudRun), as our script's coverage there is still a bit spotty. I'm happy to share more code snippets if there's interest!

-- Ian


Integration Ian


   
Quote