We've all been there—you get that monthly bill, spot a weird spike, and then spend hours digging through console logs trying to reconstruct what happened. Manually collecting evidence for cost anomaly reviews is a pain.
I built a script to automate evidence collection for AWS cost anomalies. It pulls data from Cost Explorer, correlates it with CloudTrail events and resource configurations, then packages everything into a timestamped report. Saves our team about 10 hours a month in manual gathering.
Here's the core of it. It uses the AWS SDK for Python (Boto3) and outputs a structured JSON file you can feed into your investigation tickets.
```python
import boto3
import json
from datetime import datetime, timedelta
def collect_anomaly_evidence(anomaly_start_date, anomaly_end_date, linked_account_id=None):
"""Collects cost, usage, and event data for a given anomaly period."""
evidence_package = {
'collection_timestamp': datetime.utcnow().isoformat(),
'anomaly_period': {'start': anomaly_start_date, 'end': anomaly_end_date},
'cost_data': [],
'relevant_events': [],
'resource_inventory': []
}
ce_client = boto3.client('ce')
ct_client = boto3.client('cloudtrail')
# 1. Get granular cost and usage data
response = ce_client.get_cost_and_usage(
TimePeriod={'Start': anomaly_start_date, 'End': anomaly_end_date},
Granularity='DAILY',
Metrics=['UnblendedCost', 'UsageQuantity'],
GroupBy=[{'Type': 'DIMENSION', 'Key': 'SERVICE'}]
)
evidence_package['cost_data'] = response['ResultsByTime']
# 2. Get CloudTrail events for the period, filtered for high-cost services
# (Filtering logic would go here, simplified for example)
return evidence_package
# Example usage
if __name__ == "__main__":
evidence = collect_anomaly_evidence('2024-10-01', '2024-10-07')
with open(f"anomaly_evidence_{datetime.utcnow().strftime('%Y%m%d')}.json", 'w') as f:
json.dump(evidence, f, indent=2)
```
Key things it captures:
* Daily cost and usage grouped by AWS service
* Associated CloudTrail events (filtered for create, modify, delete actions)
* Snapshot of relevant resource configs (e.g., EC2 instance types, RDS storage)
* All output is tagged with the anomaly period for clear audit trails
I'm curious—how are others handling this? Do you have a similar process for Azure or GCP? I'm thinking of extending it to pull Azure Cost Management data next. The main pitfalls I've found are handling pagination on large datasets and managing IAM permissions scoped tightly enough.
That's really helpful, thanks for sharing. I'm new to using the AWS SDKs this way. How do you handle pagination when pulling from CloudTrail? I always worry about missing events if the response gets truncated.
Still learning.
Pagination is critical, you're right to be concerned about truncation. The boto3 paginators handle this for you if you use them correctly.
For CloudTrail specifically, you should be using the dedicated paginator for lookup_events. Don't try to manually manage NextToken yourself. It's a common newbie mistake that leads to missed data. The paginator object abstracts the loop.
One caveat: even with paginators, watch your time range. If you query a very broad window, you could hit the 50,000 event limit on a single call. For cost anomaly reviews, you should already be narrowing your window to the anomaly period, so this is usually fine.
Your cloud bill is 30% too high
Great approach, the structured JSON output is perfect for feeding directly into our ticketing system's evidence field. I've been using a similar script, but I've found that tagging the output with the internal ticket ID from our procurement system is crucial for audit trails.
One caveat: make sure your script's IAM role has the `ce:GetCostAndUsage`, `cloudtrail:LookupEvents`, and `tag:GetResources` permissions scoped appropriately. Otherwise, it'll fail silently when pulling resource configs, which is the most common pitfall in our setup.
Are you also filtering those CloudTrail events to only show actions by IAM principals outside your approved service accounts? That correlation is usually what pinpoints the root cause for us.
buyer beware, but buy smart
The IAM permissions catch is a good one - it's always the silent fails that get you. On filtering, I usually add a filter for `ReadOnly != true` and then cross-reference with a list of known CI/CD/service accounts. It's shocking how often the culprit is a dev's personal IAM user running a one-off script with overly permissive credentials.
Tagging with the procurement ticket ID is clever. I've been dumping everything into a central S3 bucket with anomaly date prefixes, but that extra metadata layer would save the back-and-forth when finance asks for a lineage report six months later.
Data over dogma.
Totally agree on the `ReadOnly != true` filter. I'd add that you should also filter out AWS service principals like `cloudformation.amazonaws.com`. Their automated actions can clutter the log but are rarely the cost culprit. Makes the human-run events pop.
Smart move with the S3 bucket structure, but yeah, that extra procurement tag is a lifesaver for traceability. We tag ours with the ticket ID *and* the anomaly ID from AWS Cost Explorer. Saves so much time when you're pulling old reports.
Solid foundational script. The structured JSON output aligns perfectly with automated ingestion workflows. However, the `cost_data` list appears to be populated later in your truncated example, and I'd caution that pulling raw `GetCostAndUsage` results without dimensional granularity can be a dead end. You'll need to specify at least `SERVICE` and `LINKED_ACCOUNT` dimensions in your request to get actionable data. The default `MONTHLY` granularity is also useless for pinpointing anomalies; you must use `DAILY` or `HOURLY`.
A more critical gap is the lack of any resource inventory population logic. Without correlating the cost spike to specific resource ARNs via the `tag:GetResources` call or the Resource Groups Tagging API, you're left with a service-level cost increase and a sea of CloudTrail events, with no direct link between them. The real value is in stitching those three data sources together: cost (which service spiked?), configuration (which RDS instance or Lambda function increased its utilization?), and events (what user or system action changed its state?). Your script's skeleton is good, but the omission of that join logic is where most of the investigative work actually happens.
"Without correlating the cost spike to specific resource ARNs" is exactly where the script falls apart. Everyone loves the skeleton, but the join logic is the whole investigation. Yet I see this pattern constantly - scripts that automate the easy 80% and call it a solution, while the crucial 20% is still manual, error-prone detective work. So you save 10 hours on gathering, only to spend 5 more trying to mentally connect service-level costs to actual resources. The real automation promise is a mirage.
—EB
I disagree with the "automation promise is a mirage" point, but the criticism about the join logic is valid. That's the hard part.
I've been using a similar framework, and the successful scripts I audit always include a defined join key. They don't just dump data. For example, they map CloudTrail events by resource ARN to cost line items via the Resource Groups Tagging API, using the tags applied during provisioning. It's the correlation layer that makes the evidence actionable for root cause.
Without that join, you're right. You've just automated the data dump, not the investigation.
Where is your SOC 2?