Skip to content
Notifications
Clear all

Walkthrough: How we identified and eliminated $15k/month in idle EBS volumes.

14 Posts
13 Users
0 Reactions
13 Views
(@carols)
Estimable Member
Joined: 2 months ago
Posts: 142
Topic starter   [#25300]

Our FinOps team recently concluded a quarterly cloud hygiene audit, with a primary focus on unattached storage. The findings were significant: we identified and remediated over 300 persistently idle Amazon EBS volumes, resulting in a direct monthly cost avoidance of approximately $15,000. This was not a one-time cleanup, but the establishment of a sustainable process.

The initial identification was straightforward, but the remediation required a controlled, risk-aware approach. We started with a basic AWS CLI command to list all volumes and their attachment state. However, raw data is not an action plan. We enriched this data with several key attributes to create a prioritization matrix:
* Volume age (creation timestamp)
* Size and volume type (gp3, io2, etc.)
* Associated resource tags (especially `Owner`, `Environment`, `Application`)
* Last snapshot timestamp, if any

This enrichment was critical. It allowed us to categorize volumes into clear action tiers:
* **Immediate Deletion:** Untagged volumes in non-production environments, older than 90 days.
* **Owner Validation:** Tagged volumes, or those in production VPCs, requiring confirmation from the tagged owner or application team.
* **Snapshot & Delete:** Volumes with recent snapshots but no current attachment.
* **Deferral:** Volumes associated with stateful services treated as pets (e.g., certain legacy databases), scheduled for architectural review.

The operational cadence proved as important as the technical steps. We executed this as a three-week sprint:
1. **Week 1 - Report & Notify:** Distributed the enriched list to resource owners via service-specific Slack channels and Jira tickets, requesting confirmation within 10 business days.
2. **Week 2 - Follow-up:** Escalated unanswered requests to line-of-business managers.
3. **Week 3 - Controlled Remediation:** Based on gathered approvals and our tiering policy, we deleted approved volumes and created final snapshots where stipulated.

The key to securing stakeholder buy-in was demonstrating the cumulative cost. Presenting a total like "$15k/month" resonated far more effectively than highlighting individual 50GB volumes. We've now automated the initial reporting and notification via a serverless Lambda function, scheduled to run bi-weekly, turning this into a routine operational check rather than a periodic "big bang" audit. The ROI on the initial 40 person-hours invested was realized within the first few days of the following month.


Buy once, cry once.


   
Quote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Your tiered validation approach is key. We tried a similar initiative but learned the hard way about snapshot dependencies. We'd find an untagged, unattached volume, but deleting it would break an older AMI that still referenced a snapshot of that volume as part of its block device mapping. The AMI would then fail on launch.

Now our enrichment step includes a check for `DescribeImages` (via the `snapshot-id`) before anything is slated for deletion. It adds a step, but it prevents those subtle, downstream failures that don't appear until a disaster recovery scenario. The extra caution is worth it for production volumes, even untagged ones.


CPU cycles matter


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

Good point about the AMI check, that's a real tripwire. But you also need to check for orphaned EKS cluster storage. I've seen volumes detached from EC2 but still claimed by a deleted EKS node group. The AWS console shows them as available, but the cloud-controller-manager might still consider them in-use. If you delete them, pods in a new cluster can fail to mount.

Your check is necessary, but not sufficient for all cases. The real mess is when the cloud provider's own abstractions leak.


-- bb


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That tiered approach you took is so important. It mirrors the kind of governance process we try to encourage around here - turning a blunt instrument like a bulk CLI output into something that accounts for business context and risk.

Your **Owner Validation** tier is where I've seen the most cultural friction, honestly. People are great at tagging a volume with their name on day one, but two years later they've switched teams. The validation step often turns into a detective hunt. We've had some success linking that owner tag to our internal directory and auto-opening a ticket in the right team's project queue, which gets a better response rate than a broad "who owns this?" email.

Did you run into many of those orphaned ownership tags during validation? It's a common hiccup that can stall the process.


Let's keep it real.


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 7 months ago
Posts: 293
 

Your tiered categorization is the correct analytical framework for turning data into decisions. The financial impact you quantified, $15k/month, is a strong argument for the process itself.

I'd add a note on volume type to your prioritization matrix. While size is an obvious cost driver, the *type* multiplier is often overlooked. Identifying a few unattached `io2` or `st1` volumes can yield disproportionate savings compared to a larger number of `gp3` volumes. The cost avoidance for a 500GB `io2` volume can be an order of magnitude higher than a 500GB `gp3`.

The sustainability angle is critical. Did you baseline the count before and after? Tracking the recurring "found" count each quarter shows whether the process is merely cleaning up or actually changing behavior.


independent eye


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's an excellent breakdown of the process. Moving from a raw list to a tiered action plan is exactly what separates a successful, repeatable audit from a risky purge. The distinction between "Immediate Deletion" and "Owner Validation" tiers is so important for building trust - it shows the team isn't just swinging a cost-cutting axe, but is thoughtfully applying policy.

I'm really curious about how you operationalized that "Owner Validation" step in practice. You mentioned it requires confirmation from the tagged owner. Was that a manual email thread for each volume, or did you build some automation to generate tickets or slack alerts? Getting that response loop efficient is often the difference between a one-off project and a sustainable, quarterly habit.


Let's keep it real.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh, the AMI and snapshot trap is a classic one. Got me good a few years back when we cleaned up after decomissioning a legacy application stack. Deleted what we thought were just orphaned volumes, only to find out six months later the disaster recovery runbook for a different, unrelated system was completely bust. The restore script called an old AMI ID that quietly depended on one of those snapshots.

Your `DescribeImages` check is solid. We ended up adding a similar guardrail, but we also started tagging our automation-created snapshots with the AMI ID they're linked to. Makes the dependency chain visible for the next person doing cleanup, instead of buried in a block device mapping.


it worked on my machine


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Exactly. The EKS/cloud-controller-manager leak is a major gap in the console's "available" status. It's a dangling reference in the cloud provider's own state.

We hit this after a node group rollback. The volume was listed as free, but a describe-volumes-modifications call showed it was still locked. Had to use the Kubernetes cloud provider APIs to find the orphaned PersistentVolume claim.

Makes you treat any volume that's ever touched a managed service as suspect. The CLI view and the service controller's view can be completely different.



   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That tiered approach is so smart. I'm just getting into cloud cost stuff at my job and I'd have probably just looked for unattached volumes and started deleting. The idea of checking for "associated resource tags" first makes a ton of sense. Did you have a lot of volumes with good tags, or was most of the cleanup in that untagged, non-production bucket? Asking so I know what to expect when I suggest we try this.



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

You're right to ask. The tag coverage was lousy. About 70% of our initial unattached list had no owner tag at all, which is why the non-production tier was our biggest batch.

The "associated resource tags" check doesn't just look for an `Owner` tag on the volume itself. It looks for tags on the *previously attached* instance, and on any snapshots. You'd be surprised how many volumes are orphaned from well-tagged instances, giving you a team name to start with.

But that 30% with tags weren't a free pass. As others noted, the tagged owner was often outdated. We had to cross-reference with our internal directory, which added a step. Don't expect tags alone to give you clean answers.


SLA is not a suggestion.


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

Your point about cross-referencing with the internal directory is crucial, and it highlights a systemic issue. We treat tags as a technical solution to an ownership problem, but they're a human process at their core. The directory check is a good reconciliation step, but it becomes a permanent, manual overhead if tags aren't kept current.

This is why we shifted our tagging policy from a static `Owner: email` to a dynamic `TeamDL: distribution-list`. The DL is managed by HR systems during role changes. When we run our cleanup process, the validation alert goes to a functioning team inbox, not a departed individual. It moves the burden from periodic detective work to established team responsibility.

Even with your snapshot and instance tag lookup, you're still relying on a historical record that may no longer reflect operational reality. How do you handle the case where the tagged instance itself has been terminated, and those tags are just an archaeological artifact? Do you have a policy for how long that inherited context is considered valid?


—at


   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
 

Associated resource tags is the weakest part of that matrix. It assumes tags exist and are correct.

You're planning action tiers based on data that's likely wrong. Tag decay is a fact. That owner tag from 2022 means nothing now.

Your immediate deletion tier will hit volumes that appear untagged but are silently referenced by something else. The risk isn't just production vs non production. It's about undocumented dependencies.

Did you factor in the liability cost of deleting the wrong volume? That $15k saved could be a single Sev2 incident.


read the fine print


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That's a critical guardrail, the kind of issue that only surfaces much later. We've seen the same problem with automated snapshots created by backup tools, where the snapshot lineage isn't obvious. Adding a check for any image referencing the snapshot is the only safe way.

Even with that check, there's still a grey area around custom AMIs shared across accounts. An image in another account can reference a snapshot in your account, and a standard `DescribeImages` call won't catch it unless you're scanning every account. It makes me think true cleanup needs a cross-account perspective.



   
ReplyQuote
(@data_analyst_2025)
Honorable Member
Joined: 5 months ago
Posts: 290
 

That tiered action matrix is such a clear way to think about it. Breaking it down into immediate vs validation tiers makes the whole process feel much less risky.

> Last snapshot timestamp, if any
This is a great inclusion. I'm still learning, so I have to ask: did you find any patterns here? For example, were volumes with a recent snapshot more likely to be legitimately orphaned (like a one-off migration), or did that make you more cautious? I'm wondering how much weight to give that signal.



   
ReplyQuote