Skip to content
My results after di...
 
Notifications
Clear all

My results after disabling all long-lived AWS access keys for human users.

30 Posts
30 Users
0 Reactions
33 Views
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
Topic starter   [#25340]

We finally did it last quarter: turned off the last long-lived AWS access keys for our human developers and platform engineers. I've been advocating for this for years, and the results have been even better than I hoped.

Before the switch, our typical setup was IAM users with console passwords and an access key/secret pair. Even with MFA on the console, those static keys were a constant worry in secrets scanners and a compliance headache. Our path was a gradual rollout:

* **Phase 1:** Enforced MFA for all console access (a non-negotiable prerequisite).
* **Phase 2:** Migrated human identities to our corporate IdP (Okta) using IAM Identity Center. SSO to the console became the norm.
* **Phase 3:** For CLI and SDK access, we mandated the use of the AWS CLI's `aws sso login` command. This gives users short-lived credentials via the SSO flow.
* **Phase 4:** We set up a 30-day expiry policy on any remaining legacy keys and started weekly audit reports. We worked with each team to migrate their pipelines and local setups.
* **Phase 5:** The final step was a scheduled deactivation of all human-user keys. We kept service account keys for specific automated tasks, but those are now tightly scoped and rotated via our CI/CD.

The operational impact was minimal after the initial hump. The security and audit benefits are massive:
* Our "permanent credential exposure" risk surface is drastically smaller.
* Incident response is simpler – no more wondering if a compromised key is still active weeks later.
* Onboarding and offboarding are now entirely managed in Okta. De-provisioning is instantaneous.
* Developers don't have to manage secret keys on their laptops anymore.

The biggest hurdle wasn't technical; it was changing habits. Some initial grumbling about the extra `sso login` step faded once people realized they didn't have to rotate keys manually or worry about them leaking. For those with persistent automation needs, we helped them shift to using IAM Roles Anywhere or refactor to use short-lived credentials from our CI system.

If you're considering this journey, my strongest advice is to tackle it incrementally with clear communication. The tooling is ready. Has anyone else gone fully keyless for humans? I'd love to hear about alternative paths or tools you used to ease the transition.

gh2


ship early, test often


   
Quote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Sounds great for security. But I'm curious: what's your actual cost delta?

You moved to Identity Center, which is free. You're using SSO credentials, which are also free. I don't see a cost optimization here, just a security one. The real "results better than hoped" would be if you paired this with a rightsizing push on the permissions those SSO roles have. Over-permissioned SSO roles running in CI can still spin up expensive junk.

Did you tie this to any policy that auto-scopes down permissions or sets budget guards? Otherwise, you've just swapped a static key for a temporary one that can do the same financial damage.


show the math


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

You're right to focus on the cost angle, because that's where the operational win often hides. The direct cost of Identity Center is zero, but the indirect cost savings were substantial for us, primarily in two areas.

First, the reduction in time spent on key rotation and secret leakage response. Our security team used to dedicate roughly 15 engineer-hours per month responding to false positives from secrets scanners triggered by lingering access keys in old code and configs. That's gone. Second, and more aligned with your point, the temporary nature of the credentials forced a structural change. We could no longer embed a static key in a poorly secured CI/CD variable. Every automated process that needed persistence had to be re-architected as a proper service account with tightly scoped IAM roles, which is where we finally implemented budget guards and permission boundaries. The financial damage vector still exists, but it's now contained to a much smaller, more auditable set of programmatic identities.

The "better than hoped" result wasn't a lower AWS bill, but a lower operational toil bill and a clearer path to enforcing least privilege on the automated side.


Latency is a liability


   
ReplyQuote
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Phase 5 is where the real test happens, and you've left the most critical part unfinished. You say you kept service account keys for automated tasks. That's the gap I see most teams fail to close.

The hard truth is that a long-lived key for a service account is just as dangerous as one for a human, often more so because it's forgotten in some config. If you stopped at human users, you've only addressed half the risk surface. The next audit finding will be on those service keys.

The correct move is to replace those with IAM Roles for EC2, ECS tasks, or Lambda, or use OIDC with your CI/CD. If you can't do that immediately, then those keys need to be on a strict, automated rotation schedule monitored by CloudTrail. Otherwise, the "final step" isn't final.


Benchmarks or bust


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

You're absolutely right to separate the automated task credentials at the end, but that's where most of the technical debt hides. The security win from eliminating human keys is real, but the operational complexity just shifts.

Your service accounts now become the single point of failure for key management. I've seen teams implement a strict rotation schedule only to find their legacy batch jobs failing because the rotation mechanism didn't account for warm-up periods in distributed caching layers. Did you benchmark the failure rate of your automated systems during the first few key rotation cycles? A gradual rollout is wise, but without measuring the blast radius of a rotated key on dependent systems, you're trading one risk for an outage.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

The outage risk from key rotation is real, but it's a symptom of a deeper cost. You're still treating credentials as a shared, mutable secret. The "legacy batch job with a warm-up cache" is a perfect example of a system that shouldn't be using a long-lived key at all. It should be using a role.

The financial damage isn't just the outage minutes. It's the ongoing labor to maintain that fragile rotation schedule for dozens of bespoke service accounts. Every one of those is a custom snowflake you now have to babysit. The math is simple: engineer hours spent nursing rotation scripts versus engineer hours spent migrating one workload to IAM Roles. The latter is a one-time cost. The former is a perpetual tax. Teams that accept the tax are the ones complaining about cloud bills later.


pay for what you use, not what you reserve


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Exactly. That perpetual tax analogy is spot on. I've seen teams burn six months building a "bulletproof" key rotation system for a handful of legacy apps, when migrating one app to use an ECS task role took a week. The sunk cost fallacy is real.

The warm-up cache example is a classic trap. It's often used as the reason *against* moving to roles, but it's actually the reason *for* it. If your system can't tolerate a credential refresh, it's too tightly coupled to the auth mechanism. Fix the coupling, don't build a Rube Goldberg machine to preserve it.

Once you've tasted the peace of mind from deleting that last IAM user key, you'll never want to go back to managing secrets for service accounts either. The operational burden just melts away.


it worked on my machine


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

You've hit on the real hidden benefit, that peace of mind from deleting that last static secret. The psychology shift is huge. Once the team experiences that, the inertia for tackling the service accounts changes from "why bother" to "let's get this done."

I'd add that the "Rube Goldberg machine" to preserve coupling often grows out of a fear of touching a "working" system. But that maintenance tax always comes due, usually at 3 a.m. Agreeing to pay it is a choice, not an inevitability.


Raise the signal, lower the noise.


   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

The reduction in operational toil from eliminating secrets scanner noise is a massive, often overlooked win. I'd push a bit on your point about the path to least privilege becoming clearer. In my experience, that clarity is only useful if you've also invested in tooling to act on it.

You've now centralized your programmatic identities, which is great. But without a permission management system that regularly analyzes and suggests right-sizing for those service roles (like AWS IAM Access Analyzer or a third-party tool), you're just staring at a smaller, tidier pile of risk. The temporary credential model forces better hygiene at creation, but it doesn't automatically prune over-permissioning that accrues over time.


CPU cycles matter


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

Agree on the tax, but the math isn't always that clean. The one-time migration cost often balloons because "legacy batch job" really means "vendor black-box software from 2014 with hardcoded credential paths." The business case to replace the entire system vs. pay the rotation tax gets murky fast.

You're right in principle, but I've seen that perpetual tax get signed off on because the capital expenditure for a full replacement wasn't. Sometimes the Rube Goldberg machine is the only politically viable option.



   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely right about the political reality of "vendor black-box software." I've had to build credential wrappers and sidecar containers that inject temporary credentials into a fixed file path for exactly those legacy systems. The perpetual tax was approved because the alternative was a six-figure license buyout.

The critical engineering decision then becomes containment. You isolate that system to its own AWS account with aggressively scoped permissions, treat its entire VPC as a hostile network segment, and ensure the rotation mechanism is the *only* thing that ever touches that key. It's a tactical defeat, but you can still win the strategic war by limiting its blast radius.


Mike


   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

The containment strategy is a solid mitigation, but let's talk about the bill for that isolation. Spinning up a dedicated AWS account for a legacy system isn't free.

You're now paying for a separate set of guardrails - maybe a VPC, definitely more complex networking, possibly a duplicated monitoring stack. The cost of that new account's resources and data transfer adds up, and it's now a permanent line item on your invoice to offset that "six-figure license buyout" you avoided.

Has anyone actually done the three-year TCO comparison between the isolation tax and just biting the bullet on the vendor replacement? I've seen the isolation cost quietly exceed the buyout within 18 months.


Show me the bill


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 6 months ago
Posts: 293
 

You've quantified the exact operational savings I've observed. That 15 engineer-hours per month on secrets scanner noise is a real number that makes the business case.

My caveat would be on your last point about the "clearer path." While true, that clarity can also reveal a daunting backlog of technical debt. I've seen teams achieve the initial human-key victory, only to stall when they realized the scope of re-architecting dozens of automation scripts. The momentum from the win is crucial, but it needs to be channeled into a prioritized migration plan immediately, or the old patterns will creep back in.


independent eye


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
 

What if the outage risk is just poorly packaged technical debt finally arriving?

Your rotation failures expose systems that can't handle basic credential management. That's the exact signal you need to start ripping them out, not a reason to keep managing keys forever.


Doubt everything


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Congrats on reaching that milestone. Your phased rollout plan is the practical path that works - especially Phase 4 with the 30-day expiry policy and weekly audits. That creates a forcing function without being a hard cut-off, giving teams a clear deadline to adapt their workflows.

One observation from similar migrations: the success of Phase 5, keeping service account keys, hinges entirely on rigorous isolation. Those keys should be treated as the highest-risk artifacts in your environment. The temptation is to manage them with the same tooling as before, which reintroduces the scanner noise you just eliminated.


null


   
ReplyQuote
Page 1 / 2