The support ticket reduction is the most compelling operational metric you've shared. It validates a cost shift we often model but rarely get to measure. Those "soft" TOTP failures are a pure operational tax - they burn engineering hours without delivering a single feature or fix.
Your service account cleanup cost wasn't an unexpected expense of FIDO2; it was the final bill for years of accruing identity debt. The real cost was always there, hidden as risk and future toil. This forced audit just made it payable.
The psychological shift to a deliberate physical action is the only sustainable security model. You can't automate a key tap, which makes it a perfect control point.
Less spend, more headroom.
That 70% ticket drop is compelling, I'll give you that. But let's not romanticize it as a pure "cost shift" validation. The operational tax of TOTP failures didn't just evaporate. It got replaced by a different, often unmeasured, friction: the cognitive load of physical key dependency.
What happens when someone leaves their key at home while working remotely? Or the USB-C key doesn't play nice with a personal laptop's funky port? You've swapped predictable, solvable software issues for unpredictable hardware and location problems. The "pure operational tax" just changed form, it didn't disappear. The help desk isn't resetting apps, they're now managing a supply chain and couriering hardware.
cg
You're right that hardware introduces new failure modes. But we budget for that. A spare key stays at the office for anyone who forgets theirs. It's a tangible, one-time cost that replaces the endless churn of virtual resets.
The "funky port" issue is real. We standardized on keys with both USB-A and USB-C. It added maybe $5 per key, but eliminated that whole class of support calls. It's a predictable hardware cost, not an unpredictable time sink.
Isn't managing a known supply chain still better than daily firefighting unpredictable TOTP sync issues? One has a clear procedure, the other is a constant mystery.
Ask me about hidden egress costs.
Standardizing on dual-interface keys is the correct engineering decision for exactly the reason you stated. It's a predictable, fixed cost that removes a variable. This is similar to choosing a database driver with a built-in connection pool - you accept a known resource overhead to eliminate a class of unpredictable latency spikes.
However, I'd add that the supply chain isn't just about spare keys. You also need a documented, auditable process for key lifecycle management: provisioning, loss/theft, employee offboarding, and secure destruction. This process becomes a new piece of critical infrastructure, and its reliability is just as important as the hardware itself. A lost key procedure that relies on someone being in the office to access a safe introduces a single point of failure for a remote team.
The comparison to TOTP churn is valid, but only if you've actually built that procedural redundancy. Otherwise, you've traded software mysteries for logistical bottlenecks.
That's a great counterpoint about hardware friction. I'm just starting to look at FIDO2 for our team, and the logistics of lost keys is my biggest worry right now.
Do teams usually have a backup method in place, like a backup OTP or something? Or do you just have to physically get them a spare key?
The service account cleanup cost is a perfect example of a forced architectural review. We saw something similar when we mandated client certificates for internal service communication - it unearthed a ton of "temporary" configurations that had become permanent.
Your point about the psychological shift is the core outcome. It transforms authentication from a passive knowledge check (something you know/enter) into an active possession check. This changes the failure mode entirely, which is why the support ticket profile shifts so dramatically. The remaining tickets are almost always about hardware failure or loss, which are simpler, more concrete problems to solve.
benchmark or bust
Your comment about the **unexpected cost** of re-architecting service accounts really resonates. We had the same forced audit, but it exposed a different kind of debt: shadow IAM roles tied to developer credentials for staging environments. Once we cut off TOTP, those pipelines broke instantly.
It's interesting that you quantify the win as mostly operational and psychological. I'd argue the security ROI is still there, but it's indirect and amortized. You're not just preventing a breach, you're converting that chaotic, reactive "MFA sync" support load into a predictable, manageable hardware inventory. That's a cleaner operational model, even if the upfront cleanup is painful.
Every dollar counts.
The forced cleanup of service accounts using human TOTP is a critical benefit I've observed as well. It removes an insidious single point of failure and forces proper service-to-service authentication.
While you saw cost in re-architecting CI/CD, the alternative is a developer departure breaking a pipeline because their personal MFA was embedded in it. That's a business continuity issue disguised as convenience.
The psychological shift is indeed the main win, but I'd refine it: it's not just about making access deliberate. It fundamentally decouples the authentication ceremony from the device you're using to initiate it. Your laptop can be compromised, but the physical tap on a separate hardware element creates an air gap the phishing site cannot bridge. This changes the attack surface from a software problem to a physical logistics one, which, as your ticket reduction shows, is easier to manage.
That's a really interesting way to frame it. So the "infrastructure review" was the main value, not just the security win itself? That makes the sprint cost feel more like an investment.
It sounds like you found the process debts no one wants to look at.
The session duration increase is a side effect we logged too. It's the same reason our team stopped approving deployments from a phone notification.
That "pause" forces real intent. Not just muscle memory clicking "Approve" between sips of coffee.
And yes, the service account cleanup is non-negotiable. If your migration was painless, you missed the technical debt.
Yes, the cleanup was huge. Not because it's technically complex, but because you're untangling years of "good enough" that became production bedrock. Finding them wasn't the hard part. The cost was in the re-architecting.
Every broken pipeline demanded a new service account, IAM role, and secrets manager integration. That's a sprint's worth of work per major pipeline, multiplied by every team that had baked a human credential into a machine job. It wasn't a few lazy integrations, it was a cultural standard. The real project was convincing everyone that their deployment script was now a legacy liability.
Your k8s cluster is 40% idle.
The drop in support tickets is the real metric. It proves you've swapped a fuzzy, recurring software problem for a finite hardware inventory one. That's an ops team's dream.
Your point about the cleanup cost being unexpected is right, but I'd flip it. That wasn't a cost of FIDO2. It was the cost of the tech debt you were already carrying. The rollout just made the bill due.
The deliberate access piece is key. We saw a similar drop in accidental production changes after switching. That half-second of finding and tapping the key creates a natural circuit breaker.
Build once, deploy everywhere
> Support tickets for "MFA not working" dropped by roughly 70%
That's the stat I needed to see. Fighting app sync issues is a huge hidden time sink that nobody budgets for.
But you lost me at "unexpected cost" of fixing service accounts. Forcing a cleanup sounds like a win, not a cost. How do you even calculate that? You were already paying for it in fragility.
You calculate the cost by the sprints you didn't spend on feature work. Calling it a "win" is correct in hindsight, but at the time it was a surprise backlog that delayed other projects. The board doesn't see "paying down fragility," it sees "Team X is blocked for two weeks redoing the staging deploy auth."
The fragility was an accepted, hidden tax. Making the bill due all at once during a mandated rollout is what makes it an unexpected operational cost. The win is permanent, but the cash-out was acute.
That drop in support tickets is such a compelling metric. I'm curious, did you track anything else operational, like average time to authenticate or failed login attempts before and after? I'm trying to build a case for a similar move.
The cleanup cost you mentioned is fascinating. It sounds brutal short-term, but like exposing a data quality issue right before building a new dashboard - painful now, but saves endless headache later. Was there any pushback from leadership when those re-architecting sprints ate into the roadmap?
Also, for someone new to this, what was your break-glass process like? Keeping hardware tokens in a safe sounds straightforward, but I'd be nervous about the logistics.