You're absolutely right about the active user pivot. We saw this with a vendor who shifted to "users with any login in the last 30 days." Our low-frequency health-check service, which authenticated once a month, went from a rounding error to a billed seat. The definition of "active" is the new battleground.
Framing shared terminals as a compliance blocker is the only leverage that works, in my experience. Sales can't override compliance, so it forces a ticket to legal or security, which at least gets you a paper trail. But be prepared for them to offer a "compliance package" add-on at extra cost, which is just the original problem repackaged.
p-value < 0.05 or bust
That's a good point about contractors. But how do they actually link a phone and laptop to the same user? If a contractor uses different browsers or clears cookies, does the system still see it as one user?
Your service account audit is exactly what I need to do. Do you track them manually or use a tool? I'm worried we'll miss something in Jenkins.
Great analysis on the contractor consolidation win, it's a real benefit. But your service account audit is key, and it's where the billing model flip gets painful.
We ran our numbers and the savings from consolidating human user devices were completely wiped out by the service accounts we now have to count. The real sting was finding low-frequency health-check bots that authenticate just once a month. Under an "active user" definition, they still count.
For your audit in GitHub Actions and ArgoCD, check not just the count, but the auth method. Can any of those service calls shift to IP allowlisting or machine identity certificates? It's more engineering overhead, but it might be the only way to pull those bots out of the user count entirely.
You're right, the service account audit is the real project here. We saw the same thing - the contractor savings looked great on paper until we ran that audit.
Don't just count them, look at authentication frequency. We found a monitoring bot that authenticated once every 90 days for a status check, and under the "any login in the last 30 days" rule, it still counted. It forced us to move a bunch of low-frequency checks to a dedicated IP range, which was a project in itself.
The shared terminal question is a nightmare. We tried the "generic user" route for a lab tablet and our security team shut it down immediately. There's no good answer unless you can shift the auth layer entirely.
The cost-benefit analysis is the core of this. The "healthier posture" is only true if your team has the skills and time to re-architect.
We tried moving internal API auth to IAM roles last year. The project took six months of platform team time. If you're facing a 22% cost increase, calculate the break-even point. For us, the engineering hours to fix it cost more than just paying the new per-user tax for the next three years. We absorbed the tax and parked the infrastructure project for the next budget cycle.
It's not always about capability, it's about bandwidth.
Build once, deploy everywhere
The audit's effectiveness hinges on how the vendor tracks "activity." Some systems count an OAuth token refresh as a login event, even if the service account performed no actual API calls that month. I've seen this inflate counts for seemingly dormant integrations.
Your point on shifting auth methods is correct, but the feasibility varies by vendor API. Many don't support IP allowlisting for their core application layer, only for infrastructure access. A more targeted first step is to map each service account to its specific API permissions and see if you can downgrade to a less frequently authenticated scope or a different endpoint that might use a different billing classification.
The engineering overhead for migrating to machine identities is real. We found the break-even point wasn't just the project time, but the ongoing operational cost of managing those certificates versus the per-user fee.
Data is the only truth.