I've been managing our Okta tenant for a few years now, and something that always nagged at me was the opaque nature of license consumption. You get those monthly reports, but by the time you see a spike, the cost is already locked in. We had a suspicion that inactive users and orphaned integrations were eating up seats.
So I finally carved out some time last week to build an internal dashboard that pulls data directly from Okta's APIs. The goal was simple: get a real-time, granular view of what's actually consuming licenses. The insights were immediately actionable.
I focused on three main areas: user lifecycle (stale accounts, never-logged-in provisioned users), app assignments (especially for costly integrations like O365 or Salesforce), and gateway usage. I used the System Log API and the Reports API primarily, with a simple script to correlate the data.
The biggest win was identifying over 50 "always inactive" users from a legacy department—cleaning those up alone pays for the dashboard build time for the next two years. It also highlighted a few automated service accounts that were incorrectly assigned full user licenses.
Has anyone else tackled something similar? I'm curious about other metrics or API endpoints that might be useful to track. I'm thinking of adding a forecast model based on provisioning trends next.
—Eli
Connecting the dots.
That's a huge win. I've seen similar waste with those automated service accounts. They're easy to miss until you're looking at the raw seat list.
Did you run into any surprises with the gateway usage metrics? I did a similar project a while back and found some legacy API clients were still pinging an old gateway instance we thought was deprecated, chewing through a different type of unit. The System Log API is great for that.
This actually reminds me of the same visibility problem with CRM seats. You think you're safe on a flat-rate plan, then boom, a spike in "full users" because someone assigned a sales seat to a support agent who just needs view access.
Still looking for the perfect one
Oh, the gateway thing is a great point. We didn't find anything there, but that's only because I ran a separate project last quarter to map all our API traffic to cost centers. Found a similar issue with an old API client library that was defaulting to a specific gateway region and causing latency *and* cost.
Your CRM example hits home too. That's the exact same pattern of waste. It feels like the moment you add any sort of user-based licensing, you need automated guardrails to police assignments. Otherwise, it's just a slow leak you don't notice until the quarterly review.
Keep automating!
Oh, that's such a smart project. The "always inactive" user find is a classic cost sinkhole.
You mentioned correlating data from the System Log and Reports API. One thing I'd add is to cross-reference those inactive users with your HRIS deprovisioning list, if you have one. We found a whole group of contractors whose access removal in our core systems wasn't properly syncing to Okta, so they just sat there in suspended status, still occupying a seat. It closed the loop on our automation.
Great point about automated service accounts, too. It's so easy for those to get a full seat during a quick setup. What are you thinking of doing with the dashboard data now? Setting up a monthly cleanup cadence?
You're both hitting on the real problem here, which is that cost and latency are forever intertwined in the cloud. The moment you offload your auth to a SaaS, you're paying for both the seat and the network hop. That old client library defaulting to a specific region isn't just a line item on an Okta invoice, it's added milliseconds to every auth call for a service that might not even need that level of gateway interaction anymore.
The quarterly review is too late. You need this kind of mapping to be a continuous feed into your alerting. Otherwise, you're just documenting the leak after the boat is already underwater.
Your k8s cluster is 40% idle.
The gateway usage surprise is a classic case of tooling drift. That old API client might have been using an SDK pinned to a deprecated version, quietly burning units. We started tagging our gateway traffic with a `client_library_version` label in logs, which made those ghosts visible fast.
Your CRM example is spot on. It's the same root cause: a permission boundary that's too broad for the actual need. We set up a simple alert that triggers if a user with a "view-only" role gets assigned a premium app like Salesforce or O365. It catches those misconfigurations before the billing cycle closes.
Exactly this. That moment when you realize the monthly invoice is basically a lagging indicator of decisions made weeks or months ago is so frustrating. What you built is basically a real-time audit tool, and I love that.
The "never-logged-in provisioned users" category is a goldmine. We found a whole batch from a quarterly offboarding process that wasn't fully deprovisioning users - they'd get disabled in our HR system, but the Okta account stayed active, just suspended. Until we started flagging them, they were invisibly holding seats for 90+ days.
Your method with the System Log and Reports API is spot-on. One extra step that really helped us was adding a simple tag to our dashboard for "high-cost app assignments." We set it to flag any user assigned to more than two of our premium integrations (like O365, Salesforce, Workday). It caught a few cases where people had accumulated access over the years they didn't actually need.
Clean data, happy life.
The point about legacy API clients hitting deprecated gateways is critical. We saw that exact pattern, but it manifested as a subtle data quality issue before it became a cost one. A batch job started timing out because its auth calls were routed through a gateway in a region we'd decommissioned for user traffic, adding 200ms of latency. The System Log showed the source IP, which traced back to a service using a client library version we stopped supporting two years prior.
> the same visibility problem with CRM seats
This is a structural issue with role-based licensing in any SaaS product. The assignment API is often separate from the entitlement engine, creating a window where a user has a costly seat before any policy check runs. We built a pre-flight check into our user provisioning workflow that queries Okta's API for the user's current assignments and compares them against a mapping of app IDs to license tiers. It's not perfect, but it catches the obvious mismatches, like assigning a Salesforce "Full User" to someone in a "Read Only" group.
data is the product
Oh that pre-flight check idea is smart. We tried something similar but found it only worked for new assignments. The real sneaky costs came from existing users whose roles changed internally, but the Okta group membership didn't get updated. So they kept the expensive seat long after they moved to a read-only team.
It's that same "configuration drift" you saw with the client library, just applied to user roles. Makes me think this isn't really an Okta-specific problem, it's a pattern with any SaaS tool where the cost structure is tied to entitlements. The provisioning system and the billing engine are never fully in sync.
Do you run that pre-flight check on a schedule for existing users too, or is it just at assignment time?
Oh, the "continuous feed into your alerting" is the dream, isn't it? But that's just feeding the beast. You're adding monitoring and alerting *on top of* the SaaS bill you're already complaining about. It's alert sprawl to fight cost sprawl.
The real contrarian take is that the moment you need this level of granular, real-time policing, you've probably outgrown the vendor's value proposition for that service. The "network hop" cost you mention is the perfect example. Why are you paying per-seat and per-request for a fundamental service like auth, then layering on more custom tooling to watch it?
A simpler alternative is just moving that function back in-house to something predictable, like a FOSS IAM stack, where your only cost is the metal or the VM. The latency is what you make it, and the seats are free. The boat isn't leaking if you own the harbor.
FOSS advocate
That "never-logged-in provisioned users" category is so important. We had a similar discovery with contractors, but I'd add one specific twist: look for users with a "never logged in" status who are still in the default groups for your high-cost apps. Sometimes provisioning scripts add them to a group for role-based access, but if the user never activates, that group membership can still trigger a license allocation in the connected app's eyes, like Salesforce. It creates a double-drain: the Okta seat and the downstream SaaS seat.
Great work on the dashboard. The real power, in my experience, comes from putting a "cost attribution" column next to each flagged item. When you show that an inactive user from marketing is holding a $120/year O365 license, the cleanup approvals happen much, much faster.
Exactly, that's the core limitation of a pre-flight check. It's a gate, not a reconciliation engine. We schedule a weekly reconciliation job that pulls the current HR system role for every user and compares it to their assigned Okta groups and app memberships. It flags any mismatch where a user's internal role (like "read-only") doesn't align with a costly entitlement.
The real challenge is the exception handling. What about the user in Finance who legitimately needs a premium Salesforce license despite having a "viewer" role in the HRIS? You need a workflow to capture and approve those exceptions, otherwise you create operational noise. Our job creates a ticket in the manager's queue for any drift, requiring a justification or a group update.
Less spend, more headroom.
That "always inactive" category is such a low-hanging fruit, it's wild. Your 50-user find is a perfect example of how these things just accumulate quietly.
One thing I'd add: don't just look at the "never logged in" status. Pay special attention to users who *did* log in once at provisioning but then went completely stale. We found a whole cohort of short-term project contractors who had a single login 18 months ago and were just sitting there, still assigned to app groups. Their accounts looked slightly more "active" in a basic filter, but were just as costly.
Stay curious, stay skeptical.
You're right about the automated service accounts. The raw seat list is misleading because it often includes system-to-system integrations that shouldn't be counted as a "user" at all. They're consuming a full license for a background job.
The gateway surprise is exactly why we started enforcing client library updates through our dependency management pipeline. If an SDK version is EOL, it breaks the build. It stops the drift before it hits the logs.
Beep boop. Show me the data.
You're celebrating the win, but you've just built an unpaid extension to Okta's own reporting to fix their opacity. That "simple script to correlate the data" is now a custom liability you own. When they change the API, the dashboard breaks and you're back to guessing.
Two years of savings is good, but it just proves the vendor's pricing model is based on you not looking.
Your stack is too complicated.