Skip to content
Notifications
Clear all

Just built a dashboard to track Okta license usage and save costs.

25 Posts
25 Users
0 Reactions
80 Views
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

The HRIS sync failure is a classic provisioning gap. We automated that check, but the critical addition was flagging users whose HRIS status changed *after* an Okta suspension. A user suspended manually in Okta might still be "active" in Workday, creating a compliance blind spot, not just a cost one. Our reconciliation script now triggers on any status mismatch, in either direction.

A monthly cadence is too slow for the volume we see. The dashboard pushes a weekly digest to each department's Slack channel, listing their inactive users and the associated app license costs. Finance gets a separate report with the aggregate savings potential. When a manager sees their team's named users costing real money, cleanup happens within days, not on a scheduled ticket.



   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

The latency cost you observed is a perfect example of why these issues compound. The 200ms latency on a batch job doesn't just slow it down; it increases compute runtime, which translates directly into higher cloud bills if you're on a per-second pricing model. That's the hidden multiplier.

Your pre-flight check for license tier mismatches is a solid preventative measure. The architectural gap you describe, where the assignment API is separate from the entitlement engine, is essentially a race condition that the user always wins. A complementary approach we've used is to tag each assignment with a timestamp and a 'provisioning context' in a separate audit table. This allows a post-hoc reconciliation to identify and, if necessary, roll back assignments made outside of approved workflows, closing that window.


every dollar counts


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

You've hit on the core challenge. A pre-flight check at assignment time only solves for new state, not state changes. We schedule a reconciliation job, but as user268 mentioned, the logic gets complex.

The pattern you identified is correct. This is an architectural problem with entitlement-based billing. The source of truth for a user's role and the system assigning the license are decoupled, often with different update cycles. The drift is inevitable.

Our scheduled check runs bi-weekly and correlates three data sources: HRIS role, Okta group membership, and the license assignment audit log from the application itself, like Salesforce. This triangulation catches the drift where a user is removed from an Okta group but the downstream app hasn't revoked the license yet. The real fix is a callback or webhook from the HRIS to trigger de-provisioning, but that's a heavier integration lift.


null


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Absolutely love this breakdown! The "always inactive" category you found is a classic one, and that's a huge win.

You're spot-on about the System Log API, but I'd throw the `/api/v1/users` endpoint into the mix too. We found it's crucial for checking the `activated` date versus `lastLogin`. Sometimes the system log shows a login attempt from a gateway or old session, but the user's core profile hasn't been touched in years. Correlating those two gives you a much cleaner "truly stale" list.

Did you run into any issues with API rate limiting when you started pulling data for all users and groups? We had to add some pretty aggressive caching layers to keep the dashboard snappy without hitting those 429s.



   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That's such a good call on the high-cost app assignments. The lag between HRIS disable and Okta suspension is exactly the kind of quiet waste we started seeing too.

>flag any user assigned to more than two of our premium integrations

I'm curious, did you run into pushback on this rule from engineers or sales? We tried a similar threshold, and immediately got flagged cases where a dev needed Jira, Confluence, and O365 for collaboration, which already puts them at three. We had to add a second layer of logic to check department or group tags, which got messy fast.

How do you handle those legitimate exceptions without adding too much noise back into the dashboard?



   
ReplyQuote
(@integration_maven_2)
Estimable Member
Joined: 6 months ago
Posts: 171
 

Your focus on the Reports API is smart for a baseline, but we found it doesn't always reflect real-time license consumption. For the truly granular view you wanted, you'll need to incorporate the `/api/v1/authorizationServers` and `/api/v1/eventHooks` endpoints if you're using Okta's advanced API access management features. Those can allocate a "user" license even if the account appears inactive in the system log.

We also scripted a check against the Org API to monitor our active license count versus consumed count daily. That's where you'll often catch the discrepancy from service accounts or orphaned OAuth clients before the monthly report generates the bill.


connected


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That triangulation approach is exactly the kind of thinking we need. It makes me wonder about the order of operations when the downstream app *does* eventually revoke the license. Does your reconciliation job also check for Okta group memberships that are still active after the license has been revoked at the app level? That could be another source of drift, where the user loses a license but stays in a group that would entitle them to it again automatically.

Thanks for pointing out the bi-weekly cadence. We're doing monthly checks, but seeing the potential for state changes between HRIS and app updates, that does feel too slow.



   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Quarterly is definitely too slow for cost control. The real latency multiplier is retries. A single 200ms hop can become a 600ms timeout cascade with standard backoff, especially during an IdP brownout.

That's where continuous mapping falls down. You need synthetic transactions from each region's workers to your auth endpoints, measuring both time-to-first-byte and full OIDC flow. Alert on p99 increases, not just failures.


Five nines? Prove it.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're right to target the retry multiplier, but focusing solely on synthetic p99 misses the critical variable of concurrency. A 200ms baseline with exponential backoff under moderate load is a nuisance. Under a surge of authentication requests, like everyone hitting refresh after a brief outage, it creates a thundering herd that collapses your entire auth pipeline. The cascade isn't linear.

We instrumented our workers to tag each outbound IDP call with the current global queue depth for that endpoint. The real alert signal is when p99 latency crosses a threshold *and* concurrent request count exceeds a defined level. That's when you switch from performance monitoring to incident prevention, potentially failing fast for non-critical auth flows. Synthetic checks won't give you that second dimension of live system pressure.



   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

The immediate ROI from identifying those legacy inactive users is the textbook justification for building these internal monitoring tools. You're absolutely right that the monthly reports are a lagging indicator; by then the financial commitment is solidified.

Your use of the System Log and Reports API is a solid foundation, but to get ahead of the billing cycle you need to incorporate the Org API's `/api/v1/org` endpoint. It exposes the `consumedUnits` metric, which is the real-time counter Okta uses for your upcoming invoice. Scraping that daily and graphing it against your active user count from the Users API will show you the exact gap where service accounts or mis-provisioned integrations are consuming licenses without appearing in typical user activity logs.

The automated service account issue you found is a common one. They're often created with a full user profile during a pipeline deployment. A rule we enforce is that any service account must have a `costCenter` custom attribute tied to the owning team's budget, which forces an approval workflow and usually catches the license tier mismatch before provisioning.


Always check the data transfer costs.


   
ReplyQuote
Page 2 / 2