You're absolutely right about the SCIM spec edge cases becoming a hidden project in themselves. We had a similar experience where "partial updates" turned into a major headache when only a user's department changed. The sync logic had to get much smarter to avoid unnecessary churn.
That control over the source of truth is the ultimate payoff, though. Once you can tag service accounts at creation, not only does it solve the audit problem, but it also prevents those accounts from ever polluting your HRIS data join. It's a permanent fix, not just another cleanup.
—HR
You're right on the money about it being a direct cost center now. That script is a good first step, but like others have said, `last_login` isn't always reliable for cleaning up.
When you run that SQL, be prepared for a lot of manual review to avoid breaking something tied to a CI/CD pipeline or a forgotten reporting tool. The time cost of that validation is real and has to go into the TCO model, not just the scripting time. Have you factored that in?
Keep it civil, keep it real.
>manual review to avoid breaking something tied to a CI/CD pipeline
That's the killer. We built a script to flag candidates for cleanup, but the validation took twice as long as the scripting. You need a registry for service accounts, period. Without it, every flagged account is a potential landmine you have to manually trace. Who owns that Slack bot from 2021? No one knows, but it'll break the weekly report.
For our TCO, we started logging those investigation hours. After three months, the "maintenance" cost was higher than the original license hike we were trying to offset.
Exactly. That's the trap of the DIY cost model.
You meticulously log engineering hours for the script, but the real bleed is in the endless, unplanned tribal knowledge calls. "Who wrote this integration? Does anyone remember the Friday deploy script service account?" Suddenly you're paying senior dev rates for enterprise archaeology.
And the vendor knows this. They price their hikes just under the threshold where that investigation work *feels* cheaper than paying up. It's a tax on your organizational memory.
Trust but verify.
That permanent fix angle is key, but I'd question if any fix is truly permanent with service accounts. The registry and tagging solve the creation problem, but you still need a deprecation process for when those services are sunset. Without one, you're just pushing the cleanup cost down the road.
We built tagging at creation and still ended up with a graveyard of tagged-but-unused accounts because nobody owned decommissioning the old reporting stack. The audit shifts from "what is this?" to "why is this still here?", which is easier but not free.
Less spend, more headroom.
Totally agree on the payoff of that source of truth control. Tagging at creation was a game-changer for us too, but the "permanent fix" part only clicked once we wired the registry into our deployment pipelines.
Now, when a Terraform module spins up a new service, it automatically registers the account with the required tags in our directory. No human forgets. The audit script just reads the registry; no more guessing games.
We did find a new headache, though - keeping the registry itself clean when services get deprecated. That's the next automation project!
Keep deploying!
Wiring it into Terraform is such a smart idea. It makes the tag automatic instead of just a new manual step to forget.
When you say keeping the registry clean for deprecated services, does that mean your de-provisioning script isn't automated yet? I'm guessing you'd have to tie it into whatever tears down the service itself, which sounds tricky if the tools are different.
Yes, that's a crucial first step. Your point about modeling the total bill, not just the per-user price, can't be overstated. We've seen teams miss the compounding effect of the platform fee across their user base and get a nasty surprise.
Your SQL filter is a solid start, but I'd caution that `last_login` can be a tricky metric for service accounts. Some might not log in conventionally but are essential for background jobs. Have you considered cross-referencing with access logs or API call history to reduce false positives?
—HR
You're highlighting the critical difference between using activity logs and relying on the source of truth. I've seen this exact pattern fail.
Joining against the HRIS feed is essential, but it introduces its own data lag. If your HR system updates nightly, you're still carrying terminated accounts for up to 24 hours in your cost calculations. For a large organization, that's a constant, recurring cost footprint that your TCO model often misses.
Your point about creating a manual ETL job is spot on. We documented the initial build time but completely underestimated the validation and maintenance cycle. Every schema change in the HRIS feed or the JumpCloud API broke the script, requiring unplanned work. The total labor over a year dwarfed the initial build effort, which is the hidden cost that makes the "lighter-weight service" argument compelling.
Data > opinions
Ah, the predictable allure of the 18-month break-even point. That's the classic line used to justify every migration project, isn't it? I'm curious about the "full-time-equivalent multiplier" for the first two quarters. Did that multiplier account for the inevitable project drag, or was it based on a clean-room estimate from the team that wanted to build it?
My experience is that the break-even horizon tends to stretch once you realize you're now running a software business. Your team that built the sync workers gets promoted, leaves, or gets pulled onto the next fire. Then you're paying to maintain a critical, home-grown system with institutional knowledge that just walked out the door. The managed service's annual price hike starts to look a lot like a fixed-cost insurance policy against that attrition.
cg
That SQL filter is a logical first step, but relying solely on `last_login` is a common pitfall. Many critical service accounts authenticate via certificates or API keys, never generating a traditional login event.
You need to cross-reference with application logs. For example, join your directory data against the audit trail from your primary applications to see which accounts are actually making requests. An account with no login for a year but daily API calls to your data warehouse is not a cleanup candidate; it's a production dependency.
Without that data layer, you risk deprovisioning accounts that will silently break nightly batch jobs or monitoring alerts.
benchmark or bust
Bingo. This is the exact mental shift that took us six months to make. We spent all that time polishing the "find orphans" query when the real problem was we had no authoritative list of what *wasn't* an orphan.
> treating it as a render target
That phrase is gold. We moved our source of truth to a Terraform module (HCL in Git) and it completely reframed the cost. Suddenly the budget discussion was about merge request velocity and peer reviews for new service accounts, not about the mystery bill from the IDP. The vendor's price increase email became almost irrelevant, because our provisioning logic was ours.
pipeline all the things
Your action plan is spot on, especially the emphasis on modeling the total bill. We ran those numbers and found the platform fee effectively doubled our per-unit cost for our core admin group, which was a brutal realization.
The SQL cleanup script is a necessary first step, but I'd add a caveat from our own audit: you need a parallel process for validating what you find. We nearly deactivated a handful of service accounts that hadn't logged in for over a year but were tied to legacy on-premise backup jobs. The `last_login` field in the directory was null, but the backup system's own service account was still authenticating via an old key. We now cross-check any candidate list against a consolidated report from our monitoring and job scheduler platforms before taking action.
It turns a simple script into a manual review process, but it prevents the silent breakage user200 and user404 mentioned. The price increase forces this hygiene, but the operational cost of doing it right is the real budget impact.
Support is a product, not a department.
The partial updates headache is real. We spent a week optimizing our sync just to stop it from rewriting every user's entire profile for a changed phone number.
That permanent fix idea works until your HRIS decides to change its attribute naming convention. Then your clean source of truth is suddenly full of unmapped fields, and you're back in the soup.
Yep, that mapping layer seems to be the real problem. We tried to build one ourselves and it felt like playing whack-a-mole every time the source system updated.
Does anyone just bite the bullet and use a commercial sync tool for this? I'm wondering if the maintenance cost of our homemade script actually outweighs a vendor license.