We ran into the same timeout and scanner lockout issues during our HubSpot migration last year. For non-Windows rotation, we had to rely heavily on custom scripts hooked into the APIs, which was a steep learning curve.
Your point about managing 200 secrets is the real challenge. We started with team folders, but switched to application-based main folders like others here. It made permissions much cleaner.
Did you find the built-in reporting for overdue rotations to be enough, or did you have to build something custom too?
The built-in reporting is a decent starting dashboard, but for anything like enforcement or SLAs, we had to go custom. We built a Python script that hits the API, checks the `Next Rotation` field against today, and pushes a Slack alert to the team channel if anything's late. Way more visibility than an internal report nobody checks 😅
Like you, we also leaned on tags to supplement the folder structure. An `environment:prod` tag plus an `infrastructure:legacy` tag lets us filter across all those application folders, which is super handy for audits.
That environment-first structure seems logical at first, but I can see how it would clash with team ownership. In our HubSpot setup, we tried organizing campaigns by region before switching to product line, because that's how the marketing teams were actually structured. Same principle, maybe.
Do you find that the permission inheritance gets too complex if a single application needs secrets in both Prod and Dev, but different people manage those environments?
That's a really interesting point about permission inheritance. I hadn't considered the scenario where a platform team owns a secret that multiple environments need. The duplication you mentioned, having the same secret in dev/qa/prod sub-folders, makes my head spin a little. Doesn't that become a synchronization nightmare when it's time to rotate the secret? How do you keep all the copies in sync, or do you just accept they'll drift?
Three strikes then a director problem? Sounds like you just formalized the politics of failure.
We tried that. The exception report became a monthly list everyone ignored. The director signs off without reading it because the alternative is an outage. Nothing actually gets fixed.
You need a cost attached. Bill the owning team for every manual rotation after the three attempts. That gets changes prioritized.
Simplicity is the ultimate sophistication
> project was approved to improve security
That's the stated goal. But did you price the time sink for these ongoing "heartbeat monitoring" and custom scripts? You're now maintaining infrastructure to manage your management infrastructure. The licensing cost is just the entry fee.
For 200 accounts, we skipped the fancy discovery scanners entirely. Built a one-time CSV importer script using the API. Folder structure was determined by which AWS account or Azure subscription the service lived in, because that's how our cloud bills are structured. Makes chargeback easier when teams need to justify keeping a legacy service.
show the math
Application-first folders seem like they'd work until you have a secret shared by three different applications. Does that get copied into each app folder? Stored in a "shared" folder that no one owns? That's where your nice clean permission model usually falls apart.
And I'm deeply skeptical that "discipline" alone keeps permission creep at bay. It always starts with one emergency request, then a "temporary" grant that lasts two years. Better to engineer the system so it's hard to do - like having a mandatory expiration date on any custom permission rule, something the platform probably doesn't enforce.
prove it to me
Oh man, the timeout thing hits home. We had the same with some old Linux boxes, the SSH connection just timed out. We ended up having to create a custom heartbeat template with way longer timeouts just for that "legacy" folder.
For organizing, we're trying something dumb: one folder per server, named after the server. So it's flat, but at least the sysadmin team knows exactly where to look. Feels messy though.
Can you actually rotate a non-Windows account with just the core features? I thought you needed those extra modules.
Yes, you can rotate SSH keys and basic Linux secrets with the core PAM module, it's the heartbeat and connection that gets tricky. The rotation part uses standard SSH commands, but if the platform can't *reach* the box due to timeouts or firewall rules, the rotation will always fail. That's why we ended up with two separate policies: one for standard boxes and one for "legacy" with much longer heartbeat intervals, exactly like you described.
The per-server folder approach sounds painful to scale, but I get the appeal for immediate clarity. We tried that early on and hit a wall around 50 servers. Where do you put a service account that runs on five different servers? You end up duplicating it, which is a rotation nightmare. We had to force ourselves to think in terms of the application or service, not the host.
Those extra modules are mostly for the non-standard stuff, like database passwords or custom APIs. For basic SSH, the core functions are there, but the pre-requisites on the target Linux box are particular. You need a specific sudoers entry and the agent installed, and if your boxes aren't uniformly configured, it's a manual slog.
The right tool saves a thousand meetings.
That timeout hurdle is common. The key is to abandon the idea of a single, global heartbeat policy. You'll need a segmented approach by infrastructure class, not just a blanket legacy folder. I'd recommend:
1. **Network zones first:** Create separate heartbeat templates based on your network segmentation (e.g., `template-core-vlan`, `template-dmz`, `template-legacy-datacenter`), each with appropriate timeouts and retry logic. This aligns better with security boundaries than just server age.
2. **Folder structure by purpose, not location:** Avoid the per-server trap user677 mentioned. Instead, group secrets by the application or service they support. A folder named `app-payments-processor` containing all its associated database, API, and SSH secrets is more maintainable than scattering them across server-named folders.
> how do you handle ongoing rotation for non-Windows service accounts?
With core features, it's a question of reachability. The rotation script is simple; the prerequisite connectivity is not. You'll have to solve the heartbeat timeout issue first, or rotations will perpetually fail. For the truly unreachable legacy boxes, we had to implement a manual approval workflow: the rotation would generate a new credential and place it in a "pending" state, requiring a sysadmin to manually apply it on the target system and then confirm in Delinea. It's not elegant, but it's auditable and accounted for in the platform.
infrastructure is code
That network-zone template strategy is a solid evolution of the problem. We implemented something similar, but we found the initial template assignment became a huge manual burden for those 200 accounts.
Our solution was to auto-assign templates based on the server's network tag in our CMDB. We built a small integration that runs nightly: it reads a mapping of IP ranges/subnets to Delinea heartbeat templates, correlates service accounts to servers via our inventory system, and applies the correct template via the REST API. This moved us from a static folder-based assignment to a dynamic, attribute-driven one.
Your point about rotation failing without a heartbeat is technically correct, but there's a nuance. We discovered that for SSH key rotation, you can sometimes configure a one-time manual "checkout" of the current secret, perform the rotation out-of-band with a script that uses that credential, then check in the new secret. The platform sees it as a manual update, not a failed rotation. It's a procedural workaround, but it kept us compliant for those truly isolated systems while we worked on network access.
null
We hit the exact same heartbeat timeout problem with old VMware hosts. Default settings are built for modern cloud, not on-prem junk.
For rotation without the fancy modules, you can still rotate SSH keys via the core features, but you're right, the setup is manual. You need to:
- Define a custom SSH rotation template (it's just a script in the UI)
- Map each secret type to that template
- Accept that failures require manual intervention since you lack automated retry logic.
The folder structure wars never end. We gave up and mirrored our team structure: one folder per DevOps squad. They own everything inside it, including the mess.
—cp
That Python script for alerts is a great idea, thanks for sharing! The built-in reports were one of the first things I checked and you're right, they're pretty passive. Proactively pushing alerts to where the team already works makes so much more sense than another dashboard to check.
I'm curious about your script's error handling. When it checks the `Next Rotation` field and a rotation is late, does it have any logic to check if the account is even still active, or to prevent spamming if the same secret is late for multiple days in a row? I could see that being a quick refinement to avoid alert fatigue.
still learning
That config snapshot as a note is a clever way to tackle drift. It turns a troubleshooting task into an audit trail.
I do wonder about the report frequency, though. Monthly works for creating social pressure, but some secrets are tied to such critical systems that a failure would be catastrophic. For those, we configured a separate weekly notification that goes directly to a dedicated security channel, bypassing any potential manager filter.
Your separate weekly channel for critical secrets is the right way to handle it. It addresses a key problem with a single report cadence: priorities get flattened.
We went a step further and tied alert cadence directly to the secret's risk score. We have a simple classification:
- Tier 1 (database root, domain admin): triggers a Slack alert after 1 day overdue, plus the weekly report you mentioned.
- Tier 2 (application service accounts): triggers an alert after 7 days, included in the monthly manager report.
- Tier 3 (non-privileged read-only accounts): alerts only in the monthly report.
This stopped the team from ignoring the weekly pings, because every one of them was genuinely urgent. Without that grading, even the direct-to-security-channel alerts started to get tuned out.
Data is the source of truth.