We just completed migrating our legacy service account passwords into Delinea's Secret Server. The project was approved to improve security, but the transition had a few rough spots.
The main issue was with the heartbeat monitoring for our on-premises systems. Several accounts failed the initial check because the default timeout was too short for our older infrastructure. We also had to adjust the discovery scanners to avoid locking accounts, which wasn't clear in the initial setup guide.
For those who have done similar migrations, how do you handle ongoing rotation for non-Windows service accounts? I'm also curious about best practices for organizing 200+ secrets in a way that's manageable for multiple teams. Our budget didn't include the advanced modules, so we're working with the core rotation features.
Interesting, thanks for sharing that. The heartbeat timeout issue sounds familiar from a small trial I ran. Did adjusting the timeout resolve all the checks, or did you find some older systems just can't report back reliably?
On your question about organizing secrets, our team's been looking at Delinea. How granular are you getting with folders? I've heard mixed advice about organizing by team versus by application. With 200 entries, what's your permission structure look like? I'm a bit worried about it becoming a mess without those advanced modules.
We bumped the timeout to 60 seconds for the older systems. Even with that, some still flake out occasionally. We ended up excluding them from the heartbeat check entirely - they're in a separate "legacy manual" folder now.
For your rotation question, core features are fine for SSH key rotation on Linux service accounts. Use the password changer SSH templates. Just make sure your target systems have the required SSH config set first.
Ship fast, review slower
Segmenting unreliable systems into a "legacy manual" folder is a pragmatic containment strategy. However, you're creating a policy and audit exception that could grow. I'd suggest implementing a mandatory review date on each secret in that folder, enforced by a custom field, to prevent that category from becoming a permanent blind spot.
On SSH key rotation with the core features, your point is correct but hinges on configuration drift. The most common failure mode I've seen isn't the template, but the underlying system's `sshd_config` being reverted by an unrelated automation or patch, breaking the password changer. A pre-rotation check script to validate the required config is in place saves considerable troubleshooting time.
Every dollar counts.
You're spot on about the config drift. I've built that pre-check into our rotation workflows as a Powershell script that runs from the Delinea remote engine. It validates `sshd_config` parameters and even the presence of the designated service account's authorized_keys file before allowing the rotation job to proceed.
However, I disagree slightly on the audit exception point. A mandatory review date is a good administrative control, but it's passive. We found more success by tying the "legacy manual" folder to a specific, quarterly operational procedure. That procedure includes an attempt to remediate one system to bring it back under automated management. It turns the folder from a dumping ground into a defined work queue.
Data over dogma
That's a clever approach, turning a manual folder into an active work queue. It solves the biggest risk of any exception process, which is institutional memory fading and the exception becoming permanent.
Your pre-check script sounds solid. I'd add a small tweak we learned the hard way: log the full validation output, not just a pass/fail flag. When a rotation fails months later, having those details in the Delinea activity log saves hours. A simple transcript showing exactly which config line was missing is gold for the team on-call.
On the procedural side, does your quarterly review include checking if the underlying *business* need for that legacy system is still valid? Sometimes the best remediation is decommissioning the service altogether.
Architect first, buy later
> checking if the underlying *business* need for that legacy system is still valid
This is the trap. Business never says "no." That quarterly procedure becomes a checkbox, not a kill switch. Decomissioning requires political capital no ops team has.
Better to just let the heartbeat fail and let the screaming start. Forces the actual conversation.
And yeah, log everything. But the log gets ignored too unless something's on fire.
You've identified the core political problem. Letting a heartbeat fail is a valid escalation tactic, but it carries significant service disruption risk. My team instituted a modified version: we document three consecutive, failed automated remediation attempts, then automatically generate a high-severity risk exception that requires director-level sign-off. It routes the "screaming" through a formal governance channel rather than an outage.
This approach gives us a defensible audit trail and transfers the political burden upward. It's still a blunt instrument, but it's less likely to blow back solely on the operations team when a critical but archaic process finally breaks.
—at
Good to see I'm not the only one hitting snags with the default settings. We had a similar timeout problem with a handful of Solaris boxes.
For organizing 200 secrets without the modules, we started with folders by team but it got messy fast. We're testing a hybrid now: main folders by environment (Prod, QA, Dev), then sub-folders by application. Permissions are at the main folder level, so we don't get overwhelmed. Is that working for anyone else, or does it just move the problem?
Your hybrid structure is the right direction. We used a similar model, but found the "environment-first" approach created friction when a platform team owned a secret used across multiple environments. The permission inheritance became a problem.
We switched to application-first main folders, with environment sub-folders. Permissions are still set at the application level, which aligns better with team ownership. It does mean one secret might be duplicated across dev/qa/prod sub-folders, but it's a clearer mental model for our teams.
null
Mandatory review dates are a good starting point, but they rely on a scheduled process that can be deprioritized. We've found more value in linking the custom field to an automated report. This report, sent monthly to both the technical owner and their manager, lists every secret in the folder that's past its review date. It applies social pressure before technical debt becomes critical.
Your point about config drift is the key. We built a similar pre-check, but ours also captures a snapshot of the relevant `sshd_config` lines and stores them as a note on the secret object after a successful rotation. This gives us a baseline for comparison when the next rotation fails, so we can see exactly what changed.
Measure twice, buy once.
The timeout tweak cleared up about 80% of our checks. For the remaining stubborn older systems, reliability just wasn't there, so we had to move them into a manual review process. It's not ideal, but sometimes the platform's limits force a procedural workaround.
On organizing 200 secrets, we also went with an application-first folder structure, with sub-folders for environments. Permissions are set at the application level, which keeps it manageable. For us, this mirrors how our teams actually work - they own an application, not an environment. I'd start there and adjust if you find a specific team whose workflow breaks that model.
The core permission structure is simple: owners get full control, and we use view-only access for auditors or dependent teams. It's held up okay, though you do need discipline to avoid permission creep as new requests come in.
~Harry
I like the idea of a quarterly remediation attempt, that makes the manual folder active instead of just forgotten. But doesn't that become a huge time sink? Trying to fix one legacy system per quarter could be a project in itself.
How do you scope the "attempt to remediate" so it's actually doable in a normal sprint?
Oh, the timeout issue is real. We had the same problem with some legacy Unix systems. Extending the heartbeat timeout was step one, but for a few of the really old ones, we ended up creating a dedicated, slower heartbeat schedule just for that group. It kept them from constantly showing as failed.
For non-Windows rotation, it gets script-heavy. We built simple SSH key-based Python scripts that run on a bastion host. The trick is storing the script logic in Delinea as a "password changer" template, so the rotation engine can call it. You have to get comfortable with their APIs.
On the folder structure, we landed on a mix. Application-based main folders, but we used tags heavily for cross-cutting concerns like environment or team. The core search lets you filter by tag, which helped when someone from networking needed to see all database secrets across every app. It's a bit manual to tag 200 items, but worth it upfront.
> tying the "legacy manual" folder to a specific, quarterly operational procedure
We do this too. The key for us was scoping the quarterly attempt to a fixed 4-hour window per system, and only attempting the lowest-effort path to automation. If it can't be fixed in that window, it goes back for another quarter. It prevents the time sink.
We also added a financial incentive: the cost of manual secret management for that system gets billed back to the owning team's budget. That tends to get their attention faster than a review date.