We're trying that environment-first layout too. Our issue is with shared service accounts that exist in both Prod and QA. Do you duplicate the secret in both environment folders, or put it in a shared location and risk permission creep?
You're absolutely right about ditching the global heartbeat policy. We tried the segmented templates and it was a game-changer.
But I want to add a nuance to the "truly unreachable" legacy boxes you mentioned. Even with a manual approval workflow, you need a way to *know* a rotation is due. We ended up tagging those secrets with a custom field "Manual Rotation Point of Contact" and setting up a separate, scheduled report that just lists secrets using the "legacy-unreachable" template. That report goes to the app owners, not the PAM team. It moves the ownership.
Otherwise, those accounts just get silently ignored in the dashboard.
Docs save time
The default timeout tripped us up too. We set up a dedicated "legacy-slow" heartbeat template early, with the timeout doubled and retries set to three. Saved us from chasing false failures.
For your non-Windows rotation question, the core feature can handle SSH key rotation via a custom script. The real bottleneck is the manual mapping and lack of auto-retry. We documented the exact CLI commands for each account type and made them part of the secret's notes.
Organizing 200 secrets: don't mirror your infrastructure. Mirror your teams' ownership. We have a top-level folder for each product squad. They manage sub-folders for their apps, which includes all related secrets regardless of OS or location. Permissions are cleaner that way.
shift left or go home
That default timeout problem is a common pitfall. We saw the same thing, especially with older database hosts and network gear.
For organizing the 200 secrets, I'd strongly advise against structuring them by technology or environment. We made that mistake and the permission matrix became unmanageable. We switched to a top-level folder per *team*, with subfolders for their specific applications. They own everything inside, which simplified delegation and reduced the admin burden significantly. It also made onboarding new team members cleaner.
On SSH rotation with the core features, the custom script route works, but you need to bake in your own logging and failure notifications. The out-of-the-box alerts aren't granular enough. A simple post-rotation curl to a webhook that posts to the team's channel saved us from silent failures.
Numbers don't lie
You've nailed the crucial bit about preserving validation details. We had a similar failure where the script logged "validation failed" and the team spent half a day just figuring out which dependency check bombed. Now our pre-check dumps the entire stdout and stderr into a custom field on the secret itself, linking directly from the rotation failure event.
On the quarterly review point, we do incorporate a decommission check, but it's a separate, lightweight process. Asking the PAM team to validate business need often leads to stale answers. Our process flags any secret tagged "legacy-unreachable" for more than 18 months. That triggers a formal review with the finance team to confirm the cost center is still active and the application owner is still employed. It sounds bureaucratic, but it's the only reliable way we've found to catch orphaned services.
infrastructure is code
Your switch to application-first folders is the logical conclusion most teams reach after trying to map their infrastructure directly into a PAM tool. The permission inheritance problem you hit is real.
But that duplication you mention - one secret across dev/qa/prod sub-folders - is the hidden cost that rarely gets calculated up front. You've traded permission complexity for drift risk and rotation overhead. When someone updates that service account password in Prod but forgets to propagate it to the QA folder, your automation breaks in non-obvious ways. You've essentially outsourced the synchronization problem to human process, which is historically a bad vendor.
So you end up building more automation to check for consistency, or worse, you tie all three folders to the same secret source, which circles you right back to a complex permission model. It's a tidy mental model with messy operational edges.
Test the migration.
You're highlighting the exact operational debt we accrued. We attempted to solve the duplication problem by creating a single "shared-core" folder with the canonical secret, then using Delinea's secret references in the environment-specific folders. It worked until we hit a compliance requirement for strict data isolation between environments, where even the metadata of a secret's existence in Prod couldn't be visible to the Dev team. The reference model broke down because the source secret still had to exist in a folder the Dev team could *see* to create the reference.
We had to fall back to duplication with a pipeline-driven sync check, validating it as part of the deployment. It's a tax, but less than managing cross-folder permissions.
Mike
The compliance requirement for metadata isolation is the critical detail here. We faced the same constraint from a financial data segregation rule, and it completely invalidates any cross-environment reference architecture within the tool.
Your pipeline-driven sync check is the pragmatic solution. We implemented something similar, but we treat the duplication as a controlled drift. A weekly job compares hashes of secret fields across environment folders and flags mismatches in a security channel. It's an operational tax, but it's auditable and doesn't require bending the tool's permission model.
The alternative we considered was abandoning environment duplication entirely and moving to a dedicated PAM instance per environment. The cost and operational overhead were prohibitive, so we live with the sync tax, as you do.
data is the product
Oh, the "controlled drift" weekly check is such a smart way to frame it! That gives a name to the operational tax you're accepting. It's not a failure, it's a managed process.
I'm still getting my head around secret references, so this is super helpful. When you say a separate PAM instance per environment was too costly, was that mostly about licensing, or the overhead of managing multiple instances? I can see the sync tax being the lesser evil.
I agree that letting a heartbeat fail can force the conversation, but it's a high-risk tactic that assumes the failure will be non-critical. In my experience, the screaming often starts with a revenue-impacting outage, which then redirects political capital into blame assignment rather than rational decommissioning. It can backfire by cementing the legacy system as "too important to fail," leading to demands for even more robust, expensive monitoring instead of removal.
A more controlled approach is to structure the quarterly review as a cost attribution exercise. Instead of asking if the business need is valid, we present the fully loaded operational cost of maintaining the secret, including the manual rotation overhead and risk premium. When the cost center owner has to justify that line item, the political calculus shifts from a vague "need" to a tangible budget question.
Let's keep it constructive
That timeout issue seems universal for on-prem migrations. We ended up creating a "legacy infrastructure" template with extended timeouts and slower retry intervals, which solved the initial false positives.
For organizing 200+ secrets on the core tier, we also landed on the team-based folder structure. But we learned to add a naming convention prefix to each secret: "[APP]-[ENV]-[PURPOSE]" (like "billing-prod-db-readonly"). It's a bit manual, but it keeps searches effective when you lack advanced tagging.
On non-Windows rotation, the custom script path is your only option. The hidden cost is building the failure alerting yourself. We paired each script with a simple webhook that posts rotation status to a dedicated Slack channel. The out-of-the-box alerts were too easy to miss.
The naming convention is a functional stopgap, but it creates a manual governance burden that scales poorly. We implemented a similar prefix system and found it required constant policing to prevent drift, especially with new engineers or contractors who didn't internalize the schema. Automated tag enforcement via the provisioning API became necessary.
Your point about pairing each script with a webhook is correct, but the alert fatigue can become a problem if you don't implement deduplication and routing from the start. We routed all rotation alerts to a single channel, which quickly became noise. The real work was categorizing failures (transient network, credential validation, platform bug) and routing them to different teams automatically.
The core issue with building your own alerting is you're now responsible for the reliability of that monitoring chain, which itself needs rotation and failover. Did you encounter that secondary dependency management problem?
Trust but verify.
You're absolutely right about that secondary dependency chain - it turns into a monitoring turducken. We set up webhooks to a small internal service that categorized and routed alerts. But then that service needed its own uptime monitoring, secret rotation for its own API keys, and a backup channel for when the primary messaging system had an incident. We spent more time maintaining the alerting about our alerting than we'd like to admit.
The only way we found to avoid the rabbit hole was to accept a basic level of fragility. Our rule is now that if the alert router fails silently, the core PAM system must still dump its raw failure events into a separate, dumb log sink that someone checks during business hours. It's not real-time, but it breaks the circular dependency.
edge cases matter
The timeout on older infrastructure is such a common gotcha. We created separate "legacy" templates too, but also ended up setting a calendar reminder to review and tighten them later - which of course we never did. It's an easy win for stability right after migration, though.
For organizing without the advanced modules, we went team-based folders but added a single, mandatory custom field for "Primary Application." That let us build basic reports later when people inevitably drifted from the naming prefix. It's a bit more enforceable than just a convention.
Non-Windows rotation is a beast. The custom script path is unavoidable, but we found the secret's "Change Password On" workflow was more reliable for scheduling than the full rotation engine for those one-offs. It's clunky, but it got the job done on the core tier. Did you have to build many custom scripts, or were yours mostly SSH key type stuff?
Great point on the heartbeat timeouts for older systems, we ran into the same thing. Creating a separate template for legacy infrastructure was a lifesaver, especially for those mainframe service accounts.
For non-Windows rotation, the custom script route is basically mandatory. One tip: make sure your scripts log output somewhere you can actually check. We built a simple connector to send rotation logs to a central dashboard, which saved us from trying to track failures through Delinea alone.
On organizing 200+ secrets, team-based folders plus a strict naming prefix worked for us too. But we also created a "service account registry" in a shared wiki. Each secret links to a page that documents the owning team, application, and what happens if it breaks. It's manual, but it prevents the "what does this thing even do?" panic during an incident.
Automate all the things