Your question about alert routing is the operational key to making this shift sustainable. In our deployment, the initial flood went to a central platform team, which created a critical bottleneck and delayed remediation.
We implemented a tiered model after the first month: critical lease failures or root token usage still page the platform team, but routine warnings like token TTL expiration are routed to a dedicated Slack channel tagged with the owning service team. This forces developer ownership of their service's secret hygiene without waking them up for minor issues.
The harder part was defining "critical." We tied it to business impact: an alert is platform-team critical only if it affects a service's ability to serve production traffic. Everything else is a ticket or a notification for the dev team. This required mapping every secret's purpose, which was tedious but ultimately reduced our alert volume by about 60%.
That initial flood of alerts is your true audit log starting to speak. Your "set and forget" policies were a best-guess hypothesis, and the alerts are the experiment's results coming in. The trick is to treat them as data, not just alarms.
Most teams initially over-rotate on strict TTLs and tight policies. The noise forces you to actually look at the usage patterns. We found a third of our "token nearing expiry" alerts were for batch jobs that could safely use a static role with a one-year lease, which we'd never have considered without the pressure of the alert volume. The policy became a living document.
Your comment about integration points becoming critical infrastructure is the real cost. You've now made Vault a single point of failure for observability. A blip in its health can drown out actual service issues. Correlating Vault alerts with your APM metrics is the next, un-budgeted phase.
Trust but verify – and audit