Rotating service principal secrets (client secrets/certificates) for non-human identities in Entra ID is one of those necessary evils that can silently blow up your cloud budget if it goes wrong. A failed rotation can lead to application outages, which in turn can cause auto-scaling failures or resource sprawl that racks up cost. I'm always looking for patterns to automate this safely, especially across hundreds of principals.
My current approach uses a combination of Azure DevOps pipelines and key vaults, with a mandatory overlap period. The core idea is to never have a single point of failure during the rotation.
Basic workflow for a single service principal:
1. Generate a new secret in Azure Key Vault (with an expiration date).
2. Update the consuming application's configuration *in a staged manner* (e.g., via a feature flag or a secondary config slot).
3. Validate the new secret works.
4. Only then, remove or disable the old secret after a grace period (e.g., 72 hours).
The real challenge is doing this at scale. You need a reliable inventory. I use a PowerShell script to audit all principals and their secret expiry dates, which feeds into a rotation schedule.
```powershell
# Sample to find expiring secrets (requires Microsoft.Graph module)
Get-MgServicePrincipal -All | Where-Object {
$_.PasswordCredentials.ExpiryDateTime -lt (Get-Date).AddDays(30)
} | Select-Object DisplayName, Id, @{Name="Expiry";Expression={$_.PasswordCredentials.ExpiryDateTime}}
```
My questions for the community:
* How do you maintain a canonical list of *where* each service principal is used? This metadata management is the hardest part.
* For large-scale, do you prefer:
* A centralized orchestration (like a scheduled Azure Function) that rotates based on tags?
* Or a decentralized, per-application pipeline that teams own?
* What's your fail-safe mechanism to avoid simultaneous secret invalidation across multiple dependent services?
I've seen cost spikes from VM scale sets stuck in a failed provisioning state due to a bad secret, so getting this right is a FinOps issue too.