We've recently concluded a multi-phase migration of our identity stack to AWS, centralizing on PingFederate and PingAccess in ECS Fargate, fronted by an Application Load Balancer. The infrastructure is defined via Terraform, and we've implemented a relatively automated certificate management process using AWS Certificate Manager (ACM) for public certificates and a combination of AWS Secrets Manager and HashiCorp Vault for private key material. The issue arose during our first scheduled rotation of the SAML signing certificate used by our PingFederate instance for service provider connections.
The rotation itself, from a procedural standpoint, succeeded. The new certificate was provisioned, the private key was stored securely, and the PingFederate configuration was updated via its administrative API to start using the new key pair. However, this broke existing sessions and, more critically, new sign-on attempts for a subset of our service providers. The errors in the PingFederate logs were cryptic (`RuntimeException: Invalid signature on SAML response`), but the root cause was a mismatch in the certificate presented in the metadata versus the one used for signing.
Our architecture and the misstep are as follows:
* **Primary Certificate:** Stored in ACM for the ALB listener (HTTPS).
* **SAML Signing Certificate:** Private key in Secrets Manager, certificate in a Vault KV store. A Lambda function triggered by the Secrets Manager rotation updates PingFederate.
* **Metadata Management:** We use a custom, internally hosted metadata endpoint (behind the same ALB) that dynamically serves the federation metadata XML, pulling the active signing certificate from Vault.
The failure occurred because our metadata endpoint was caching the previous certificate's public key for a short period (60 seconds) after the rotation completed. During this window, and for any SPs that had fetched metadata earlier (depending on their `validUntil` cache duration), they were validating signatures against an old public key. The rotation script's steps were:
1. Generate new key pair in Secrets Manager.
2. Trigger rotation Lambda.
3. Lambda calls PingFederate Admin API to add new certificate, set it as active, and remove old.
4. Lambda updates the certificate reference in Vault.
The critical missing step was **forcing an immediate metadata refresh and invalidating SP caches**. We also failed to consider the `validUntil` attribute in our generated metadata.
Our current, working rotation procedure now includes:
* A pre-rotation step that sets the metadata `validUntil` to a near-future time (e.g., now + 5 minutes).
* Immediate HTTP cache invalidation for the metadata endpoint URL via a purge call to our edge CDN.
* A post-rotation step that actively notifies our major SPs (via their admin APIs) to refresh metadata.
* Staggered activation: The new certificate is added and used for new signatures, but the old one is retained (not marked as inactive) for a 48-hour overlap period to allow for cache propagation.
The corrected Lambda function logic snippet (Python) for the post-rotation step now looks like this:
```python
def update_pingfederate_signing_cert(new_secret_arn):
# ... retrieve new cert from secrets manager ...
pf_admin_api_url = os.environ['PF_BASE_URL'] + '/pf-admin-api/v1'
# Add new certificate (remains secondary, not yet active for all connections)
add_cert_payload = {
"fileData": new_cert_data,
"certificateId": new_cert_id
}
requests.post(f"{pf_admin_api_url}/certificates", json=add_cert_payload, auth=(user, pass))
# Update the signing certificate settings for the SAML IdP Connection
connection_id = "ouridpconnection"
conn_config = requests.get(f"{pf_admin_api_url}/idp/adapters/{connection_id}", auth=(user, pass)).json()
conn_config['signingSettings']['alternateSigningKeyPairRefs'] = [{"id": new_cert_id}]
# Keep the primary signing key as the old one for now
requests.put(f"{pf_admin_api_url}/idp/adapters/{connection_id}", json=conn_config, auth=(user, pass))
# Update Vault and metadata endpoint
update_vault(new_cert_data)
purge_metadata_cache()
# Notify major SPs via their APIs (implementation specific)
notify_service_providers(new_cert_id)
```
My question to the community is this: While we've patched our process, I'm concerned about the inherent fragility of this certificate lifecycle management, especially as we scale the number of relying parties. Has anyone implemented a more elegant, declarative approach for PingFederate certificate rotation, perhaps using its built-in key pair rotation features in tandem with external secret stores? Specifically, I'm interested in patterns that:
* Integrate with external KMS/HSM (like AWS CloudHSM) for key generation and storage.
* Provide a truly zero-downtime rotation for a large portfolio of SP connections without manual notification steps.
* Leverage PingFederate's "Incoming Profiles" or "Key Pair Rotation" features in an automated, infrastructure-as-code workflow.
Our current solution feels like a series of procedural scripts glued together, which is a maintenance burden and a potential single point of failure. I am looking for architectural insights or review of our revised process for hidden pitfalls.
That's a classic metadata propagation failure. Your rotation procedure updated the active signing key but the service providers were still validating against the old certificate fingerprint fetched from your SAML metadata endpoint. The metadata must be updated and propagated *before* the new key becomes active for signing, otherwise you create this exact validation window.
You need to sequence the steps: publish metadata with the new certificate (and ideally keep the old one listed for a grace period), wait for SP metadata refresh cycles, *then* switch the active signing key in PingFederate. Automating this requires a state machine that tracks the metadata publication timestamp and enforces a minimum delay before flipping the active key.
You've hit on a key operational challenge that goes beyond just the cryptographic swap. user1330 is right about the metadata sequence, but I'd add that your automation's success might be its own blind spot. When the API call to update the active key succeeds, the script thinks "job done," but PingFederate's internal state and the distributed state of all relying parties are a different matter entirely.
For a future rotation, consider baking in a verification step that polls a sample of your critical SPs (if they have test endpoints) or at least confirms your own metadata endpoint reflects the new certificate *before* flipping the active key. This turns a silent failure into a controlled halt in the pipeline.
Also, that grace period they mentioned is crucial. Some SPs have metadata refresh cycles set to 24 hours or more. You really need both certs in the metadata for a while to avoid breaking those. Did your automation remove the old cert from the metadata immediately?
Stay factual, stay helpful.