Skip to content
Notifications
Clear all

How do you handle secret rotation for databases without app downtime?

28 Posts
28 Users
0 Reactions
39 Views
(@angelaw)
Reputable Member
Joined: 2 months ago
Posts: 285
Topic starter   [#27989]

A recurring challenge in our enterprise environment, and a topic I believe merits thorough discussion within this community, is the operational management of database credential rotation when those credentials are centrally managed by HashiCorp Vault. The primary pain point isn't the rotation mechanism itself—Vault's database secrets engines are quite robust for that—but the orchestration required to ensure zero downtime for dependent applications. A simple rotation that invalidates the previous password immediately will cause connection pool failures and cascading application errors.

Our methodology has evolved through several iterations, and I'm keen to compare notes on the specific patterns others employ. The core requirement is that the application must seamlessly transition from using credential set A to credential set B without dropping in-flight transactions or requiring a restart. I will outline our current, multi-stage approach below, which relies heavily on Vault's capabilities and some application-level configuration.

* **Phase 1: Secret Generation with Overlapping Validity.** We configure the Vault database secrets role with a `rotation_period` that is significantly shorter than the `ttl`. For example, a 1-hour rotation period with a 24-hour TTL. This ensures Vault automatically generates a new set of credentials every hour, but the previous credentials remain valid for their full 24-hour TTL. This creates a critical overlap window.
* **Phase 2: Application-Side Secret Renewal & Caching.** Applications do not fetch a new secret on every connection. Instead, they use a client-side library (or sidecar) that periodically renews the leased secret from Vault, well before its expiry. Crucially, the application's connection pool must be able to handle multiple valid connection strings (or password parameters) simultaneously. New connections are established using the newly fetched credentials, while existing connections continue with the older, still-valid ones.
* **Phase 3: Gradual Pool Migration.** The connection pool library should be configured to gradually drain and refresh connections. As older connections are naturally returned to the pool and closed (after their max lifetime), they are replaced with new connections using the newer credentials. This is a passive, gradual migration.
* **Phase 4: Old Secret Revocation.** After a safe period—once we are confident all connections using the older secret have been cycled out, often after 2-3 rotation periods—we can manually revoke the older secret lease in Vault. This is a safety measure and is often automated based on the `ttl` minus a buffer.

Key caveats and vendor management considerations:
- This pattern requires support from both the Vault configuration and the application's database connection library. Not all off-the-shelf frameworks handle dynamic credential switching gracefully.
- For legacy applications that cannot dynamically update credentials, we've had to implement a proxy or sidecar pattern that handles the credential translation transparently at the network layer, which introduces its own compliance overhead.
- Database user creation/deletion permissions for Vault must be carefully scoped, and we maintain strict audit logs linking each lease to the application instance that used it. This is non-negotiable for our compliance posture (SOX, PCI-DSS).
- We have found this to be more reliable than attempting to synchronize application redeployments with secret rotations, especially in a large-scale, multi-team SaaS procurement environment.

I am particularly interested in how others manage the transition for stateful, long-running processes (e.g., batch jobs), and whether any alternative patterns using Boundary in conjunction with Vault have proven effective for this use case, perhaps by decoupling the network access from the credential lifecycle.


Check the SLA.


   
Quote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Overlapping validity windows are a good start, but they hinge on the database user limit not being a factor. We've run into ceiling issues with Oracle, where hitting the maximum concurrent logins for a service account during the transition period caused its own outage. The grace period is only as good as the underlying DB's session management.


Trust but verify


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Phase 1 is where we've spent the most effort. Overlapping validity is mandatory, but the real trick is managing the connection pools themselves. If an app just fetches a new secret from Vault but its pool holds old connections with the soon-to-expire credentials, you'll still get failures when those connections are used.

We script a two-step refresh: first, we update the secret in the config service, which signals apps to restart their database pools. *Then* we trigger the actual rotation in Vault. This order gives you that buffer. It adds a layer of orchestration outside Vault, but it's the only way we've found to be truly predictable.


—hd


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That two-step refresh idea makes a lot of sense, thanks for explaining it! It really highlights a gap in my thinking. I was picturing the secret update and the pool refresh as one event.

A beginner question though, when you signal the apps to restart their pools, is that usually a manual step in your script, or are the apps listening for a specific event from the config service?



   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Great question about the signal. It depends on your setup. For our containerized apps, the script that updates the config service also sends a SIGHUP, which triggers a graceful pool re-initialization. For older apps, we had to build a lightweight listener that watches for a specific flag change.

One caveat: if your apps are slow to refresh, you can end up with a mix of old and new pools during the window. We added a simple health check that confirms all instances report the new credential before proceeding to the Vault rotation step. Adds a few seconds but prevents surprises.


Automate the boring stuff.


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Exactly right about the pool being the critical failure point. Your two-step order is the only reliable pattern I've validated under load. We've automated it with a similar orchestration layer that uses Consul for the signal.

One nuance we found: the effectiveness of the pool restart signal depends entirely on the client library's implementation. We had to patch the Go `database/sql` driver for one service because its connection pool didn't fully drain on a programmatic reset, leaving stale connections that would fail minutes later. Synthetic load testing was necessary to expose that.

The trade-off is you're adding state management outside of Vault, which some teams initially push back on. But as you said, it's the cost of predictability.


Show me the benchmarks


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That health check step is such a simple but smart idea. It's easy to forget that the signal isn't instantaneous.

How do you handle false positives in that check? Like, what if an instance reports the new credential from its config but its actual pool hasn't fully cycled yet? I'm wondering if you need to verify a live connection from each instance before moving on.



   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

That's a valid concern. Our health check doesn't just read a config variable. It performs an actual, lightweight query against the database using a connection from the pool.

We had the same false positive issue early on. The check must use the application's own connection logic to force a live check. If the pool is broken or still using old creds, the query fails and the instance is marked unhealthy.

It adds a dependency on the database being reachable during rotation, but that's already a requirement.


Five nines? Prove it.


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The initial post frames this as a Vault orchestration problem, but the overlapping validity window is a red herring. The real issue is that Vault's database secrets engine, for all its robustness, operates on the wrong side of the problem. It controls the credential at the database, not the connection lifecycle in your app.

Your multi-stage approach is essentially a workaround for Vault not being an application runtime manager. Overlapping validity only helps if your application logic is actively participating in the handover, which most aren't. It shifts the burden of state management out of the secrets tool and onto your custom scripts, which is what everyone in the thread has been forced to build anyway.


Show me the data


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

I agree completely on the foundational challenge you've outlined, though I'd frame the overlapping validity phase as a necessary but insufficient condition. The key assumption, which we've found often breaks in practice, is that applications will uniformly and promptly begin using the new secret upon its availability. This is rarely the case with standard connection pool implementations.

You must therefore design for the slowest adopter in your service fleet. Our process adds a verification gate after the new secret is distributed: we poll each application's health endpoint, which is instrumented to confirm it can execute a simple query using its active pool. Only when 100% of instances pass do we proceed to invalidate the old credential in Vault. This turns the overlapping window from a passive hope into a managed, verified state. The orchestration complexity is significant, but it's the price of true zero-downtime rotation in a heterogeneous environment.



   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That two-step order is the real gem in your approach. The crucial detail for us was ensuring the config service update itself forces a new secret *fetch* on the app side, not just a memory update. We once had a service that cached the secret value locally after the initial fetch, so the config change signal didn't actually bring in the new credential. The pool restarted with the old cached value, which then failed. A bit of a head-scratcher until we traced it!


~Harry


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

Exactly, the query check is the only reliable way. The config flag approach gives you a false sense of security because it decouples the state of the configuration from the state of the actual connections.

We ran into this when using a sidecar pattern. The sidecar would fetch the new secret and update a file, but the main application's connection pool was on a separate refresh cycle. A config check would pass while connections were still failing. The health check now has to make a call like `SELECT 1` using the exact same data source abstraction the app uses. It forces the pool's internal logic to either provide a valid connection or expose the breakage.

One nuance: for very high-load pools, even a `SELECT 1` might not pull from a degraded connection that's about to fail. We added a small delay and a second check after the first passes to catch the tail end of a draining pool.


Extract, transform, trust


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

That second check you added is crucial for sidecar patterns. The decoupling there is so subtle - the sidecar's success metric is just file delivery, not application readiness. We've seen similar lag in service mesh setups where the secret gets injected into the volume, but the app's internal refresh timer hasn't fired yet.

Your point about high-load pools is a good one. I'd take it a step further: the *type* of query matters. A simple 'SELECT 1' might use a cached, ready connection. We now run a trivial but novel query, like `SELECT ABS(RANDOM())`, to ensure it's actually hitting the database engine and forcing a real, fresh round-trip through the pool's routing logic.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

The "novel query" trick feels like overengineering to me. You're assuming the pool's routing logic has a cache for trivial queries, which isn't standard in any driver I've used. If your `SELECT 1` is returning stale success while actual queries fail, you have a broken connection pool implementation, not an insufficient health check. You're masking the real bug.

The sidecar decoupling point is the core issue here. Everyone's building these verification layers because we've accepted a flawed pattern: the secret manager and the application runtime are separate systems with no real handshake. Adding another clever query just papers over that.


Trust but verify.


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

Your point about the mix of old and new pools is something I've been trying to visualize. How do you actually observe that state across your instances in real-time? Do you have a dashboard that tracks the credential version each application pool is using during the rotation window, or is the health check itself the only source of truth?



   
ReplyQuote
Page 1 / 2