Hey everyone! 👋 I've been deep in the trenches with HashiCorp Vault lately, specifically using the Vault Agent sidecar pattern in our Kubernetes clusters to handle dynamic secrets for our PostgreSQL and MongoDB backends. It's been mostly smooth sailing, but I've hit a recurring snag with the Vault Agent cache that's causing some headaches during pod rotations and scaling events.
Here's the gist of our setup: we have the Vault Agent running as an initContainer to get the initial token, then as a sidecar with a template that renders database credentials. The config looks something like this:
```hcl
vault {
address = "https://vault.service.consul:8200"
}
auto_auth {
method "kubernetes" {
mount_path = "auth/kubernetes"
config = {
role = "app-database-role"
}
}
}
cache {
use_auto_auth_token = true
}
template {
destination = "/vault/secrets/db-creds"
contents = <<EOT
{{- with secret "database/creds/app-role" -}}
postgres://{{ .Data.username }}:{{ .Data.password }}@postgres-service:5432/appdb
{{- end }}
EOT
}
```
The problem surfaces when a pod is terminated (during a rollout, scaling in, or a node drain) and a new one spins up. Sometimes, the new pod's Vault Agent seems to get "stuck" trying to use what I think is a cached token or lease from the previous instance, leading to permission errors or delays in fetching new secrets. We've seen messages about leases being expired or "permission denied" on the database, even though the Vault role and policies are correct.
Has anyone else run into this? I'm trying to pinpoint if it's:
* An issue with the cache persistence volume (we're using an `emptyDir` volume shared between the initContainer and sidecar).
* The Kubernetes auth method not fully invalidating the old pod's service account token before the new one tries.
* Or perhaps a race condition where the template tries to render before the cache is fully ready after the auto-auth step.
I've tried adding `exit_after_auth = true` to the initContainer config and playing with `cache` persistence settings, but the intermittent nature makes it tough. Any war stories or debugging tips would be hugely appreciatedβespecially around ensuring clean, isolated cache states per pod lifecycle. Did you have to add any specific lifecycle hooks or volume cleanup tricks?
Looking forward to comparing notes! This kind of orchestration nuance is exactly why I love-and-get-frustrated-by cloud-native systems sometimes. 😅
βB
Backup first.
We ran into something similar last year! The cache can get stuck in a weird state when the agent shuts down abruptly during a pod termination. It leaves behind files that the new sidecar tries to read, but they're stale or corrupt.
Try adding `exit_after_auth = true` to your initContainer config. That forces a clean exit after getting the token, instead of lingering. Then make sure your sidecar has `cache { persist = false }`. It means more calls to Vault on startup, but it prevents those cache conflicts during rapid pod rotations. It smoothed things out for us.
Are you seeing any specific error in the sidecar logs right before it fails? Sometimes it's a permission issue with the leftover cache directory.
The config snippet cuts off, but I suspect your issue stems from the shared cache persistence directory between the initContainer and the sidecar. The `cache` block without an explicit `persist = false` will default to using `$HOME/.vault-token`, which is often a mounted volume shared across containers in the pod. If the initContainer writes a token there and then terminates uncleanly, the sidecar can inherit a stale or locked cache.
A more deterministic pattern is to use distinct `cache_file` paths for each container and disable persistence in the sidecar. For example:
```hcl
# initContainer config
cache {
use_auto_auth_token = true
cache_file = "/vault/agent-cache/init-cache"
}
# sidecar config
cache {
use_auto_auth_token = true
persist = false
}
```
This ensures the sidecar starts with a clean slate, though it does incur a fresh token fetch on each startup. Are you observing any specific error logs, like "permission denied" on the cache file or "failed to read persisted cache"?
Nullius in verba