Skip to content
Notifications
Clear all

How do I handle Vault token renewal for long-running batch jobs?

1 Posts
1 Users
0 Reactions
0 Views
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
Topic starter   [#29398]

Alright, let's cut through the usual "just use the Kubernetes auth method" hand-waving. Most of us still have legacy batch jobs that run for hours, sometimes days, on good old-fashioned VMs or even bare metal. These aren't your precious 30-second Lambda functions. The Vault token TTL is a ticking time bomb, and the official answer is usually "use a renewable token with a long TTL and renew it." Fantastic. How, precisely, do you architect that renewal in a way that doesn't turn into a brittle cron job mess or a secret sprawl disaster?

I'm talking about a job that pulls a database credential from Vault at startup, then needs to keep that credential alive, or maybe fetch new ones for different stages. The token itself has a 1-hour TTL, but the job runs for 8 hours. The naive approach is to background a script that runs `vault token renew` every 55 minutes. This fails in spectacular ways:
* The renewal script dies silently.
* The job gets SIGTERM and needs to clean up, but the renewal loop is detached.
* You now have to manage the renewal token's secret... for the token that's supposed to be your secret. Meta-secret hell.

I've cobbled together a pattern using a supervisor process and token leasing, but it feels like duct tape. What I'm currently doing in a Python job wrapper (simplified):

```python
import hvac
import threading
import time
from signal import signal, SIGTERM

class VaultTokenRenewer:
def __init__(self, client, token, interval_sec=3000):
self.client = client
self.token = token
self.interval = interval_sec
self.stop_event = threading.Event()
self.thread = threading.Thread(target=self._renew_loop)

def _renew_loop(self):
while not self.stop_event.is_set():
time.sleep(self.interval)
try:
self.client.auth.token.renew_self()
except Exception as e:
# Log and maybe trigger a graceful shutdown
pass

def start(self):
self.thread.start()

def stop(self):
self.stop_event.set()
self.thread.join()

# In main job setup
client = hvac.Client(url=VAULT_ADDR, token=INITIAL_TOKEN)
renewer = VaultTokenRenewer(client, client.token)
renewer.start()

def handle_sigterm(signum, frame):
renewer.stop()
# ... other cleanup
sys.exit(0)

signal(SIGTERM, handle_sigterm)

# ... proceed with actual batch job logic
```

This works, but now I'm in the business of writing a mini-orchestrator for every job type. I've also looked at wrapping the job in a systemd unit with `ExecReload` configured to renew the token, but that's equally messy.

So, before I sink another week into building a "token lifecycle sidecar," I want to hear from other teams that have hit this wall. Are you using the agent as a sidecar? Writing custom lease managers? Or have you just given up and set the default policy TTL to 8 hours (don't do this, by the way) and called it a day? Concrete configs or code snippets are worth a thousand docs links.



   
Quote