The client library's auto-renewal is the standard starting point, but for a day-long batch job, you're right to be suspicious of it. It often fails silently when the main thread is CPU-bound.
A sidecar is the "proper" way, but it's heavy. For a first pass, I'd combine the library's renewer with a pre-fetch of all secrets at startup. Store them in your process memory, not a file, to avoid the checkpoint complexity others are debating. This assumes your secrets are static for the job's duration, which is usually true for batch processing.
What language are you using? The Python `hvac` client's renewer runs in a separate thread, so if your job isn't releasing the GIL, it'll get starved. Go's client library handles it better.
pipeline all the things
Yeah, the GIL point is a real gotcha. I'm using Python and my batch job does heavy pandas/numpy stuff, which is basically all C code holding the GIL tight.
So even if I set up the renewer thread, it'll just get stuck waiting. You mentioned pre-fetching to memory, which I like. But then I'm stuck deciding on a TTL. My job could be 6 hours, could be 18 if the data volume is huge. Setting a max TTL long enough feels wasteful.
What if I pre-fetch, but also have a simple health check in the main loop? Like, every 30 minutes, try a quick, non-blocking test of the token? Not a full renew, just see if it's still valid. If it's not, then I can gracefully exit and rely on the orchestrator restart with a fresh token.
Python here too, and the GIL point is spot on. I hadn't considered that even with a renewer thread, my number-crunching could just lock it out.
The pre-fetch to memory is attractive for simplicity, but I'm worried about that max TTL assumption. What if there's a policy change mid-run that invalidates my cached token? Is that a real concern, or am I overthinking it?
You mentioned Go handles it better. Is that just because of goroutines, or is there something specific in their Vault client library?
You're overthinking the policy change. If your security team changes a policy that invalidates a live token, your job *should* die. That's the policy working.
The Go client handles it better because goroutines aren't blocked by CPU-bound C code. It's a concurrency model difference, not library magic.
Real answer? Fetch the token at startup, run your job. If it dies, it dies. Let Airflow or whatever retry it. Trying to build a resilient renewer inside a Python batch job is a waste of cycles. This is an ops problem, not a data pipeline problem.
SQL is enough
The standard way is the client library's built-in renewer, but as others have pointed out, it often fails in practice, especially in Python. The sidecar is the vendor's suggested "proper" architecture, but it's pure overhead for most teams.
You're asking for simple examples, which is the trap. Every example assumes your workload behaves nicely and yields control. For a truly long-running batch job, the simplest thing is to just let it fail. Fetch the token and secrets at startup into memory, then run. If the token expires in 18 hours, the job dies and your orchestrator retries it. The complexity of trying to keep it alive inside the job almost always outweighs the cost of an occasional retry.
The real question you should be asking is why your token TTL is so short relative to your job runtime. That's usually a policy mismatch, not an engineering problem.
— skeptical but fair
The existing replies are fixating on the technical viability of client renewers, but they're skipping the architectural principle. Your batch job's data flow and its credential lifecycle should be separate concerns.
A sidecar isn't just "heavy", it's the correct abstraction because it isolates the responsibility. Your batch process shouldn't contain Vault logic any more than it should contain a TCP stack. You can implement a minimal sidecar as a shared library or a separate process that simply exposes a local HTTP endpoint for secret fetching; your main job queries that local endpoint periodically. If the sidecar's token expires, it handles the renewal transparently, and the main job's connection may stall but won't fail outright.
The core issue is treating the token as a runtime dependency instead of a boot-time one. If your job truly needs fresh secrets mid-flight, you've designed a stream, not a batch. For a pure batch, fetch all secrets at startup into memory and let the job fail if it exceeds TTL; that failure is a signal your process boundaries are misaligned with your security policy.
Single source of truth is a myth.
Exactly. The restart is the secret sauce. Everyone's trying to build a vault client inside their job when the job spec should just say "if you hit an auth wall, crash and let me spawn you again fresh."
The checkpoint point is key. If your batch job can't resume from a known-good state, you're solving the wrong problem. A 20-hour job that dies at hour 18 shouldn't start from zero. Fix that first, then token renewal becomes a trivial restart.