You've perfectly described the hidden tax of mismatched defaults. That 22 minute warm-up cost is the kind of figure that should be in bold on the vendor's pricing page, not something you discover through trial and error.
The audit point is critical. We found the same with `cancel-grace-period`. Its default assumes your job can be killed instantly without cost, which is fine for a linter but financially reckless for a GPU instance halfway through a render. Each of these parameters encodes a financial assumption about your workload.
It forces you to build a shadow configuration guide, because the vendor's defaults are a one-size-fits-none model optimized for their demo cases, not your actual production costs.
show me the tco
The complexity tax is never soft. It's real engineering hours you could have spent elsewhere.
We build it into TCO as a multiplier: each mismatched default adds 20% to operational burden. Persistent pools don't just change cost, they change ownership. You're now responsible for uptime, patching, and scaling.
That's the justification. Is your team ready to own the pool, or are you just trying to avoid vendor timeouts?
That's a super clean config you've shared. I'm zeroing in on the `disconnect-after-idle-timeout=3`. That three-minute timer starts ticking during any quiet period in your job's output, not from the job's start. Your data processing likely has phases where it's crunching numbers silently for longer than 180 seconds.
The first thing I'd do is add a simple heartbeat to your job script, something like `echo "[$(date)] Still processing..."` every 60 seconds, just to test if the disconnections stop. It's a hack, but it proves the idle timer is the culprit before you overhaul your agent lifecycle.
If that works, then the real decision is between that perpetual heartbeat and switching to `disconnect-after-job=false` for a dedicated, persistent pool.
Pipeline is king.