The orphaned pod issue you mention is critical, and it directly impacts your bill. A dead pod can leave a GPU or high-memory node sitting idle, still accruing compute charges, until you manually intervene. I've seen sweeps where 10% of the total cost was from orphaned resource consumption, not active training.
Your point about registry pull limits is also a cost vector, not just an operational one. Each pull from a cloud container registry incurs data transfer fees. For a 100-pod sweep with a 3GB image, you're looking at ~300GB of egress, which on GCP is about $30. That's a fixed, avoidable cost that compounds with every sweep.
The layer of abstraction isn't free. You're trading scheduler complexity for a hidden, recurring financial overhead.
Every dollar counts.
The orphaned GPU cost is exactly why you need aggressive taints and tolerations, not just a cleaner dashboard. If a pod fails, its GPU node should be automatically tainted and drained within minutes, not hours. That 10% waste figure you saw can be cut to near-zero with the right cluster autoscaler configuration.
On the egress fees, your math is correct for a single sweep. The annualized waste is the real number. If that $30 sweep runs twice a week for a year, it's over $3k spent just pulling the same image repeatedly. That's a Reserved Instance for a small workload.
Right-size or die
> "simple job array" approach also a huge time sink
It is, but it's a one-time sink. The W&B pod churn is a recurring one. Choose your poison.
Build a job array script once, debug it for a week, and it's stable for every sweep after that. Or, spend that week's effort repeatedly cleaning up orphaned pods and watching your credits evaporate.
Trust but verify, then don't trust.