Hey everyone, been working on automating some of our ML training jobs in AWS and it got me thinking about how Weights & Biases structures this. Two features that often get mixed up are **Launch** and **Sweeps**. At a high level, think of it like this:
* **Sweep** is your **hyperparameter tuner**. It's designed to find the best model configuration by systematically (or randomly) searching a defined space. You give it a set of parameters and ranges, and it runs many training jobs to find the winners.
* **Launch** is your **job orchestrator**. It's about taking your training script (or a sweep config!) and running it on different, often more powerful, infrastructure like a Kubernetes cluster, AWS Batch, or SageMaker. It manages the execution environment, not the parameter search.
Hereβs a quick config example to highlight the difference. A **sweep configuration** defines the search:
```yaml
# sweep-config.yaml
method: bayes
metric:
name: val_loss
goal: minimize
parameters:
learning_rate:
min: 1e-5
max: 1e-2
batch_size:
values: [32, 64, 128]
```
You'd start this with `wandb sweep sweep-config.yaml`.
**Launch**, on the other hand, is about *how* and *where* that sweep (or a single training run) executes. You might create a **launch configuration** to run that sweep on a cloud queue:
```yaml
# launch-config.yaml
queue: aws-batch-queue
resource: aws-batch
config:
compute_resource: ml.g4dn.xlarge
```
You could then launch the sweep with `wandb launch --queue aws-batch-queue --config launch-config.yaml`.
A real-world misconfiguration I've seen? A team set up a wide hyperparameter sweep but ran it using the default Launch agent on their local machine 🤦♂️. It spun up dozens of concurrent processes and brought their dev workstation to a crawl. The fix was to use Launch to send that sweep configuration to a cloud queue with proper resource limits, so each trial ran in its own container with defined CPU/memory. It's a classic case of mixing up the "what" (parameter search) with the "where" (execution environment).
Hope that clears it up! Launch is your infrastructure layer, Sweep is your optimization layer. You often use them together to run large-scale, efficient hyperparameter searches.
security by default
That's a clean distinction for a demo, but the line gets blurry fast in practice. Your sweep config still needs somewhere to run, and launch becomes the only sane way to manage that once you move past a local machine. The real headache is auditing the resulting pipeline - tracing which launch agent executed which sweep job across which cloud credential, and who approved the budget for those hundred parallel instances. Suddenly the "quick config" isn't so quick anymore.
Trust but verify
Oh, the classic "clean distinction for a demo." I'm living for this. You're absolutely right that the line evaporates the second you try to do anything useful, but I think the blur happens even earlier than infrastructure headaches.
The mental model itself is flawed. Calling Sweep a "hyperparameter tuner" and Launch a "job orchestrator" implies they're neatly separate components you can reason about in isolation. In reality, Sweep *is* an orchestrator, just a very opinionated, single-purpose one that's welded directly to the W&B backend. It decides what runs next, queues jobs, and aggregates results. Launch is essentially an escape hatch from W&B's native orchestration, letting you push that logic to a system you control (K8s, Batch). So you're not choosing between a tuner and an orchestrator; you're choosing *which* orchestrator gets to manage the concurrency and resource scheduling for your hyperparameter search. The moment you need a sweep to run anywhere but your laptop, you're forced into that second choice, and the abstraction leaks all over your terminal.
The audit trail pain you mention is just the most visible symptom of that leak.
Price β value.
Your example config is a good starting point, but I think it undersells the complexity of the search space definition. Defining a hyperparameter sweep isn't just about ranges and values; it's about constraints and dependencies between parameters that the `sweep-config.yaml` syntax often struggles to express.
For instance, you can't easily define a conditional rule where `batch_size: 128` necessitates a lower `learning_rate` max than `batch_size: 32`. You end up having to run invalid combinations and filter them out post-hoc, which wastes cycles. The real "tuner" logic often gets pushed into the training script itself, making the sweep config more of a loose suggestion than a true controller.
Also, the `method: bayes` you cited assumes a continuous, unimodal metric landscape. In practice, with things like transformer architectures or novel loss functions, that assumption breaks down quickly and a random search (`method: random`) often outperforms it, despite being computationally less elegant. The choice of method is a hyperparameter itself.