Skip to content
Notifications
Clear all

Complete newbie here - what's the absolute minimum code to start logging?

1 Posts
1 Users
0 Reactions
39 Views
(@laurah)
Estimable Member
Joined: 3 months ago
Posts: 62
Topic starter   [#8776]

Alright, I see this question a lot from developers who are rightfully wary of getting bogged down in tooling before they can even see if it's useful. The marketing sites make it look like you need to redesign your entire workflow. You don't.

The absolute minimum is three components: initialization, a single log call, and a finalization step. Here's the canonical example for a training loop, stripped of all ceremony.

First, install the client library. I'm assuming Python here, as that's the primary use case.

```bash
pip install wandb
```

Now, the minimal script. This assumes you have an account and have run `wandb login` at least once, which is a one-time setup.

```python
import wandb
import random # Just to generate dummy data

# 1. Initialize the run. This is where you set the project name.
run = wandb.init(project="my-minimal-test")

# Simulate a training loop
for epoch in range(5):
# Simulate your metrics
loss = random.random() / (epoch + 1)
accuracy = 1 - loss

# 2. The core logging call. Pass a dictionary.
wandb.log({"loss": loss, "accuracy": accuracy})

# 3. Finish the run. Crucial for scripts, optional in notebooks.
run.finish()
```

That's it. Execute this script. It will create a run in your project `my-minimal-test` on the W&B server (or local if you configured it that way) and log five steps of dummy data. You can then view the charts, see the system metrics it automatically captures (CPU/GPU/ memory), and inspect the run details.

Key points to understand from this snippet:
* `wandb.init()`: Creates a run context. All logged data is attached to this run. The `project` argument is the only one you really need to care about at first.
* `wandb.log()`: The workhorse. You pass it a dict of key-value pairs. The keys become the names of your metrics in the UI.
* `run.finish()`: Closes the run and ensures all data is flushed. In a notebook, this happens automatically when the cell finishes, but in a script, omitting this can leave the run "hanging."

What this doesn't show, but you'll need immediately after:
* **Configuration Logging:** Almost always, you want to log your hyperparameters or script configuration. Do this at init.
```python
run = wandb.init(project="my-test", config={"learning_rate": 0.01, "batch_size": 32})
```
* **Artifacts:** Logging files (models, datasets) is a separate system using `wandb.Artifact`. It's powerful but adds complexity. Don't worry about it for the first hour.

Common pitfalls from this minimal starting point:
1. Forgetting `run.finish()` in a script, leading to orphaned runs.
2. Not naming your projects consistently, ending up with a mess of test projects.
3. Logging at irregular intervals or with inconsistent metric names across runs, which makes comparison in the UI difficult.

Start with this exact code block. Run it. See the data appear in your dashboard. Then, incrementally add one more thing: log your config, then try adding a manual `wandb.summary` entry, then maybe log a matplotlib figure. Adding complexity piece-by-piece is the only sane way to learn this or any observability tool.


Measure twice, migrate once.


   
Quote