Skip to content
Notifications
Clear all

Just built a dashboard tracking 50+ experiments. Sharing my setup.

2 Posts
2 Users
0 Reactions
44 Views
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
Topic starter   [#7197]

Hey folks, been a lurker here for a while but finally have something worth sharing. I’ve been using W&B for about 8 months now, mostly for tracking our ML training runs. It’s been great, but I always felt I was missing a high-level view, especially with our team scaling up. We’re now running 50+ concurrent experiments across different projects, and it was getting hard to see the forest for the trees.

So, I spent last week building a custom dashboard that pulls it all together. The goal was to have a single pane of glass for our engineering leads to track progress, spot anomalies in metrics, and keep an eye on cloud costs associated with these experiments. I leaned heavily on the W&B API and some light scripting.

Here’s the core snippet I use to pull the summary data from multiple projects into a pandas DataFrame. I run this as a scheduled job that updates the dashboard source.

```python
import wandb
import pandas as pd

api = wandb.Api()
projects = ["project-a", "project-b", "vision-models"]

all_runs = []
for project in projects:
runs = api.runs(project)
for run in runs:
summary = run.summary._json_dict
config = run.config
all_runs.append({
"project": project,
"run_id": run.id,
"state": run.state,
"metric_x": summary.get("val_accuracy"),
"config_batch_size": config.get("batch_size"),
"created_at": run.created_at,
"sys_metrics_gpu_mem": summary.get("system/gpu_memory"),
})

df = pd.DataFrame(all_runs)
```

The dashboard itself is a simple Streamlit app that visualizes this DataFrame. I have sections showing:
- Experiment state distribution (how many failed/ finished/running)
- Top 5 runs per project by validation accuracy
- A scatter plot of accuracy vs. GPU memory usage (helps spot inefficient configs)
- A timeline of runs created, which is surprisingly useful for correlating with our AWS bill spikes.

The biggest win? Last Thursday, this dashboard helped us identify a group of experiments that were stuck in "running" but weren't consuming GPU. Turns out there was a bug in our training script's early stopping logic. Saved us about $400 in wasted GPU time that weekend 😅.

I’m curious how others are building their overview dashboards. Are you using the built-in W&B reports, or rolling your own like this? Any tips for tracking the actual cloud cost per experiment more granularly? I’m currently approximating from the run duration and instance type, but I’d love to hook it into our AWS Cost Explorer data.


cost first, then scale


   
Quote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

I've seen a few teams build dashboards like this, and the API approach is solid. One thing I always check in these setups is whether the audit trail from the job itself is captured. When you run it as a scheduled task, you're generating log events - the times it runs, which projects it polls, any permission errors or API rate limiting. Are you piping those logs to your SIEM or a central logging service?

It's easy for that script to become a blind spot, especially if a project name changes or an API key rotates and the dashboard just silently stops updating for a week. I'd wrap the core loop with some structured logging that gets pushed to, say, CloudWatch Logs or Datadog, so you have a history of its execution health. You can then set a simple monitor alerting if no successful completion log appears in the expected interval.


Logs don't lie.


   
ReplyQuote