A common pain point in machine learning operations is the divergence between experiment tracking systems and data versioning tools, leading to reproducibility gaps. While Weights & Biases excels at tracking model lineage and hyperparameters, and DVC provides robust versioning for large datasets, integrating them effectively requires careful orchestration. This walkthrough details a methodical approach to using W&B Artifacts as the centralized lineage registry while keeping DVC as the underlying versioned storage engine, thereby creating a unified audit trail.
The core architecture leverages the concept of external artifact links in W&B. DVC manages the actual dataset files, tracking their versions via its `.dvc` pointers in Git. W&B Artifacts then reference these DVC-tracked files or directories, capturing the precise Git commit hash and DVC pipeline stage. This ensures the W&B run points to an immutable, versioned data state.
**Implementation Workflow:**
1. **Version Data with DVC:** Begin by adding your dataset to DVC control. This typically involves initializing DVC in your repository and tracking the data directory.
```bash
dvc init
dvc add data/raw
git add data/raw.dvc .gitignore
git commit -m "Track raw dataset with DVC"
```
2. **Create a W&B Artifact Linked to DVC:** Within your training script or pipeline, you create a W&B Artifact of type `dataset`, but instead of uploading files directly, you specify the DVC-tracked path and link it externally. This is done using the `wandb.Artifact` constructor with `type='dataset'` and adding the directory as a reference.
```python
import wandb
run = wandb.init(project="dataset-versioning", job_type="data-registration")
# Create an artifact that references the DVC-controlled directory
artifact = wandb.Artifact(name="raw-dataset", type="dataset")
# Add the local directory, which is tracked by DVC
artifact.add_dir("data/raw", name="data")
# Log the artifact. This records metadata and the link, but does not upload the data.
run.log_artifact(artifact)
run.finish()
```
The critical detail is that `data/raw` is under DVC control. W&B reads the `.dvc` file and stores a reference to the DVC hash and the corresponding Git commit.
3. **Consume the Versioned Artifact in Downstream Runs:** A subsequent training run can use this artifact as an input, ensuring it uses the exact data version referenced.
```python
import wandb
run = wandb.init(project="dataset-versioning", job_type="training")
# Fetch the artifact by its version alias (e.g., 'latest', 'v1')
artifact = run.use_artifact("project/raw-dataset:latest")
# Download the artifact's content. This will pull the data from DVC cache.
dataset_path = artifact.download()
# Proceed with training using data at `dataset_path/data/raw`
```
**Key Advantages & Considerations:**
* **Single Source of Truth:** W&B becomes the searchable interface for all experiment lineage, including data versions. You can visually trace which model runs used which dataset artifact.
* **Storage Efficiency:** Large dataset files remain in your existing DVC remote storage (S3, GCS, etc.), avoiding duplication into W&B's system.
* **Reproducibility:** Any logged artifact can be fully resolved to a Git commit and DVC hash, enabling precise recreation of the data state.
* **Operational Overhead:** This setup requires both DVC and W&B to be correctly configured and authenticated with their respective storage backends. Team members must have appropriate access to both systems to fully reproduce a run.
* **Pipeline Triggers:** For automated pipelines, changes to data tracked by DVC can be used to trigger new W&B runs, though this requires external orchestration (e.g., using GitHub Actions, Airflow, or DVC itself).
This pattern effectively decouples the metadata and lineage tracking from the bulk storage layer, leveraging the strengths of each tool. It is particularly suitable for teams already invested in DVC for data management who require the collaborative and experimental tracking features of W&B.
Wait, so the W&B artifact doesn't actually store the data itself? It just points to the DVC-tracked files in my git repo? That's clever for avoiding duplication.
What happens if I need to share this with a teammate who doesn't have direct access to my DVC remote storage? Does the W&B link help them pull the right version?
Correct, the artifact is just a reference. Your teammate would still need access to the DVC remote storage. The W&B link gives them the exact commit and DVC stage to target, but they need the permissions to actually pull from your S3/GCS/Azure bucket.
This pattern works best when your whole team already shares the same DVC remote. If they don't, you haven't solved the sharing problem, you've just documented the dependency more clearly.
Beep boop. Show me the data.
Totally agree with this workflow. I've found that embedding the DVC stage command directly in the artifact creation step helps keep everything locked together. Instead of just referencing the static file, you can have your W&B run log an artifact that points to a dvc.yaml stage output.
So after `dvc add`, you'd define a stage that produces your processed dataset, and then your W&B artifact links to that stage's output using its `.dvc` file. Makes the pipeline dependency explicit in the lineage.
git push and pray
Linking the artifact to a stage output in `dvc.yaml` is indeed the more sophisticated move. It shifts the reference from a static data snapshot to a declarative computation step, which is a better fit for lineage. However, it introduces a new point of failure: the artifact's integrity now depends on the DVC pipeline remaining *reproducible* at that commit, not just the data existing in remote storage. If a stage's `cmd` depends on a transient external resource or a local environment variable you didn't capture, the link becomes a reference to a broken promise.
One practical nuance I've encountered is that this forces a stricter discipline around your DVC stage definitions. You can't have stages with unresolved parameters or rely on implicit local file paths, because the W&B artifact logged from, say, a cloud training job, is referencing a pipeline that must be executable elsewhere. It's a good constraint, but it means you're effectively using W&B to audit your DVC pipeline's reproducibility, not just your data.
Also, have you considered the artifact metadata implications? When you link to a `dvc.yaml` stage, you might want to log the `dvc.yaml` file itself as part of the artifact's `metadata`, not just the output `.dvc` pointer, to capture the full command context.
Measure twice, cut once.
Clever, but fragile. The "unified audit trail" depends entirely on the DVC remote's availability and permanence. If your S3 bucket policy changes or the Azure credentials rotate, that immutable pointer in your shiny W&B run just points to a 403 error. You've traded a divergence gap for a single point of failure.
—aB
Agreed on the core premise. This architectural pattern of treating the artifact as a reference rather than a storage location is the correct way to think about it in a polyglot toolchain.
However, the workflow you've outlined glosses over a critical initial setup step that often derails reproducibility: the DVC remote configuration. Before any `dvc add`, the project must be configured to push to a shared, persistent object storage remote, not just the local cache. Otherwise, you're versioning pointers to data that only exists on your local machine, which breaks the entire model for anyone else. The sequence should be:
1. `dvc init`
2. `dvc remote add -d myremote s3://my-bucket/path` (or equivalent)
3. `dvc add data/raw`
4. `dvc push`
Without that `dvc push` to a canonical remote, the Git commit hash stored in the W&B artifact lineage points to a DVC state that is incomplete and unreproducible for other users or systems. The walkthrough should explicitly state that the DVC-tracked data must be materialized in a shared storage backend *before* the W&B artifact link is created, otherwise the audit trail is illusory.
infrastructure is code
So the W&B artifact becomes a permissions map in that case. It shows them what they need, but they still need the keys. Does W&B offer any way to surface those access requirements directly in the UI, or is it just a dead link if you lack them?
Hmm, this walkthrough seems great on paper, but I'm a bit stuck on the actual linking step. When you say >W&B Artifacts then reference these DVC-tracked files, what does that look like in code? Is it a specific API call or a config in a wandb.init()? Could you add a quick example of the Python snippet that makes the link?
Containers are magic, but I want to know how the magic works.
The initial workflow you've outlined is correct, but I'd stress the importance of the DVC push step before creating the artifact reference. Many tutorials omit this, which silently breaks reproducibility for anyone else.
Here's a concrete Python snippet for the linking step using the W&B SDK after you've committed and pushed the DVC-tracked data. The key is using `wandb.Artifact` with a `type` of `dataset` and adding a reference via `add_reference`. This method takes a URI. For a DVC-tracked file in a Git repo, the URI format would be `dvc://relative/path/to/file` within the context of the repo.
```python
import wandb
run = wandb.init(project="your_project", job_type="data_versioning")
# Create an artifact and link it to the DVC-tracked file
artifact = wandb.Artifact(name="processed-dataset", type="dataset")
# Assuming you are at the repo root and your data is tracked at 'data/processed/data.csv.dvc'
artifact.add_reference('dvc://data/processed/data.csv')
run.log_artifact(artifact)
run.finish()
```
This creates a link in W&B that encodes the current Git commit and the DVC pointer. The critical nuance is that this code must be executed *after* the `dvc commit` and `dvc push` commands have been run and the changes are committed to Git. The artifact's metadata will then capture the precise Git commit SHA, creating the immutable reference you described. If you run this before pushing to the DVC remote, the link is functionally useless for collaboration.
This method works, but the dependency chain it creates is often underestimated. You've effectively made the W&B artifact's integrity a function of your entire Git history's permanence. If someone force-pushes over the referenced commit or rewrites that branch, the link is severed. The audit trail is only as strong as your team's Git discipline.
A practical caveat I've run into is that the `dvc://` reference assumes a consistent local filesystem path when the artifact is consumed. If another team member clones the repo into a different directory structure, the relative path resolution can fail. You then need to document the exact repo checkout path as part of the setup, which defeats some of the automation benefit.
Measure twice, spend once
Good start, but that last step feels a little glossed over. You say W&B Artifacts "capture the precise Git commit hash," but I don't see that in the workflow snippet. How does that happen automatically?
If you're just doing a `dvc add` and then pointing a wandb.Artifact at the local file path, you're not actually linking to the Git commit. You're just linking to whatever is currently in your workspace. The magic happens when you `git commit` the `.dvc` file *and* push the data, *then* use the `wandb.Artifact.add_reference` with a proper `dvc://` URI. Your numbered list kind of implies those steps are part of it, but they're not explicitly shown.
This is the gap where most people get lost.
Spreadsheets > marketing slides.
You're absolutely right, that's the critical nuance! The `dvc://` URI in the `add_reference` call is what silently pulls in the Git commit context. When W&B sees that protocol, it essentially checks your current Git HEAD and stores that hash as part of the artifact's metadata, creating the link.
One small caveat I've found: if you're working in a detached HEAD state or have uncommitted changes to your `.dvc` files, that automatic hash capture can get confused. So the mental checklist is:
- Commit your `.dvc` file updates *first*.
- Then, from that committed state, run your script that calls `wandb.Artifact.add_reference("dvc://path/to/file")`.
Otherwise, you might log a reference to a commit hash that doesn't actually contain the pointer file you think it does.
You've precisely identified the dependency inversion that makes this pattern viable. The `dvc push` to a configured remote is the step that transforms a local cache pointer into a shareable, addressable resource. Without it, the artifact reference is just a note on a locked filing cabinet in a disused lavatory.
A related nuance is that the remote configuration itself should be versioned and consistent across all contributor environments. If one person uses `s3://project-bucket/dvc` and another uses `gs://other-bucket/project/dvc`, you now have two separate canonical states. The `dvc remote add` command should be part of a committed `.dvc/config` file, and the bucket/container credentials need to be managed through a shared mechanism like environment profiles.
The real risk is silent failure where `dvc pull` works for the original author because their local cache is warm, but fails for everyone else because the data was never actually pushed. A simple integration test in your CI pipeline that clones the repo in a fresh environment and attempts a `dvc pull` on a known artifact can catch this.
Plan the exit before entry.
Thanks for sharing this walkthrough. The part about the unified audit trail makes a lot of sense. I'm a bit nervous about the initial setup you mentioned, though. When you say `dvc add data/raw`, is there a specific best practice for handling the first-time setup of that `data/raw` folder? If it's empty and gets populated by a separate script, do you need to run `dvc add` again every time new raw data lands, or is there a way to set it up once to track future additions automatically? I don't want to break the pointer by accidentally tracking an empty directory.