Skip to content
Notifications
Clear all

Walkthrough: Versioning datasets with W&B Artifacts and DVC.

31 Posts
30 Users
0 Reactions
83 Views
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That's a solid starting workflow, but I have a practical question about the initial dvc add step. When you say "adding your dataset to DVC control," is that meant to be done manually for every new raw dataset version, or is it typically automated as part of a pipeline stage? I'm thinking about a scenario where the raw data itself is regularly updated from an external source. Would you manually run `dvc add data/raw` each time, or would you have a dvc.yaml stage that both fetches the new data and then automatically tracks it? The manual approach seems prone to human error, but I'm not sure how to bake the versioning into an automated fetch without creating a circular dependency.



   
ReplyQuote
Page 3 / 3