Hi everyone, I'm trying to understand the real-world use of ML platforms. I'm a cloud/devops junior, so my view is more from the infrastructure side.
Our small healthcare NLP team (3 data scientists) is evaluating Weights & Biases. We need strict tracking for model versions, data lineage, and metrics due to potential compliance reviews. I've been tasked with looking at the integration from a deployment angle.
Has anyone integrated W&B in a similar regulated or healthcare context? My main concerns are:
1. Where does the actual run data live? Is it easy to keep everything within our own AWS VPC?
2. The Terraform provider for W&B – is it mature enough to manage users, projects, and service accounts as code? I saw the resource list but couldn't find examples.
```hcl
# Something like this for setting up a team?
resource "wandb_team" "healthcare_nlp" {
name = "healthcare-nlp"
description = "Team for patient note NLP models"
}
```
3. How does it handle large text datasets (de-identified patient notes)? Do you log them as artifacts, and what's the cost impact?
Would love to hear about your setup, especially if you automate experiment tracking as part of a CI/CD pipeline. Thanks!