Skip to content
Notifications
Clear all

Beginner question: Where does W&B store my data? Is it secure for IP?

11 Posts
11 Users
0 Reactions
24 Views
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
Topic starter   [#22841]

Hey everyone! I’ve been exploring Weights & Biases for tracking some ML experiments, and I really love the dashboard and the artifact lineage features. But as I’m starting to log more projects, a foundational question popped up that I couldn’t find a crystal-clear answer to in their docs.

Where exactly does W&B store my experiment data, model artifacts, and logs? More importantly, how secure is it for proprietary datasets or intellectual property? I’m thinking about:
- **Data at rest**: Is it encrypted? Who holds the encryption keys?
- **Data in transit**: I assume TLS, but any specifics?
- **Access controls**: Beyond project/team permissions, is there any private cloud or on-prem option if you need total control?
- **Compliance**: Do they have certifications like SOC 2, and how does data residency work?

For context, I’m coming from an iPaaS background (Zapier, Make) where webhook and API security are huge, so I’m always curious about the actual endpoint and data flow safety. For example, when I log an artifact via their Python SDK:

```python
wandb.log({"confusion_matrix": wandb.plot.confusion_matrix(...)})
```

Where does that matrix actually land, and who can access it if my project is set to “private”?

I’d love to hear from anyone who’s dug into their security whitepapers or has real-world experience with sensitive IP on W&B. Any gotchas or things you wish you knew earlier?

Thanks!
chloe


Webhooks or bust.


   
Quote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Great question! The security and data residency piece is exactly what made me dig deeper when I first started using them more seriously for client work.

You're right that TLS covers data in transit, but the data at rest encryption and key management is a bit of a black box in the standard docs. They do state everything is encrypted, but from what I've gathered, they hold the keys for the default cloud offering. That's the trade-off for convenience.

On your point about private cloud/on-prem, that's actually their big Enterprise offering. They call it "on-premises" or "private cloud" deployment, where you host the whole thing in your own VPC or data center. That's the route to go if you need full control over the keys and have strict data residency requirements, like GDPR or working with PHI. It's a totally different pricing ballgame, though.

I know they have SOC 2 Type II, but the data residency specifics get fuzzy. For example, if your team is based in the EU, does your data stay in an EU region automatically? I don't think it does by default - you'd need to confirm with their sales team or go the private cloud route. It's a good follow-up for them.


Happy testing!


   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

The API endpoint you're hitting with that `wandb.log()` call is typically their cloud service, which by default uses AWS S3 buckets in US-East-1 (Virginia). Your confusion matrix data lands as JSON metadata in their system, with any associated media files stored in the linked bucket. Access is controlled by your API key and the project's visibility settings you configure.

For IP concerns, the default cloud setup relies on their key management. If your company has a legal or compliance team, they usually require the private cloud deployment user1348 mentioned. That puts the storage backend and keys entirely within your own infrastructure, turning W&B into more of a self-hosted dashboard.

Regarding compliance, their public trust center does list SOC 2 Type II. For data residency, the standard cloud doesn't let you pick a region, which is a major driver for enterprises to go with the on-prem option.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Coming from an iPaaS background, you've hit on a key concern with API data flow. That `wandb.log()` call sends your confusion matrix to W&B's cloud endpoints, typically hosted on AWS. While TLS secures the transit, the at-rest encryption relies on their key management unless you opt for private cloud.

From my experience with webhook security in Zapier and Make, I always recommend verifying endpoint certificates and using API keys with least privilege. With W&B, you can set fine-grained project permissions, but for IP-sensitive data, consider encrypting artifacts locally before logging. It's a bit of a hack, but it adds a layer of safety.

Have you looked into their webhook integrations for triggering alerts? Sometimes, the data residency aspects come into play when webhooks cross borders.


api first


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Great point about encrypting artifacts locally before logging! I've done something similar for a client who needed an extra layer of protection for sensitive data before it left their VPC. One gotcha is that it can make artifact lineage a bit tricky, because W&B won't be able to read the contents for things like automatic version diffs or previews. You end up managing your own decryption key whenever you need to pull that artifact back for review.

Also, on your webhook comment - absolutely. If you're using W&B webhooks to trigger actions in Zapier or Make, remember the webhook payload data is flowing to *their* endpoints, not yours. That's another data egress point that could have residency implications, depending on where those services process the data. It's a small detail, but it's caught people off guard before.


Integration Ian


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The confusion matrix you log as a media artifact lands in their managed S3 buckets, which is fine for most projects. The IP question is totally valid, though.

For proprietary datasets, I'd echo the pre-encryption hack mentioned. Just know it'll break some native artifact features - no automatic previews or diffing on their dashboards, so you'll have to manage the decryption manually when you need to review. Kind of a trade-off.

If data residency is a hard requirement, their private cloud option is the real answer. It lets you pin the storage location and keep the keys. Their sales team can give you the specifics on which compliance certs apply to that deployment.


Ship fast. Learn faster.


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 7 months ago
Posts: 293
 

The trade-off with pre-encryption is real, especially for lineage. I'd add that it also impacts cost efficiency if you're using artifact references later. You might pull down an entire encrypted artifact just to find it's the wrong version, using bandwidth and time you could've avoided with a native preview.

For a true TCO comparison, you should factor in the operational overhead of managing your own keys and decryption workflow against the subscription cost of their private cloud. Sometimes the "hack" ends up being more expensive in engineering hours.


independent eye


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

You're right to be skeptical about that API call's endpoint. It defaults to hitting W&B's managed cloud, which means your confusion matrix metadata and plots land in their AWS S3 buckets, typically in US-East-1 unless you're on a regional Enterprise plan.

> how secure is it for proprietary datasets or intellectual property?

For real IP, the default cloud offering is a compliance headache waiting to happen. They hold the encryption keys. You're relying on their SOC 2 controls and hoping no employee with internal access misbehaves. If your iPaaS background has you paranoid about webhook data flows, apply that same logic here.

The only clean answer for total control is their private cloud. It's not cheap, but it moves the storage backend into your own S3 bucket with your KMS keys. Then the only data leaving your perimeter is what you explicitly whitelist.


Cloud costs are not destiny.


   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

The confusion matrix lands as a JSON metadata artifact in their cloud, with any plots stored in their managed S3. The default cloud uses their keys.

For your IP concerns, pre-encrypting before logging is a band-aid that breaks lineage features. Their private cloud deployment is the actual solution if you need to control the S3 bucket and KMS keys yourself.

Check their trust center for the current SOC 2 specifics. Data residency in the default cloud is typically US-East-1 unless you're on a regional plan.


—cp


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Agree on the private cloud being the solution for full control. The operational cost of managing it is significant though, often requiring dedicated SRE time for patching, scaling, and monitoring. It's not just a license fee swap.

You have to weigh that against the actual risk of using their keys. Their SOC 2 audit covers internal access controls, which mitigates some of the "employee misbehavior" worry. For many proprietary datasets, that's an acceptable trade-off.


Five nines? Prove it.


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Yeah, the API call lands your confusion matrix metadata in their US-East-1 S3 buckets, encrypted with their keys by default. For your data residency question, unless you're on a regional plan or private cloud, your data is likely in Virginia.

If you're coming from iPaaS, think of the default setup like sending data to a Zapier trigger endpoint - you're trusting their infra and controls. Their SOC 2 audit covers that, but the key control stays with them. The private cloud option is essentially moving the endpoint and storage backend into your own AWS account, giving you the keys and location control.



   
ReplyQuote