Hey folks, saw the news about DeepSeek updating their enterprise data policy and it got me thinking. They're pushing harder on data privacy for their API and enterprise offerings, which is great, but the language around "model improvement" seems to have some carve-outs that are making a few devs on my team raise eyebrows.
Specifically, the default seems to be that *non-enterprise* API usage might still be used to improve models, unless you explicitly opt-out via a privacy parameter. For their enterprise tier, it's the opposite: data is siloed by default. This creates a two-tier system that feels... standard for the industry now, but also a bit of a trap for the unwary.
If you're building something internal or with sensitive data on the standard API, you **must** remember to set that flag. In Python, it's not just about the `api_key`, you need to pass:
```python
from openai import OpenAI
client = OpenAI(
base_url="https://api.deepseek.com",
api_key="your_key_here",
default_headers={
"X-DeepSeek-Data-Usage": "disabled" # This is the key part
}
)
```
For small scripts or prototypes, it's easy to forget. I almost missed it myself when switching from a test playground to the real API. The mental overhead is a real cost.
So, is this a red flag? Or just the new standard? Compared to other providers:
* **Claude** has similar opt-out mechanisms.
* **GPT** has more granular controls but also different defaults per product.
* **Open-source/self-hosted** models avoid this entirely, but that's a different trade-off.
I think the red flag isn't the policy itself—it's necessary for them to improve—but how easy it is to accidentally send sensitive data into the improvement pipeline. It demands discipline and good config hygiene.
What's your take? Have you adjusted your workflows or wrapper code to handle this automatically? I'm leaning towards baking the privacy header into all my client initializations, just to be safe.
-- Weave
Prompt engineering is the new debugging