We've developed a standard clause, but it's rarely accepted verbatim. The pushback is consistent; vendors argue it removes their ability to improve service reliability. Our compromise is to allow metadata and aggregate performance metrics, but we define those terms exhaustively.
We explicitly list prohibited data types in an attachment: prompts, completions, retrieved context chunks, embedding vectors, and individual token counts. The key is linking this to a data deletion protocol. If the vendor can't use it, they shouldn't store it beyond the retention window needed for our observability.
This usually shifts the negotiation from data usage to data lifecycle, which is a more tangible discussion. The fight becomes about retention periods and auditability rather than abstract definitions of anonymization.
Data never lies.
Oh, I really like the shift to **data lifecycle**. That's a concrete angle to bring to the table that doesn't get as emotional as the training debate.
Our legal team gets stuck on the same pushback about service reliability. Defining "aggregate metrics" is the whole battle. How do you scope that in practice without letting token-level details slip back in?