Four different things, one sentence
"We don't train on your data" answers exactly one question: whether your content updates model weights. It says nothing about the other three fates your data can meet — logging, retention, and disclosure.
- Training: does your content improve the model? (The claim covers this.)
- Logging: is your content recorded as it passes through? (Usually yes.)
- Retention: how long do those records live? (Often 30 days to indefinitely.)
- Disclosure: who can be compelled to hand them over? (Anyone who has them.)
Questions that actually test a vendor
- How long are prompts and outputs retained, and where?
- Can retention be set to zero, in writing, for my tier?
- Which sub-processors see content, and under what terms?
- What has the vendor produced under legal process before?
The shortcut
Every one of those questions dissolves if the data never leaves your infrastructure. "No training" is a policy. "No egress" is a property. Policies change with a terms update; properties don't.
