All posts
Privacy & AIMay 18, 20268 min read

What “we don’t train on your data” actually means (and what it doesn’t)

The most reassuring sentence in AI marketing has at least four loopholes. A plain-language anatomy of the claim.

Four different things, one sentence

"We don't train on your data" answers exactly one question: whether your content updates model weights. It says nothing about the other three fates your data can meet — logging, retention, and disclosure.

  • Training: does your content improve the model? (The claim covers this.)
  • Logging: is your content recorded as it passes through? (Usually yes.)
  • Retention: how long do those records live? (Often 30 days to indefinitely.)
  • Disclosure: who can be compelled to hand them over? (Anyone who has them.)

Questions that actually test a vendor

  • How long are prompts and outputs retained, and where?
  • Can retention be set to zero, in writing, for my tier?
  • Which sub-processors see content, and under what terms?
  • What has the vendor produced under legal process before?

The shortcut

Every one of those questions dissolves if the data never leaves your infrastructure. "No training" is a policy. "No egress" is a property. Policies change with a terms update; properties don't.

Keep your documents — and your questions — yours.

ZSearch is a sovereign AI workspace that runs on your machine. Zero egress, no profiling, any model.

Download ZSearch