Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

Where Your Data Goes When an AI Application Runs

Last updated:

Four questions to answer in writing

  1. What data leaves your systems, field by field?
  2. Where does it go — which provider, which region?
  3. How long is it retained by that provider?
  4. Is it used for training under the terms you have signed?

Every AI application should have those four answers documented. Most do not, and it becomes a problem at the first security questionnaire.

Business tiers differ from consumer ones

Consumer plans and business APIs have materially different terms on training and retention. Using a personal account for company data is the most common mistake we find, and it is entirely avoidable.

Business tiers from the major providers offer no-training terms, short retention and regional processing options. Read the current terms rather than assuming.

Redact what does not need to travel

  • Replace names, account numbers and identifiers with tokens before sending
  • Restore them in the output, locally
  • Send excerpts rather than whole documents where the task allows
  • Strip metadata that carries more than you intend

Most tasks work perfectly on redacted input, and redaction removes an entire class of question from the conversation.

Retention on your side too

Prompts and outputs are usually logged for debugging, and those logs contain the same data. Decide the retention period, apply it automatically, and include the logs in any data subject request process.

A thirty-day log with automatic deletion is usually the right balance between diagnosis and exposure.

What to tell customers

If customer data is processed by a third-party model provider, your privacy notice should say so, in the same terms as any other processor.

It is a small documentation change and skipping it is a straightforward compliance gap rather than an ambiguous one.

Frequently asked questions

Can we run models on our own infrastructure?

Yes, for some workloads. Open models on your own hardware remove the transfer question entirely, at the cost of capability and operational overhead. Worth it where the data is genuinely sensitive.

Is personal data allowed at all?

Generally, with a lawful basis, an appropriate processor agreement and the usual transparency. It is ordinary data protection, not a special regime.

What about client confidentiality obligations?

Check your engagement terms. Some professional obligations restrict subprocessors, which shapes whether a hosted provider is available to you.

How do we answer a security questionnaire about this?

With the four documented answers plus your provider's terms. Having them written before you are asked turns a scramble into a paragraph.

Keep reading

Not sure what your AI tools are sending out?

It is worth knowing before someone asks. We can audit an existing setup and document the answers.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development