No Client KYC Data in Unapproved AI Tool

Keep client KYC data out of unapproved ChatGPT, Claude, Gemini, Mistral and other AI tools. Learn what an approved service must provide

Aug 13, 2026Geoffrey Safar1 min read
No Client KYC Data in Unapproved AI Tool

No Client KYC Data in Unapproved AI Tool

A familiar logo does not make a public chatbot an approved place for a passport.

An analyst may trust ChatGPT, Claude, Gemini, Mistral or another well-known AI service enough to summarise a report. That does not mean the firm has approved it to receive a trust deed, source-of-funds explanation, ownership chart or screening result. The same warning applies to DeepSeek, Qwen, Kimi, GLM and whatever service launches next.

The problem is not the provider’s country. It is the absence of a clearly defined and approved data boundary.

KYC and AML records can combine identity documents, dates of birth, home addresses, signatures, financial accounts, family relationships, ownership percentages and unproven adverse-media allegations. A prompt containing that material is not an informal question. It is a transfer of client data to another processing environment.

The rule should therefore be universal: no live or re-identifiable client data enters an AI tool until the exact product, plan, contract, configuration and use case have passed privacy, security and compliance review.

The provider name does not define privacy

Firms often approve or reject AI at brand level. That is the wrong unit of analysis.

One provider may offer:

  • A free public chatbot for individuals.

  • A paid consumer subscription with different settings.

  • A managed business workspace.

  • An enterprise product under negotiated terms.

  • An API with separate retention rules.

  • A dedicated or self-hosted deployment.

Each can handle the same prompt differently. Training use, retention, human review, processing location, sub-processors, administrative access and deletion may change between plans.

Even the word “private” is insufficient. It may mean that another user cannot see a conversation. It does not necessarily mean the provider cannot retain the prompt, use it for service improvement, send it through a moderation system or allow authorised personnel to review it.

Likewise, an opt-out from model training solves only one question. It does not establish the firm’s lawful purpose, create a data-processing agreement, approve an international transfer or prove that information can be deleted from logs and backups.

The product and contract define the boundary, not the model name.

What current product terms show

The major providers’ own documentation demonstrates why brand-level assumptions fail.

ChatGPT: OpenAI says content submitted through individual services such as consumer ChatGPT may be used to improve model performance depending on the user’s settings. Its business offerings follow a different default: API, ChatGPT Business and ChatGPT Enterprise inputs and outputs are not used for model improvement unless the organisation opts in. A personal ChatGPT account and an approved business workspace are therefore not the same processing environment.

Claude: Anthropic gives consumer users privacy and model-improvement controls, while its commercial-product policy says Claude for Work and API inputs and outputs are not used to train its models by default. Feedback and explicit opt-in programmes can be handled differently. Approval must cover the account type, settings and feedback behaviour, not just “Claude”.

Gemini: Google’s consumer Gemini Apps privacy material describes activity settings, retention and circumstances involving product improvement or human review. Qualifying Workspace products have separate protections. Google states that Gemini for Workspace content is not used to train generative AI models outside the customer’s domain without permission, while users of consumer Gemini services can be subject to different terms. A work email address alone does not prove which regime applies.

Mistral: Mistral’s published data-training controls vary by plan. Its consumer Free and Pro input and output data may be used for training unless the user opts out, while Team and Enterprise data is not used for training. Its free API plan and paid Scale plan also differ.

These examples do not show that one provider is safe and another is unsafe. They show that a product can be appropriate under one plan and unacceptable under another.

The same test applies to any hosted service, including DeepSeek, Qwen, Kimi and GLM. If the firm cannot establish the applicable terms, data location, retention, sub-processors, access and deletion controls for the exact service being used, client data does not go in.

Build one policy for every provider and country

A practical AI policy should classify the service, not its nationality.

Unapproved public or consumer tool: No live or re-identifiable client data. Use public, synthetic or properly anonymised information only.

Approved enterprise workspace: Use only the data classes and tasks named in the approval. Apply centrally managed identities, settings, retention, logging and upload controls.

Approved API or embedded vendor: Define the precise workflow, minimise each request, control data location and sub-processors, and test the full chain from source evidence to reviewer action.

Self-hosted model: The firm controls the runtime and data path, but still needs secure model provenance, restricted network access, testing, logging, patching and human oversight.

These controls should apply to every model. A provider from the United States or Europe does not receive a shortcut. A provider from any country can be considered when the exact service meets the firm’s requirements.

For tasks such as screening analysis and false-positive resolution, the AI output also needs source evidence and named human review. A private processing environment does not make an unsupported conclusion reliable.

The operating discipline is the same when AI agents perform KYC remediation task by task. The workflow should expose only permitted case data, restrict tools and preserve the evidence behind every proposed action.

Steward follows this principle inside a purpose-built, end-to-end AML/KYC platform for investment services. Customer data is never used to train AI models. AI interactions are stateless and logged, with explainable outputs and human oversight at material decisions.