How to Use AI Safely With KYC and AML Data
A practical guide to using AI chatbots and vendors with KYC and AML data while controlling access, training, retention, transfers and audit.
How to Use AI Safely With KYC and AML Data
How to Use AI Safely With Customer KYC and AML Data
One pasted passport can turn an AI experiment into an unapproved data-processing arrangement.
The risk is easy to miss because a chatbot looks like a search box. An analyst uploads a trust deed for a summary, pastes an adverse-media article for an assessment or asks a model to identify the beneficial owners in a subscription pack. The output may be useful. The input may now sit in a service whose retention, training use, processing location and access controls the firm has never approved.
KYC and AML data can include identity documents, dates of birth, home addresses, ownership relationships, source-of-funds evidence, financial account information, screening results and unproven adverse-media allegations. Treating that material like ordinary office text is the first failure.
The safe principle is simple: control the data boundary before judging the output.
An AI tool is a data-processing environment
Not every AI service creates the same risk. Firms should separate four deployment types.
Public consumer chatbots are designed for individual use under standard online terms. They may offer privacy settings or a training opt-out, but those features do not create a data-processing agreement, define an approved transfer mechanism or give the firm effective vendor oversight. Real customer data should not enter an unapproved consumer service.
Approved enterprise chatbots can be useful for general productivity when the contract, security configuration and permitted use are clear. Approval must apply to the particular product and plan. A provider may handle data differently across its consumer chat interface, business workspace and API.
AI vendors embedded in a KYC workflow process data as part of an operational service. They require the same scrutiny as any material technology supplier, plus controls for model inputs, outputs, tools and changes.
Self-hosted models keep inference inside infrastructure controlled by the firm or its contracted environment. This can materially reduce exposure to the model publisher, but it does not remove the need for access control, secure deployment, testing, logging and human oversight.
This classification prevents a common policy mistake: either banning all AI or treating every AI product as interchangeable.
Set firm rules for employee AI chatbots
A workable employee policy should be short enough to remember and precise enough to enforce.
1. No live client data in unapproved tools
Do not paste or upload passports, application forms, trust deeds, ownership charts, screening results, case notes or customer correspondence into a public AI tool. Use synthetic cases for experimentation.
Replacing a name with initials is not automatically anonymisation. A combination of nationality, date of birth, employer, fund interest and ownership role may still identify a person. The European Data Protection Board’s opinion on AI models makes the broader point: anonymity is a case-by-case assessment, not a label applied after superficial redaction.
2. Minimise every approved prompt
Even an approved service should receive only the fields needed for the task. A document-classification model does not need a full case history. A name-screening workflow does not need a passport image if structured identifiers will do.
Separate retrieval from generation where possible. Retrieve the minimum authorised evidence, pass a bounded context to the model and discard temporary working data under a defined retention rule.
3. Keep decisions and source evidence distinct
AI can classify documents, extract fields, compare records, propose an ownership map and prepare a screening analysis. It should not silently turn those outputs into an acceptance, rejection or risk-rating decision.
Require source references, mark model-generated analysis, and preserve reviewer corrections. This is especially important when AI is used to assist screening-hit analysis and false-positive resolution, where secondary identifiers and the reason for a decision matter as much as the proposed result.
4. Make the approved route easier
Staff use public chatbots when the approved alternative is unclear or inconvenient. Give them an authenticated enterprise tool, named use cases, practical examples and a way to request a new use case. Apply single sign-on, role-based access, logging and sensible upload restrictions.
Training should explain what counts as client data, how to use synthetic examples, when a prompt needs approval and how to report an accidental disclosure.
Approve the product, contract and workflow
The vendor review cannot stop at “data is encrypted” or “prompts are not used for training”. Those are useful controls, not a complete answer.
For each AI product, establish:
Role and purpose: Is the provider a processor, controller or both for different data? What precise task is permitted?
Training and improvement: Are prompts, files, outputs, feedback, logs or derived data used to train or improve any model? Is the restriction contractual or just a setting?
Retention and deletion: What is stored, for how long, in which backups and under which exceptions? Can the firm verify deletion?
Processing location: Where are prompts, embeddings, outputs, logs and support data processed? What transfer mechanism and assessment apply?
Sub-processors: Which model, hosting, moderation, analytics and support providers can access data? How are changes notified?
Access and security: Are single sign-on, least privilege, encryption, tenant isolation, audit logs and privileged-access controls available?
Incident response: How quickly must the vendor notify the firm, and what evidence will it provide?
Model governance: Which models may process the data? How are model, prompt and workflow changes tested and approved?
Rights and complaints: Can the firm support access, correction, deletion and restriction requests across prompts, files and logs?
Exit: Can the firm export its data, evidence and audit history in a usable form, then confirm deletion?
The FCA’s data-security guidance is direct: outsourcing does not remove the firm’s responsibility for protecting customer data. Due diligence must continue after procurement through monitoring, contract controls and an exit plan.
Approval should attach to the complete combination of product, plan, region, contract, configuration and use case. A vendor approved for drafting internal policy text is not automatically approved to process live investor files.
Build controls around the complete AI workflow
The model is only one component. A production workflow may also use document storage, optical character recognition, embeddings, search, plug-ins, case-management tools and external data sources. Each component can receive or retain sensitive information.
The UK National Cyber Security Centre’s secure AI design guidance recommends due diligence on external model providers and controls over data sent outside the organisation. For KYC and AML, those controls should include:
Define the permitted task, data classes and accountable owner.
Complete the privacy, security and model-risk assessment before using real data.
Test with synthetic data, then representative historical cases in a controlled environment.
Run in shadow mode beside the existing process.
Measure unsupported conclusions, missing sources, reviewer corrections and policy exceptions.
Route material decisions to named human reviewers.
Monitor model versions, vendor changes, tool calls, access and data flows in production.
Maintain a deterministic or manual fallback and a tested exit route.
This is the same operating discipline required when AI agents perform KYC remediation task by task. The agent may prepare evidence and coordinate work, but permissions, policy and accountable decisions remain explicit.
Where Steward fits
Steward is an AI-first AML/KYC platform for investment services. Customer data is never used to train AI models. AI interactions are stateless and logged, with explainable outputs and human oversight at material decisions. That data-handling design sits inside an end-to-end workflow covering onboarding, ownership analysis, native screening, ongoing monitoring and periodic review.
Safe AI adoption is not achieved by telling employees to “be careful”. It comes from giving them an approved route, minimising what the model receives and governing the full chain from source evidence to decision.
Book a demo here to see it for yourself.
Related Insights

The AML AI Readiness Gap in North America
North American firms allocate funds to AI for AML, yet 54% use 8-10 fragmented systems. Why AI adoption isn't the same as operational readiness.

AML Red Flags for Payroll and Annex 1 Firm
Identify AML red flags in payroll and Annex 1 firms: understand sector-specific risks, connect anomalies to customer context, and build effective controls.

How to Set Up AML Controls for a UK Business
A practical operating model for building AML controls that work across payroll and Annex 1 businesses