Most AI features process personal data, often without anyone deciding that they should. A support chatbot receives names, phone numbers and order details. A document pipeline reads invoices with sole proprietors' tax numbers. An agent with CRM access pulls customer histories into its context. Every one of those flows is subject to data protection law — the GDPR in the EU, Ukraine's personal data law at home, and contractual obligations to clients. This guide translates the legal requirements into engineering decisions. It is not legal advice; involve your DPO or counsel for your specific case.

Where personal data flows in an LLM application

Map the flows before choosing controls:

Flow Example Typical risk
User input Customer types their phone number and complaint Sent to a third-party model provider
Retrieved context (RAG) CRM notes, tickets, contracts in the prompt Over-sharing between users; unnecessary data in prompts
Tool results Agent fetches an order with address and payment details Data in context, logs and traces
Model output Summary containing personal data Shown to the wrong user; stored indefinitely
Logs and traces Full prompts in observability tools Long retention, broad access
Memory "Remembers" facts about a user across sessions Hard to find and erase
Fine-tuning data Past conversations used for training Memorisation, inability to erase

Each row is a processing activity that needs a purpose, a legal basis and safeguards.

GDPR principles, translated into engineering

The GDPR's core principles (Article 5) map directly to design choices:

  • Purpose limitation: use data sent to the assistant for answering, not silently for training or marketing.
  • Data minimisation: send the model only what it needs. A support bot rarely needs a full customer record — pass the relevant order, not the profile with date of birth and full address.
  • Accuracy: do not let model output overwrite master data without verification.
  • Storage limitation: define retention for conversations, prompts in logs, embeddings and memory.
  • Integrity and confidentiality: encryption, access control, protection against prompt injection and data exfiltration; see AI agent security.
  • Accountability: document what you do — records of processing, DPIAs where required, vendor assessments.

Article 25 — data protection by design and by default — is essentially an instruction to make these choices in the architecture, not in a policy document.

Choosing and configuring model providers

When you send data to an API provider, they are usually your processor and you need a data processing agreement (DPA). Check, in writing:

  1. Training on your data. Business and API offerings of major providers generally do not train on customer data by default; consumer apps may. Make sure you are on the right product tier.
  2. Retention. How long are prompts and outputs stored, for abuse monitoring or otherwise? Many providers offer zero data retention or shortened retention for eligible customers.
  3. Location and transfers. Where is data processed? For transfers outside the EEA, rely on adequacy (such as the EU–US Data Privacy Framework for certified companies) or standard contractual clauses, and document a transfer assessment. Several providers offer EU data residency or regional processing through major cloud platforms.
  4. Sub-processors and security certifications (ISO 27001, SOC 2).
  5. Features that store data: file uploads, conversation state, caching, batch jobs — each may have its own retention.

When data must not leave your infrastructure at all, self-hosted models remove third-party transfers, though all other obligations remain.

Common legal bases for AI features are contract performance (the user asked the assistant for help with their order) and legitimate interests (internal productivity tools), the latter requiring a balancing test. The EDPB's Opinion 28/2024 addresses AI models specifically — including when a model trained on personal data can be considered anonymous and how legitimate interest applies to development and deployment.

Transparency obligations apply: update your privacy notice to describe AI processing, providers and purposes. Under the EU AI Act, users must also be told when they interact with an AI system; see EU AI Act guide. If an AI system makes decisions with legal or similarly significant effects on people — credit, hiring, insurance — Article 22 GDPR on automated decision-making applies, and meaningful human review is needed.

PII redaction and pseudonymisation

Redaction before sending data to the model is a strong minimisation measure when the task does not need the identifiers:

from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine

analyzer = AnalyzerEngine()      # add custom recognizers for UA formats
anonymizer = AnonymizerEngine()

def redact(text: str) -> str:
    results = analyzer.analyze(text=text, language="en",
                               entities=["PHONE_NUMBER", "EMAIL_ADDRESS", "IBAN_CODE", "PERSON"])
    return anonymizer.anonymize(text=text, analyzer_results=results).text

Microsoft Presidio and similar tools detect common identifiers. For Ukrainian data, add custom recognisers: RNOKPP (10-digit tax number), EDRPOU, passport formats, Ukrainian phone numbers (+380…), IBAN (UA + 27 digits). Name detection in Ukrainian is harder; test recall on real samples.

Pseudonymisation is often better than deletion: replace identifiers with tokens (<CUSTOMER_1>), keep the mapping server-side, and re-insert real values into the model's answer before showing it to the authorised user. The model can reason about "customer 1's order" without seeing the name.

Be realistic: redaction is never perfect, and some tasks need the data (an assistant drafting a reply to a named customer). Combine redaction with contractual and technical controls rather than relying on it alone.

RAG and agents: access control is privacy control

The most common privacy incident in AI features is not a provider leak; it is one user seeing another user's data because retrieval or tools ignored permissions:

  • Filter retrieval by the user's permissions in the query, with metadata stored on every chunk; see vector databases.
  • Execute tools with the end user's identity, taken from the session — never from model arguments; see tool design.
  • Do not index data that the assistant's users should never see.
  • Remember that embeddings derived from personal data are personal data too; erasure must reach the vector index.

Logs, traces and memory

Observability tools are a frequently forgotten data store. Apply the same rules as for production databases:

  • Separate content from metadata; keep content for a short window (for example 14–30 days).
  • Redact before export where feasible.
  • Restrict access to people who need it.
  • Include LLM tooling (tracing SaaS, evaluation datasets) in your records of processing and vendor assessments.

The tooling side is covered in LLM observability. For assistant memory, let users see and delete what is stored, and expire memories that are not used.

Fine-tuning and memorisation

Training on personal data creates harder problems. Language models can memorise and regurgitate parts of their training data, including personal information (Carlini et al., 2021). Erasing one person's data from a trained model is not practical today; retraining is the only reliable option. Therefore:

  • Prefer RAG for anything involving personal data — deletion means removing a document; see fine-tuning vs RAG.
  • If you must fine-tune on conversations, pseudonymise thoroughly, document the legal basis and keep training sets minimal.

Data subject rights

Plan how you will handle requests:

  • Access: can you export a person's conversations, stored memories and derived data?
  • Erasure: can you delete them from the database, vector index, logs, caches and backups (within backup rotation)?
  • Objection and restriction: can a user opt out of AI processing and still use the service?

Build these capabilities when you design the storage, not after the first request arrives.

Ukraine: what applies

In Ukraine, the Law on Personal Data Protection applies, with the Ombudsman as the supervisory authority, and new legislation aligned with the GDPR has been in progress as part of EU integration. Many Ukrainian IT companies also process data of EU residents and are directly subject to the GDPR through their clients' contracts or their own services. In practice, designing to GDPR standards covers Ukrainian requirements and client expectations at the same time.

A privacy checklist for AI features

  • Data flows mapped: inputs, retrieval, tools, outputs, logs, memory
  • Provider on a business/API tier with DPA, no training on your data, known retention
  • Transfer mechanism documented; regional processing where needed
  • Only necessary data sent to the model; redaction or pseudonymisation where feasible
  • Retrieval and tools enforce the end user's permissions
  • Retention set for conversations, logs, traces, embeddings and memory
  • Privacy notice updated; users informed they are talking to AI
  • Access, erasure and opt-out processes cover all AI data stores
  • DPIA done for high-risk processing (large-scale, sensitive data, significant decisions)

FAQ

Can we send customer data to an LLM API under GDPR? Yes, with a legal basis, a DPA with the provider, appropriate transfer safeguards, minimisation and transparency. The question is how, not whether.

Is a self-hosted model automatically compliant? No. It avoids third-party transfers, but purpose limitation, retention, security and data subject rights still apply.

Are embeddings anonymous? No. Embeddings derived from personal data can often be linked back to the source text and should be treated as personal data.

Do we need a DPIA for a chatbot? Often not for a simple FAQ bot; often yes when processing sensitive data, monitoring employees, operating at large scale or supporting significant decisions. Ask your DPO.

Sources

  1. Regulation (EU) 2016/679 (GDPR).
  2. EDPB (2024). Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models.
  3. Verkhovna Rada of Ukraine. Law of Ukraine "On Personal Data Protection".
  4. Carlini et al. (2021). Extracting Training Data from Large Language Models.
  5. Microsoft. Presidio.
  6. OWASP. Top 10 for LLM Applications 2025: Sensitive Information Disclosure.