All posts
gdprcomplianceopenaiself-hosted-ai

OpenAI DPA and GDPR Article 28: What to Check

A practical review of the contract, API retention, regional processing and the decisions your team still owns before sending personal data to an AI service.

The contract and the configured data flow need to describe the same system.

Stéphane Lepain··Updated ·7 min read

OpenAI's DPA governs its processing of customer personal data. It does not, by itself, make your application GDPR compliant. I would check the service and contract you actually use, the endpoints that receive data, retention and deletion, processing locations, and the responsibilities your organisation retains.

Short answer

  • GDPR Article 28 requires a written contract between you, the controller, and any processor that handles personal data on your behalf.
  • OpenAI's DPA is that contract for the services it covers. Section 1.1 names OpenAI as the processor.
  • Signing it is not enough. You still need a lawful basis, the intended account and endpoint, checked retention and processing locations, and your own records.

This is an engineering checklist, not legal advice or a compliance certification. Your DPO or qualified adviser needs to assess the lawful basis, purpose and risks for your use. Sources were checked on 9 September 2026. The public DPA I reviewed is effective 1 January 2026; your order form and applicable agreement also matter.

What Article 28 and the DPA cover

GDPR Article 28 addresses processing by a processor on behalf of a controller. The agreement needs to set out the processing and obligations, including documented instructions, confidentiality, security, subprocessors and assistance with applicable rights and obligations.

In OpenAI's public DPA, section 1.1 identifies OpenAI as a processor for the processing covered by the agreement. Section 2.1 says instructions include the agreement, order form and configuration tools. The settings your engineers choose are part of the implementation you need to review.

CheckWhere to startEvidence to keep
Which service and contracting entity apply?Applicable agreement, order form and DPA scopeDated copies and the account using them
What instructions apply?DPA 2.1 and your configurationAllowed use, endpoint list and settings
Who else processes the data?DPA 2.9–2.10 and current subprocessor listRelevant providers and change notifications
How are access, deletion and incidents handled?DPA 2.3–2.8 and 2.11Your access policy, deletion test and incident procedure
What remains your responsibility?DPA 3.1–3.3Notices, authorisations and configuration decisions
Are international transfers involved?DPA section 4 and actual data flowApplicable transfer mechanism and assessment

The DPA describes safeguards. It does not prove that your application uses the intended account, region or endpoint. Check a representative request through the entire route, including logs, file stores and integrations.

API, ChatGPT and a local workspace are different choices

An API-backed application sends requests through a developer account. A ChatGPT subscription provides access to a product with its own controls and applicable terms. A self-hosted interface may use local models or external APIs. Hosting the interface yourself does not establish where inference happens.

Do not infer API billing, retention or approval from a ChatGPT subscription. Identify each product and account explicitly. My cost comparison separates API usage, subscriptions and owned hardware for the same reason.

No training is not the same as no retention

OpenAI's API data documentation states that API data is not used to train or improve models unless you explicitly opt in. It separately describes abuse-monitoring logs and application state.

The default abuse-monitoring period is up to 30 days, with the documented legal and safety exceptions. Some features retain application state under different rules. Check files, conversations and vector stores against their own entries in the endpoint table; one retention number does not describe every feature.

Zero Data Retention and Modified Abuse Monitoring require approval and have limitations. Under ZDR, supported endpoints may change behaviour; other capabilities may still store application state. Check the exact endpoints, models and tools you intend to use. A store: false setting alone is not evidence that your account has ZDR.

EU storage and regional processing need separate checks

The API documentation distinguishes regional storage from regional processing. Where an eligible region supports processing, supported requests can perform inference there too. It is inaccurate to say that data residency always covers storage only.

Eligibility, region, project configuration, endpoint and model support still matter. OpenAI distinguishes customer content from system data; residency does not apply to all system data. Review the current limitations for tools and integrations rather than extending a regional promise to every component.

DPA section 4.1 provides for applicable safeguards when EEA or Swiss data is transferred outside those areas, including SCCs or an adequacy decision. SCCs are not a promise that a transfer never occurs. An EU contracting entity is not, on its own, proof of EU-only processing.

What changes if you self-host?

If inference, embeddings and document handling run locally, that route does not send the workload to an external model provider. But backups, monitoring, remote support and administrators may still introduce processors or transfers. Trace those paths before making an EU-only or local-only claim.

For scanned documents, I also check the OCR and document-processing path. A local language model does not settle where the other stages run.

Self-hosting leaves your organisation responsible for access controls, patching, deletion, recovery and the lawful use of data. It does not automatically remove GDPR obligations or make generated answers correct.

RouteWhat I would verify before use
Local models onlyNo unintended outbound route; operator access; local logs and backups; deletion and recovery tests
External API permittedContract and account; minimum necessary data; endpoint retention; eligible regional processing; tool destinations
Mixed local and externalExplicit allowed routes; sensitive work stays on its approved route even when a larger model might answer better

A practical review before deployment

Write down one representative task and follow its data: upload, extraction, embedding, retrieval, model request, response, logs and backup. Record which system receives each part, who can access it, how long it is retained and how deletion works. Test the configured route rather than relying only on an architecture diagram.

Agree which tasks are allowed and which must stop or be handled locally. A difficult question must not cross a prohibited boundary because an automatic router prefers a larger external model.

I explain the approach on AI evidence, human checks and limits, including a synthetic source-check example. If you need it implemented, Services sets out the scoped build and operator handover. A free scoping call is the place to establish whether that work is needed.

Questions buyers ask

Does signing the DPA make our AI application compliant?

No. Review the applicable agreement and your own use, notices, authorisations, configuration and data flow. OpenAI's DPA explicitly assigns configuration responsibilities to the customer in section 3.3. Where a reference is behind an account, I check permission before copying it into an AI workflow. Permission to read and permission to process are separate questions.

Is Zero Data Retention enabled by default?

No. It requires approval and has endpoint-specific limitations. Verify the account settings and current API data documentation.

Does self-hosting the chat interface keep everything local?

No. The model gateway, embedding service, tools and backups determine where data goes. A local interface can still call an external API.

Can EU data residency include inference?

Yes, supported regional-processing configurations can include inference. Confirm eligibility, region, models, endpoints and exclusions rather than assuming every request is covered.

Sources and revision

Revision, 9 September 2026: clarified API versus subscription scope, corrected the earlier storage-only description of residency, and removed blanket claims that self-hosting eliminates processor or transfer obligations.