Any ERPRegionEuropean Union

GDPR + AI on ERP data

Building GDPR-Compliant AI on Top of Your ERP

Short answer

AI systems built on ERP data process personal information by default, HR records, customer and supplier contacts, so a DPO needs a documented lawful basis, a DPIA for anything novel, and data minimisation built into the design before go-live, not reviewed after the fact. Running the model on-prem or in an EU-hosted private instance removes the international transfer question that a US-hosted API would otherwise force, which is usually the single biggest GDPR simplification available.

ERP
SAP S/4HANA, Infor LN, Oracle EBS, NetSuite
Industries
Manufacturing, Aerospace, Defense, Electronics
Written for
DPO

Every ERP system is, among other things, a large database of personal data: employee records in HR modules, contact details in customer and vendor master data, and names attached to nearly every transaction a company processes. An AI layer that reads across the ERP to answer questions or draft documents inherits that personal data exposure automatically, whether or not the specific use case was designed with personal data in mind, which is why a DPO needs a seat at the table from the start of any AI-on-ERP project, not as a late-stage sign-off.

The GDPR analysis for this kind of system runs through the same core questions as any other processing activity: what is the lawful basis under Article 6, is the processing necessary and proportionate for the stated purpose, and has a Data Protection Impact Assessment been carried out under Article 35 for processing likely to result in high risk, which a novel AI system reading broadly across HR and customer data usually is. None of this is unique to AI, but AI systems tend to expose these questions more sharply because they can, in principle, surface any personal data in the ERP in response to a question, unless the design explicitly limits that.

The international transfer question deserves its own attention because it is where AI-specific risk concentrates. Since the Schrems II ruling invalidated the EU-US Privacy Shield and put Standard Contractual Clauses under closer scrutiny for transfers to US processors subject to surveillance laws, sending ERP data to a US-hosted LLM API has become a genuinely difficult transfer to defend, even with SCCs and supplementary measures in place, because the underlying legal exposure (the US CLOUD Act and FISA 702) is a matter of law, not something a contract clause can fully mitigate.

The practical response most DPOs land on, once they walk through this, is to avoid the transfer question rather than try to fully resolve it: keep the model and the data it processes inside the EU, either on-prem or in an EU-based private cloud instance. That single architectural decision resolves the hardest part of the GDPR analysis and leaves the more manageable questions, lawful basis, minimisation, DPIA, retention, to be worked through with the same rigor as any other processing activity.

What usually gets in the way

The problems we hear most from dpo teams running SAP S/4HANA.

AI exposes personal data across the whole ERP, not just the intended use case

A question-answering assistant designed for planning or quality questions can, unless explicitly restricted, also answer questions about HR or personal contact data, because the underlying access is broader than the intended use case unless it is scoped.

Cloud LLM APIs raise the Schrems II transfer problem

Sending ERP data containing personal data to a US-hosted model API triggers the same international transfer analysis that made Schrems II a landmark case, and SCCs alone are not always viewed as sufficient supplementary measure by regulators or courts.

DPIA needed for most first AI-on-ERP deployments

Because the processing is typically novel for the organization and touches data at scale, Article 35 usually requires a DPIA before go-live, and many companies discover this only after the project is already underway.

Retention and logging create their own personal data trail

Audit logs that capture every query and response, needed for governance and AI Act evidence, themselves contain personal data and need their own retention policy and access control, or they become a second compliance problem the project created to solve the first.

Article 22 automated decision-making concerns for write-back agents

An agent that could autonomously make a decision with legal or similarly significant effect on a person, an automated credit hold, for example, risks triggering Article 22's restrictions on solely automated decision-making, which is why human review of any consequential write-back is both a GDPR and a good-practice requirement.

Where AI earns its place in SAP S/4HANA

Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.

Lawful basis and purpose documentation for each use case

Documents the specific lawful basis (typically legitimate interest or performance of a contract) for each AI-on-ERP use case before it is built, rather than relying on a single blanket justification for the whole system.

Touches: use case purpose statement, lawful basis assessment, legitimate interest balancing test where applicable

Outcome: a defensible, use-case-specific lawful basis exists for each feature, which stands up much better under scrutiny than one generic justification

Data minimisation by scope

Restricts the assistant's retrieval scope to what a given use case actually needs, excluding HR (HCM) data by default from a planning or quality use case unless there is a specific, documented reason to include it.

Touches: retrieval index scoping rules, role-based access filters

Outcome: the assistant simply cannot surface personal data it was never designed to touch, removing an entire category of risk

DPIA for the AI system

Runs a full Data Protection Impact Assessment for the AI system before go-live, covering necessity, proportionality, and risk mitigation, with the on-prem or EU-hosted architecture as the primary mitigation for transfer risk.

Touches: DPIA document, risk register, mitigation tracking

Outcome: a completed DPIA that satisfies Article 35 and gives the DPO a real basis to sign off rather than a rubber stamp

Audit log retention policy

Applies a specific, documented retention period to the AI system's query and response logs, separate from and shorter than ERP transaction retention where the two do not need to match.

Touches: audit log storage, retention schedule, access control on logs

Outcome: the governance benefit of logging is kept without creating an unbounded, unmanaged store of personal data

Human review gate for consequential write-backs

Ensures any agent action with a real effect on a person, a customer credit status, a supplier payment hold, requires human confirmation, keeping the system out of Article 22's restrictions on solely automated decisions.

Touches: approval workflow logs, write-back transaction records

Outcome: clear, logged evidence that a human made the actual decision, which both GDPR and internal audit expect to see

Data subject access request (DSAR) support

Extends the company's existing DSAR process to cover the AI system's logs, so a request for 'all data held about me' can be answered completely, including what the assistant has retrieved or generated about that person.

Touches: AI query/response logs tied to data subject identifiers

Outcome: DSAR responses stay complete as AI becomes part of the ERP landscape, rather than becoming an unaddressed gap

Third-country transfer avoidance

Confirms, as a standing architectural principle rather than a one-time check, that no part of the AI pipeline calls a non-EU service by default, including secondary functions like embeddings or telemetry that are easy to overlook.

Touches: network egress rules, service configuration audit

Outcome: the transfer question is resolved by design rather than needing to be re-argued for every new feature

Reference architecture

The architecture is built around removing the international transfer question at the source, keeping model, data, and retrieval inside the EU, then applying ordinary GDPR discipline, minimisation, purpose limitation, retention, human oversight, to what remains.

  1. 1

    ERP connectors

    Read-only connectors (OData/BAPI/RFC, ION API, REST/data entities, SuiteQL) scoped per use case so retrieval only reaches the ERP objects a given feature actually needs.

  2. 2

    Data and semantic layer

    A retrieval index that explicitly excludes HR/HCM and other sensitive personal data categories by default, included only when a specific use case and lawful basis justify it.

  3. 3

    Model serving

    An open-weight model served on EU-based infrastructure the company controls, on-prem or private cloud, with no default call to a non-EU model API.

  4. 4

    Retrieval and agents

    RAG grounds answers in current, scoped data; any agent action with a consequential effect on a person stops for human review before it takes effect.

  5. 5

    Governance and audit

    Query and response logs with their own retention policy and access control, feeding both the records of processing activities and DSAR response process.

Integration notes for your ERP team

  • ERP connectors are read-only and scoped per use case, so a quality or planning assistant cannot reach HR (HCM) or unrelated personal data by default.
  • Retrieval indexes are built to explicitly exclude sensitive personal data categories unless a specific use case and lawful basis justify their inclusion, documented at build time.
  • All model inference and retrieval run on EU-based infrastructure with no default outbound call to a non-EU service, including secondary functions like embeddings or telemetry.
  • Audit logs are stored with their own retention schedule and access control, separate from general ERP data retention, and are searchable by data subject to support DSAR responses.
  • Any write-back capable of affecting a person (a credit hold, a status change with consequences for an individual) requires a logged human approval step before it takes effect.
  • Access to the AI system mirrors existing ERP role-based access controls, so personal data visibility through the assistant never exceeds what a user could already see directly.

Deployment options

Air-gapped on-prem

Organizations with the strictest interpretation of data residency or a DPO who wants the transfer question fully closed

No data leaves the company's own network at any point, which is the cleanest possible answer to both the transfer question and general data minimisation scrutiny.

Private EU-hosted cloud

Organizations without in-house GPU capacity who still need EU-only processing

A dedicated instance on EU-based infrastructure keeps processing inside the Union, removing the Schrems II analysis without requiring the company to operate GPU hardware directly.

Hybrid

Multi-entity groups with different risk tolerance or data sensitivity across subsidiaries

The most sensitive personal data categories are restricted to on-prem processing while lower-sensitivity use cases run on a shared EU private cloud instance, under one governance framework.

Compliance and data control

How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.

GDPR Article 6 (lawful basis)

Each use case has its own documented lawful basis, most commonly legitimate interest with a recorded balancing test, or necessity for performance of a contract where applicable.

GDPR Article 35 (DPIA)

A DPIA is completed before go-live for the AI system as a whole and revisited when a materially new use case or data category is added.

GDPR Article 5 (data minimisation and purpose limitation)

Retrieval scope is restricted per use case by design, not left open and governed only by policy, so the system cannot access data it was not built to need.

GDPR Chapter V (international transfers)

EU-only processing removes the transfer analysis for the AI system itself, avoiding reliance on SCCs and supplementary measures for a US-hosted model API.

GDPR Article 22 (automated decision-making)

Any AI action with a legal or similarly significant effect on a person requires human review before it takes effect, keeping the system outside Article 22's restrictions on solely automated decisions.

How an engagement runs

Phase 1 . 2-3 weeks

Discovery

  • -Inventory of personal data categories the proposed AI system could touch across the ERP
  • -Draft lawful basis assessment for each planned use case
  • -Data flow diagram confirming EU-only processing
  • -Initial DPIA scoping document

Phase 2 . 6-8 weeks

Pilot

  • -One use case live with retrieval scope restricted to what it actually needs
  • -Completed DPIA for the pilot use case
  • -Audit log retention policy implemented and tested
  • -DSAR process extended to cover AI system logs

Phase 3 . 6-10 weeks

Production

  • -Full lawful basis documentation and DPIA sign-off from the DPO
  • -Human review gates implemented and logged for any consequential write-back
  • -Records of processing activities updated to include the AI system
  • -Staff guidance on what the assistant does and does not do with personal data

Phase 4 . Ongoing

Scale

  • -DPIA and lawful basis review extended to each new use case as it is added
  • -Periodic audit of retrieval scope to confirm no drift toward broader personal data access
  • -Annual review of retention schedules against current guidance
  • -DSAR process tested periodically against the growing AI system footprint

Questions to ask any vendor, including us

A short list that separates real SAP S/4HANA AI work from a chatbot demo.

  1. Can the vendor show us exactly which ERP objects and personal data categories a given use case's retrieval scope includes?
  2. Does any part of the pipeline call a non-EU service by default, including for embeddings, telemetry, or logging?
  3. Will the vendor support us through a DPIA with real documentation, not a generic compliance statement?
  4. How are audit logs retained, and do they have their own retention schedule separate from ERP transaction data?
  5. Is there a real, logged human review step for any action that could affect a person, or does the system act autonomously?
  6. Can the AI system's logs be included in our existing DSAR response process without custom engineering each time?
  7. What is the lawful basis the vendor recommends for each use case, and is it use-case-specific or one blanket justification?

Frequently asked questions

Does adding AI to our ERP automatically require a new DPIA?

In most cases yes, because a new AI system processing personal data at scale in a way the organization has not done before is exactly the kind of processing Article 35 targets. The DPIA does not need to be feared as a blocker; done early, alongside the architecture decisions, it usually confirms that a well-scoped, EU-hosted design is low enough risk to proceed with appropriate mitigations documented.

Is it possible to use a cloud LLM API and still be GDPR compliant?

It is possible in principle with the right SCCs and supplementary measures, but since Schrems II, transfers to US-based processors subject to laws like the CLOUD Act or FISA 702 face real scrutiny, and some European regulators have taken a strict line on specific cloud AI transfers. The simpler and increasingly common path is to avoid the question by keeping processing inside the EU, on-prem or via an EU-hosted private instance, rather than relying on contractual safeguards to fully close a legal gap.

What is the lawful basis for an AI assistant reading ERP data?

Most AI-on-ERP use cases rely on legitimate interest, the interest being operational efficiency, with a documented balancing test weighing that interest against the impact on data subjects, or on necessity for performance of a contract where the processing directly serves an existing contractual relationship. Consent is rarely practical for internal operational tooling like this, since it needs to be freely given and easily withdrawn, which does not fit an employee-facing operational system well.

Does the AI system need to exclude HR data entirely?

Not necessarily, but HR (HCM) data should be excluded by default from use cases that do not need it, planning, quality, procurement, and included only where there is a specific, documented lawful basis and purpose. This default-exclude approach is both good data minimisation practice and a straightforward thing to demonstrate to a DPO or auditor.

How does Article 22 apply to an AI agent that can take actions in the ERP?

Article 22 restricts solely automated decision-making that has a legal or similarly significant effect on a person, such as an automated credit decision. The practical mitigation is straightforward: design any agent capable of a consequential action so a human reviews and approves it before it takes effect, which keeps the decision with a person and the system firmly in an assistive rather than decision-making role.

Do audit logs of AI queries create their own GDPR problem?

They can, if left unmanaged, since a full log of every query and response is itself a rich store of personal data. The fix is treating the logs as their own processing activity with a defined purpose (governance and audit), a specific retention schedule, and access control, rather than assuming that because they support compliance, they are automatically exempt from it.

What happens with DSAR requests once AI is part of the ERP landscape?

A subject access request for 'all data you hold about me' should be understood to include what an AI system has retrieved about that person and any content it generated referencing them, not just the underlying ERP records. Extending the existing DSAR process to search AI system logs by data subject identifier, planned for from the start, keeps this answerable rather than becoming a gap discovered under a real request.

Talk it through with an engineer who knows SAP S/4HANA

Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.