SAP S/4HANA + private AI
AI for SAP S/4HANA, Running On-Prem or in Your Private Cloud
Short answer
Adding AI to an on-prem or private-cloud SAP S/4HANA system means running an open-weight model next to HANA, grounding it in the CDS views and OData services your Basis team already exposes, and keeping every prompt and completion inside your own network boundary. The pattern is read-only by default: the model answers questions and drafts documents, and a person approves anything that writes back through a BAPI or IDoc. This fits S/4HANA shops that cannot, or will not, send MM, SD, or PP data to a public LLM API.
- ERP
- SAP S/4HANA, SAP S/4HANA Private Cloud
- Industries
- Manufacturing, Aerospace, Defense, Electronics
- Written for
- CIO
If your S/4HANA landscape runs on-prem or on a private-cloud edition (including RISE with SAP private edition), most of the generative AI conversation aimed at SAP customers was not written with you in mind. Demos assume a cloud-connected tenant and a comfort level with sending queries to SAP-managed or third-party inference services that your security team has not signed off on, and may never sign off on given your industry.
That does not mean AI is off the table. S/4HANA already exposes a documented, well-understood integration surface: OData v2/v4 services, CDS views, BAPIs, and IDocs. A private, self-hosted language model can be grounded in that same surface, running on GPUs inside your data center or a dedicated private-cloud tenancy, with nothing sent to a public model API at any point.
The shift that makes this practical now is the quality of open-weight models. Llama, Qwen, Mistral, and gpt-oss-class models, served with vLLM or Ollama, are capable enough for grounded question-answering and document drafting over structured ERP data when paired with retrieval against your actual CDS views and tables, not just their training data.
This page covers the architecture, deployment options, and integration mechanics for adding AI to an on-prem or private-cloud S/4HANA system, plus the honest trade-offs against SAP's own Joule and Business AI offerings, which are covered in more depth on a separate comparison page.
What usually gets in the way
The problems we hear most from cio teams running SAP S/4HANA.
Joule assumes a cloud-connected tenant
SAP Business AI and Joule are pitched as a single story, but several capabilities are built around S/4HANA Cloud and SAP-hosted AI services, which on-prem and some private-cloud customers either do not have or do not want their data flowing through.
MD04 and ME2M are still manual
Planners and buyers dig through the stock/requirements list and PO worklists by hand because natural-language search across SD, MM, and PP tables does not exist without exporting data somewhere first.
Export-controlled data can't leave the boundary
ITAR-marked material master and BOM data cannot leave the network boundary, which rules out most SaaS copilots outright, regardless of how the vendor describes their security posture.
Years of Z-transactions are undocumented
Custom Z-tables, Y-transactions, and ABAP enhancements built up over a decade are not documented anywhere a general-purpose assistant could read them, and the people who wrote them are retiring.
Basis and security want an architecture review first
Any AI touching S/4HANA needs a review from Basis and security before rollout, and a generic 'we're SOC 2 compliant' answer from a SaaS vendor rarely satisfies that review on its own.
Where AI earns its place in SAP S/4HANA
Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.
Delivery risk triage from MD04
A planner asks which open sales orders are at risk of late delivery this week; the agent reads the stock/requirements list and open order data and returns a ranked list with reasons.
Touches: MD04 stock/requirements list, VBAK/VBAP, CDS-based ATP check
Outcome: turns a 20-30 minute manual MD04 sweep into a ranked list in seconds, still reviewed by the planner before action
PO follow-up drafting
The agent reads overdue purchase order lines and drafts a vendor follow-up email; the buyer reviews and sends it.
Touches: ME2M worklist, EKKO/EKPO/EKET, vendor master LFA1
Outcome: cuts the time to draft routine PO follow-ups from several minutes each to seconds, with the buyer still in control of what goes out
Quality notification and 8D drafting
A quality engineer opens a notification and asks for a first-pass 8D or CAPA draft based on the notification text and similar prior cases.
Touches: QM01/QMEL notification, QMFE, QMUR
Outcome: gives quality engineers a working draft instead of a blank template, cutting the time to start a CAPA investigation
Finance close variance commentary
During close review, the agent reads GL line items and drafts a plain-language explanation of a variance for the controller to check and refine.
Touches: FBL3N, CDS view I_GLAccountLineItem, FAGLFLEXT
Outcome: shortens the manual pull-and-explain cycle that usually eats a day or two during close review
Engineering change impact analysis
When an ECN is raised, the agent identifies affected BOMs, open sales orders, and open production orders touched by the change.
Touches: CS01/CS02 BOM maintenance, MAST, STPO, open AFPO production orders
Outcome: surfaces affected open orders and BOMs in minutes instead of a cross-functional meeting to work it out manually
Shop floor work instruction assistant
An operator asks a question about a production order's routing or attached work instructions instead of interrupting a supervisor.
Touches: CO01/CO03, AFVC routing operations, DMS-attached work instructions
Outcome: reduces routine interruptions to supervisors for questions the routing documentation already answers
Custom code documentation search
A new team member or consultant asks what a specific Z-transaction or user-exit does; the agent answers from indexed ABAP source and comments.
Touches: Z-tables, Y-transactions, ABAP source, code inspector export
Outcome: gives new team members and consultants a searchable map of years of custom enhancements instead of tribal knowledge
Reference architecture
The stack sits beside HANA, not inside it: a connector layer reads and writes through S/4HANA's own APIs, a retrieval layer indexes both structured data and documents, and the model itself runs on infrastructure your team controls, whether that is on-prem GPUs or a dedicated private-cloud tenancy.
- 1
ERP connectors
OData v2/v4 services and CDS views exposed through SAP Gateway, RFC/BAPI calls for actions not yet exposed as OData, and IDoc listeners for events, all running inside the customer network.
- 2
Data/semantic layer
A cached read replica or HANA calculation views for fast queries, plus a vector index for unstructured documents (work instructions, config notes), with row-level security mapped to SAP authorization objects.
- 3
Model serving
vLLM or Ollama running an open-weight model (Llama, Qwen, Mistral, or gpt-oss class) on customer-owned or private-cloud GPUs, sized to expected query volume.
- 4
Retrieval/agents
RAG grounding against the semantic layer, with an agent framework whose tools are scoped to specific OData services and BAPIs so it cannot act beyond what those calls expose.
- 5
Governance/audit
An approval workflow for any write-back action, a full query and answer log, role mapping to PFCG authorizations, and an audit trail export for security or SOX review.
Integration notes for your ERP team
- Read paths use OData v2/v4 services exposed through SAP Gateway or embedded directly in S/4HANA; where a service does not exist, a CDS view can usually be released as one without new ABAP development.
- Write-back (creating a PO, updating a notification) goes through the same BAPI or OData service a Fiori app would use, for example BAPI_PO_CREATE1, never a direct table write.
- Event triggers such as a new sales order or a goods receipt come from IDocs or SAP Event Mesh, not from polling tables on a schedule.
- Custom Z-tables and Y-transactions can be indexed for retrieval, but the agent should treat them as read-only context unless a maintained API exists for them.
- Authentication uses a dedicated communication user with SAP authorization objects scoped to the exposed services, following the same pattern as any other Fiori or integration user.
- For S/4HANA private cloud under RISE with SAP, the connector runs inside the customer's own landscape or a peered VPC; access to the SAP-managed hyperscaler tenant itself is not required.
- SAP's own Business AI features, including Joule, keep working alongside a private layer; the two are complementary rather than exclusive, and the trade-offs are covered on a separate comparison page.
Deployment options
Air-gapped on-prem
defense and ITAR-covered suppliers, DFARS-covered data, plants with no outbound internet on the ERP network
GPUs sit in your existing data center; model weights and the retrieval index never leave the network, and updates move through controlled media rather than the internet.
Private/sovereign cloud
S/4HANA private cloud (RISE private edition) or a customer-managed VPC
The model runs in a dedicated VPC alongside or beside the SAP private-cloud tenant; data stays within the contracted region and tenancy boundary, with no calls to a public model API.
Hybrid
S/4HANA on-prem today, with an eventual cloud move on the roadmap
Inference and the retrieval index run on-prem now, but the connector layer is built so it can be repointed at a cloud-hosted S/4HANA tenant later without rewriting the agent logic.
Compliance and data control
How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.
Export control (ITAR, EAR)
The model and retrieval index run inside the same network boundary as S/4HANA, so technical data used for grounding never crosses to a public model API or an unreviewed third-party service.
SAP authorization model
The agent's tool layer calls OData services and BAPIs under a service user scoped by the requesting user's existing PFCG roles, so a user cannot see through the AI anything they could not already see in SAP GUI or Fiori.
Basis and change governance
Read paths are built first and reviewed by Basis and security before any write-back tool is enabled; write actions go through an approval queue that mirrors SAP's own workflow approval patterns.
CMMC 2.0 / NIST SP 800-171
For defense suppliers, an on-prem deployment keeps CUI-adjacent ERP data inside the existing assessment boundary instead of adding a new SaaS vendor and data path to the System Security Plan.
Audit trail
Every question, retrieved record, and generated answer is logged with a timestamp and user ID, exportable alongside standard SAP change documents for internal or external audit review.
Where Netray fits
ERPray
Question-answering, dashboards, and read-only agents over S/4HANA are exactly ERPray's target: grounded answers with the underlying OData or CDS query shown, respecting existing SAP roles.
Custom build
Write-back agents for PO creation or notification updates, custom Z-table retrieval, and bespoke workflow automation typically need a scoped custom build on top of the same on-prem architecture.
How an engagement runs
Phase 1 . 2-3 weeks
Discovery
- -Inventory of OData services, CDS views, and BAPIs already exposed
- -Data flow and authorization mapping for target use cases
- -GPU and hosting sizing estimate
- -Architecture review with Basis and security
Phase 2 . 6-8 weeks
Pilot
- -Read-only RAG/agent for one or two priority use cases such as MD04 triage or PO status
- -Model serving stack deployed on customer or dedicated GPU capacity
- -Query log and accuracy review with real users
- -Go/no-go criteria for production
Phase 3 . 8-12 weeks
Production
- -Hardened deployment with access mapped to PFCG authorizations
- -Approval workflow for any write-back actions
- -Monitoring, logging, and audit export
- -Runbook handed to internal IT and Basis teams
Phase 4 . ongoing
Scale
- -Additional use cases onboarded against the same architecture
- -Model upgrades evaluated against accuracy and cost benchmarks
- -Usage and ROI reporting
- -Optional handoff to an internal team with full documentation
Questions to ask any vendor, including us
A short list that separates real SAP S/4HANA AI work from a chatbot demo.
- Does the proposed architecture ever send S/4HANA data to a public model API, even for a preview feature?
- Which SAP authorization objects and PFCG roles govern what the AI can read and write?
- Is write-back done through supported BAPIs and OData services, or direct table updates?
- Where do the model weights and vector index physically run, and who has access to that infrastructure?
- What happens to accuracy and support if the underlying open-weight model is upgraded or deprecated?
- Can we see the underlying OData or CDS query behind every AI-generated answer?
- How does this coexist with Joule and SAP Business AI features we already have licensed?
Frequently asked questions
Can we add AI to SAP S/4HANA without using SAP Business AI or Joule?
Yes. Joule and SAP Business AI are SAP's own offerings, generally aimed at cloud-connected tenants. A separate, on-prem AI layer reads CDS views and OData services directly and can run entirely inside your network, independent of whether Joule is licensed or reachable. The two are not mutually exclusive; many S/4HANA shops end up running both.
Do we need S/4HANA public cloud to get generative AI on our ERP data?
No. On-prem and private-cloud S/4HANA systems, including RISE private edition, expose the same OData services and CDS views a public cloud tenant does. A private LLM served on your own or dedicated GPUs can ground answers in that data with no dependency on which S/4HANA edition you run.
What open-weight models are realistic for SAP data today?
Models in the Llama, Qwen, Mistral, and gpt-oss families, served with vLLM or Ollama, are capable enough for grounded question-answering and document drafting over structured ERP data when paired with retrieval. The right choice depends on your GPU budget and language requirements; benchmarking against your own S/4HANA data during a pilot is more reliable than a published leaderboard.
How does the AI avoid seeing data a user isn't authorized to see?
The agent's tool layer authenticates as the requesting user, or a service user scoped identically, and calls the same OData services and BAPIs a Fiori app would, under the same SAP authorization objects and PFCG roles. If a user cannot see a plant's inventory in Fiori, the AI cannot surface it either.
Can the AI create or change records in S/4HANA, like a purchase order?
It can draft one, but write-back should go through the same BAPI or OData service a Fiori app uses, and through an approval step before it posts. Read-only deployment is the sane default for a first phase; write actions are added once the read layer has proven itself with real users.
How long does a first SAP S/4HANA AI pilot take?
A focused pilot on one or two use cases, such as MD04-based delivery risk triage or PO status drafting, typically runs six to eight weeks after a two to three week discovery phase that maps the relevant OData services and authorizations.
What does this cost compared to a SaaS copilot subscription?
The cost structure is different: GPU capacity, whether owned, colocated, or private cloud, plus integration work up front, versus a recurring per-user SaaS fee. For teams that already need on-prem infrastructure for data-sovereignty reasons, the GPU cost is often close to what a comparable SaaS seat count would run, without the data leaving the network.
Related guides
Get AI Value From SAP ECC Now, and Use It to De-Risk the Move to S/4HANA
SAP ECC 6.0 mainstream maintenance ends in 2027. Add AI value on ECC now, using its existing BAPIs and IDocs, and use it to de-risk the S/4HANA migration.
SAP BTP + self-hosted AISAP BTP and a Private LLM: Where the Generative AI Hub Fits and Where It Doesn't
Where SAP BTP's Generative AI Hub fits and where a self-hosted LLM belongs instead. CAP, Integration Suite, and an architecture decision framework.
SAP Joule + on-prem alternativeSAP Joule or a Private LLM Beside SAP: An Honest Comparison
An honest comparison of SAP Joule and Business AI against a private, self-hosted LLM beside SAP: what each covers, where they overlap, and where CIOs run both.
SAP + on-prem AI for A&DAI on SAP for Aerospace and Defense Manufacturers, Without the ITAR Exposure
AI on SAP S/4HANA or ECC for aerospace and defense manufacturers: serial genealogy, configuration control, and quoting, without technical data leaving your boundary.
On-prem AI, any ERP, A&DOn-Prem AI for ERP in Aerospace, Defense, and Electronics Manufacturing
A hub guide to on-prem AI across SAP, Infor LN, Costpoint, IFS, and Oracle EBS for aerospace, defense, and electronics manufacturers under ITAR, CMMC, and AS9100.
Agents + approval gatesAI Agents for ERP, Running On-Prem
A practical guide to on-prem AI agents for ERP: what they can safely automate, where human approval belongs, and how to design the guardrails.
Plan it with numbers
On-Prem LLM Total Cost of Ownership Calculator
Model the full multi-year cost of running LLMs on your own hardware, including GPU capex, power, cooling, support contracts, and operations staffing.
Free ToolERP AI Maturity Assessment
Benchmark how deeply AI and automation are embedded in your ERP operations, from data foundations to autonomous agents, across four maturity levels.
Free ToolSelf-Hosted LLM Hardware Estimator
Estimate the VRAM footprint, GPU count, and hardware budget required to self-host an open-weight LLM with your concurrency and context needs.
GuideEnterprise RAG Architecture: The Full 2026 Blueprint
A practitioner's blueprint for enterprise RAG in 2026: ingestion, chunking, embedding, retrieval, rerank, generation, and the eval loop that keeps it honest.
GuideOn-Prem LLM Deployment Architecture: Reference Guide
Reference architecture for on-prem LLM deployment: inference servers, GPU sizing, RAG pipelines, and security zones for regulated manufacturers.
GuideWhy On-Prem AI Is Back in 2026
On-prem AI is back in 2026 as data sovereignty, CMMC 2.0, and GPU economics shift the math. Why manufacturers are moving LLMs behind the firewall.
Talk it through with an engineer who knows SAP S/4HANA
Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.