Any ERPUse Case

RAG + SQL + permissions

A Private LLM Grounded on Your ERP Data

Short answer

A private LLM for ERP data is a language model served on infrastructure you control, connected to your ERP through two complementary mechanisms: retrieval-augmented generation (RAG) over documents and unstructured content, and structured querying, often called text-to-SQL, over transactional tables. The model never trains on your data and never sends it to a third-party API; the harder engineering problem is making sure the model only ever sees, and only ever answers with, data the asking user is actually permitted to see in the ERP itself.

ERP
SAP S/4HANA, Infor LN, Infor SyteLine, Oracle EBS, Dynamics 365
Industries
Manufacturing, Aerospace, Defense, Electronics
Written for
CIO

CIOs evaluating generative AI on ERP data usually start with the model question (which open-weight model, how big, hosted where) but the two decisions that actually determine whether the system is trustworthy are how it retrieves data and how it enforces permissions. Get those two wrong and a large, capable model just becomes a confident way to leak or fabricate ERP data faster.

There are two distinct retrieval mechanisms, and most real deployments need both. Retrieval-augmented generation, RAG, indexes documents and unstructured content, work instructions, specs, prior tickets, so the model can find and quote relevant passages. Text-to-SQL, or more precisely, a governed structured-query layer, translates a natural-language question into a query against ERP tables, so the model can answer 'what is our current on-hand for this part' with an actual, current number rather than a plausible-sounding guess. Treating a transactional question as a RAG problem is a common and costly mistake, because a model retrieving a stale, indexed snapshot of inventory will confidently state a number that was correct last night, not now.

Permissions are the second hard problem. In most ERPs, what a user can see is already governed by role: a shop floor supervisor sees different data than a controller, a regional sales rep sees different customers than a VP. A private LLM that ignores this and answers every question from a single unrestricted service account defeats the ERP's own access control model. The correct design has the model's query layer execute with the asking user's actual ERP permissions, or an equivalent mapped role, so the answer respects the same boundaries the user already operates within.

None of this requires sending data to a public AI vendor. Open-weight models served with vLLM or Ollama on customer-owned or private-cloud GPUs can do both RAG and governed structured querying entirely within the customer's network, which is the deployment model this page assumes throughout.

What usually gets in the way

The problems we hear most from cio teams running SAP S/4HANA.

Transactional questions get answered from stale document snapshots

A pure RAG setup indexes a periodic export of ERP data, so a question about current inventory or order status returns a plausible but outdated answer instead of live data.

A single service account sees everything

The easiest way to build a demo is to give the model one broad database credential, which also means every user gets answers based on data they might not be authorized to see in the ERP itself.

The model fabricates a specific number when it is unsure

Without a governed query layer forcing it to actually run a query, a language model asked for a specific figure will sometimes generate a confident, wrong number rather than admitting it does not know.

Public AI tools create a shadow data path

Employees frustrated with slow official channels paste ERP screenshots or exports into a public chatbot to get quick answers, moving sensitive data outside the company's control without anyone deciding that should happen.

IT cannot explain how an AI answer was produced

When a generated answer is questioned, there is often no clear trail showing which records were retrieved or which query ran, making it hard to confirm the answer was right or fix it if it was wrong.

Where AI earns its place in SAP S/4HANA

Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.

Natural-language question answering over transactional data

A planner asks 'which open orders for part X are past their due date' in plain language, and the system translates that into a governed query against the ERP's order and scheduling tables, returning current, specific results.

Touches: Open Orders, Scheduling records, Item Master

Outcome: gives current, specific answers instead of a stale, pre-indexed approximation

Document and work instruction retrieval

An operator asks a question best answered by a document rather than a table, like a work instruction or a spec, and RAG retrieval finds and quotes the relevant passage, citing the source document and revision.

Touches: Work Instructions, Specifications, Document Control records

Outcome: surfaces the right passage from the right document instead of requiring a manual search

Role-scoped dashboard summaries

The same question, 'how are we doing on on-time delivery this month,' returns different underlying data to a plant manager, a regional VP, and a corporate executive, each scoped to what their ERP role already permits them to see.

Touches: Order/shipment history, role and permission tables

Outcome: one assistant serves every role correctly instead of building a separate report per audience

Hybrid RAG plus structured query for compound questions

A question like 'what does the spec say the tolerance should be, and is the last inspection result within it' requires pulling from a document (the spec) and a table (the inspection result) and combining them in one answer.

Touches: Specifications, Inspection/Quality records

Outcome: answers compound questions that neither RAG nor structured query alone could fully resolve

Confidence-aware refusal

When a question cannot be answered from data the system has access to, or when the governed query layer cannot produce a confident result, the assistant says so explicitly instead of generating a plausible-sounding guess.

Touches: Query execution layer, access control mapping

Outcome: reduces the risk of a confidently wrong answer being acted on

Shadow AI displacement

Once a sanctioned, private assistant reliably answers the questions employees were previously taking to a public chatbot, usage of unsanctioned tools with company data drops because there is a faster, safer, equally convenient option.

Touches: Usage data across whichever ERP objects the sanctioned assistant covers

Outcome: reduces the incentive to paste sensitive ERP data into an ungoverned public tool

Query audit and explainability

Every answer that used the governed structured-query layer can be traced back to the exact query executed, and every RAG answer can be traced to the specific source passages retrieved, giving IT a way to verify or correct any given answer.

Touches: Query logs, retrieval logs

Outcome: gives IT a concrete trail to investigate a questioned answer instead of a black box

Reference architecture

The architecture keeps RAG and structured querying as two distinct paths that the model chooses between (or combines) based on the question, with a permissions layer that governs both paths using the asking user's actual ERP role, not a shared service account.

  1. 1

    ERP connectors

    Read-only connections into the ERP's transactional tables (for structured queries) and its document/content stores (for RAG indexing), each scoped to specific modules rather than the whole database.

  2. 2

    Data and semantic layer

    A mapping from ERP-native table and field names to plain-language business concepts, plus a maintained index of documents for retrieval, kept current on a defined refresh schedule for structured data.

  3. 3

    Model serving

    An open-weight model (Llama, Qwen, Mistral, or similar) served with vLLM or Ollama on customer-owned or private-cloud GPUs, with no data sent to any external inference API.

  4. 4

    Retrieval and query execution

    A router that decides whether a question needs RAG retrieval, a governed structured query, or both, and executes each with the asking user's mapped ERP permissions applied.

  5. 5

    Governance and audit

    Every answer logs which query ran or which documents were retrieved, under whose permission scope, giving IT a concrete, reviewable trail for any answer.

Integration notes for your ERP team

  • Structured queries against transactional tables (inventory, orders, scheduling) should hit current data at query time, not a periodic export, since these questions have a correct answer that changes hour to hour.
  • RAG indexing of documents (specs, work instructions, prior tickets) runs on a refresh schedule appropriate to how often that content changes, typically daily or on a document-update trigger rather than real time.
  • Permissions mapping requires translating the ERP's native role and security model (e.g. SAP authorization objects, Infor ION API roles, NetSuite role-based permissions) into an equivalent scope the query layer enforces, which is usually the most detailed part of implementation.
  • Where a question needs both RAG and structured data (a spec plus an inspection result), the router combines both retrieval paths in one answer rather than forcing a second query.
  • Text-to-SQL accuracy depends heavily on how well table and field names are mapped to business terms; a poorly documented ERP schema requires more upfront semantic-layer work than a well-documented one.
  • A confidence threshold governs refusal: if the query layer cannot resolve a question to a specific, executable query, or if RAG retrieval returns nothing sufficiently relevant, the assistant says so rather than generating an answer anyway.
  • Usage and query logs are retained according to the same policy the ERP itself uses for access logs, so AI-mediated access does not create a separate, shorter-lived, or less rigorous audit trail.

Deployment options

Air-gapped on-prem

Manufacturers with export-controlled or CUI data where no data can leave the network under any circumstance

Model serving, retrieval index, and structured-query layer all run inside the customer's network boundary, with no outbound API dependency, satisfying the strictest data residency and export-control requirements.

Private sovereign cloud

Organizations wanting managed infrastructure with data residency guarantees but without operating their own GPU hardware

A single-tenant deployment in a chosen region or jurisdiction, isolated from other customers, trading infrastructure ownership for reduced operational overhead while keeping data out of a shared multi-tenant service.

Hybrid

Organizations wanting to reserve on-prem capacity for sensitive queries while handling lower-sensitivity questions more cheaply

A policy layer routes questions touching sensitive tables or documents to the on-prem or private-cloud model, while clearly non-sensitive general questions may use a separate, less restricted path, based on explicit classification rather than the user's judgment alone.

Compliance and data control

How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.

GDPR / data minimization

The structured-query layer only retrieves the specific fields needed to answer a given question, and retrieval indexes are scoped to defined document sets, rather than the model having standing access to entire tables or file shares.

ITAR / EAR (export-controlled technical data)

On-prem or private-cloud deployment with no external inference API means export-controlled data referenced in a query or document never transits infrastructure outside the customer's control.

CMMC 2.0 / NIST SP 800-171 (CUI)

Because the model's query and retrieval layers enforce the same role-based permissions as the ERP itself, CUI access through the assistant is governed and auditable the same way direct ERP access already is.

SOC 2 / internal access control audits

Query and retrieval logs, tied to the asking user's identity and role, give an auditor the same evidence trail for AI-mediated access that they would expect for direct application access.

How an engagement runs

Phase 1 . 2-3 weeks

Discovery

  • -Inventory of the highest-value question types, split between structured (transactional) and unstructured (document) needs
  • -Mapping of the ERP's existing role and permission model to what the query layer will need to enforce
  • -A data handling and hosting decision based on data sensitivity and residency requirements

Phase 2 . 6-8 weeks

Pilot

  • -Structured-query layer connected to a defined set of transactional tables
  • -RAG index built over a defined document set
  • -Permission mapping validated with test users across at least two different ERP roles

Phase 3 . ongoing

Production

  • -Assistant rolled out to the full intended user base with role-based access enforced
  • -Query and retrieval audit logging fully operational
  • -A defined process for reviewing and correcting any answer flagged as wrong

Phase 4 . ongoing

Scale

  • -Additional tables and document sets added based on observed question patterns
  • -Confidence and refusal thresholds tuned based on real usage
  • -Periodic re-validation that permission mapping still matches the current ERP role configuration after any ERP changes

Questions to ask any vendor, including us

A short list that separates real SAP S/4HANA AI work from a chatbot demo.

  1. Does the system distinguish between a question that needs a live query and one that can be answered from an indexed document, or does it treat everything as retrieval?
  2. How does the model's data access map to the ERP's own role and permission system: does a user get the same visibility through the assistant as they have logging in directly?
  3. What happens when the system cannot confidently answer: does it say so, or does it generate a plausible guess?
  4. Can we see the exact query or documents behind any given answer?
  5. Where does the model actually run, and does any query or document content leave our network during processing?
  6. How current is the data behind a transactional answer: real-time, hourly, or a nightly export?
  7. How was the permissions mapping tested, and can we validate it ourselves with real test accounts before go-live?
  8. What is the process for correcting the system when an answer turns out to be wrong?

Frequently asked questions

What is the difference between RAG and text-to-SQL for ERP data?

RAG retrieves and quotes relevant passages from documents and unstructured content, like a spec or a work instruction. Text-to-SQL, or governed structured querying, translates a question into an actual database query against transactional tables, like current inventory or order status. Most useful ERP assistants need both, chosen based on the type of question asked.

Why not just use RAG for everything, including live data?

RAG retrieves from an index, which is a snapshot, not live data. A question with a specific, current, changing answer, like on-hand inventory, needs an actual query at the moment it is asked; answering it from an indexed snapshot risks returning a plausible but outdated number.

How do you stop the AI from showing data a user is not supposed to see?

The query and retrieval layer executes with the asking user's actual ERP permissions, or an equivalent mapped role, rather than a single shared service account. This means the assistant's visibility mirrors what that user already sees when logged into the ERP directly, not a broader default.

Does a private LLM ever send our ERP data to a public AI company?

Not in the deployment model described here. The model is an open-weight model served on infrastructure the customer controls, on-prem or in a private single-tenant cloud, with no calls to a third-party inference API. Data stays inside the customer's own network boundary throughout.

What happens when the model does not know the answer?

A well-designed system says so explicitly rather than generating a confident but fabricated answer. This depends on the query layer being able to detect when it cannot resolve a question to an executable query or sufficiently relevant retrieved content, and returning a clear 'cannot answer this confidently' response instead.

How long does the permissions mapping take to get right?

It is typically the most detailed part of implementation, often several weeks depending on how complex the ERP's existing role structure is, and it is validated with real test accounts across multiple roles before go-live, not assumed correct from documentation alone.

Is this the same as just giving employees access to a public chatbot with our data pasted in?

No, and that comparison is exactly the risk this is designed to remove. A public chatbot has no awareness of ERP permissions, no live connection to transactional data, and sends whatever is pasted to a third party. A grounded, permission-aware, on-prem assistant answers from live or indexed ERP data without leaving the network and without exceeding what the user is already authorized to see.

Talk it through with an engineer who knows SAP S/4HANA

Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.