RAG + SQL + permissions
A Private LLM Grounded on Your ERP Data
Short answer
A private LLM for ERP data is a language model served on infrastructure you control, connected to your ERP through two complementary mechanisms: retrieval-augmented generation (RAG) over documents and unstructured content, and structured querying, often called text-to-SQL, over transactional tables. The model never trains on your data and never sends it to a third-party API; the harder engineering problem is making sure the model only ever sees, and only ever answers with, data the asking user is actually permitted to see in the ERP itself.
- ERP
- SAP S/4HANA, Infor LN, Infor SyteLine, Oracle EBS, Dynamics 365
- Industries
- Manufacturing, Aerospace, Defense, Electronics
- Written for
- CIO
CIOs evaluating generative AI on ERP data usually start with the model question (which open-weight model, how big, hosted where) but the two decisions that actually determine whether the system is trustworthy are how it retrieves data and how it enforces permissions. Get those two wrong and a large, capable model just becomes a confident way to leak or fabricate ERP data faster.
There are two distinct retrieval mechanisms, and most real deployments need both. Retrieval-augmented generation, RAG, indexes documents and unstructured content, work instructions, specs, prior tickets, so the model can find and quote relevant passages. Text-to-SQL, or more precisely, a governed structured-query layer, translates a natural-language question into a query against ERP tables, so the model can answer 'what is our current on-hand for this part' with an actual, current number rather than a plausible-sounding guess. Treating a transactional question as a RAG problem is a common and costly mistake, because a model retrieving a stale, indexed snapshot of inventory will confidently state a number that was correct last night, not now.
Permissions are the second hard problem. In most ERPs, what a user can see is already governed by role: a shop floor supervisor sees different data than a controller, a regional sales rep sees different customers than a VP. A private LLM that ignores this and answers every question from a single unrestricted service account defeats the ERP's own access control model. The correct design has the model's query layer execute with the asking user's actual ERP permissions, or an equivalent mapped role, so the answer respects the same boundaries the user already operates within.
None of this requires sending data to a public AI vendor. Open-weight models served with vLLM or Ollama on customer-owned or private-cloud GPUs can do both RAG and governed structured querying entirely within the customer's network, which is the deployment model this page assumes throughout.
What usually gets in the way
The problems we hear most from cio teams running SAP S/4HANA.
Transactional questions get answered from stale document snapshots
A pure RAG setup indexes a periodic export of ERP data, so a question about current inventory or order status returns a plausible but outdated answer instead of live data.
A single service account sees everything
The easiest way to build a demo is to give the model one broad database credential, which also means every user gets answers based on data they might not be authorized to see in the ERP itself.
The model fabricates a specific number when it is unsure
Without a governed query layer forcing it to actually run a query, a language model asked for a specific figure will sometimes generate a confident, wrong number rather than admitting it does not know.
Public AI tools create a shadow data path
Employees frustrated with slow official channels paste ERP screenshots or exports into a public chatbot to get quick answers, moving sensitive data outside the company's control without anyone deciding that should happen.
IT cannot explain how an AI answer was produced
When a generated answer is questioned, there is often no clear trail showing which records were retrieved or which query ran, making it hard to confirm the answer was right or fix it if it was wrong.
Where AI earns its place in SAP S/4HANA
Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.
Natural-language question answering over transactional data
A planner asks 'which open orders for part X are past their due date' in plain language, and the system translates that into a governed query against the ERP's order and scheduling tables, returning current, specific results.
Touches: Open Orders, Scheduling records, Item Master
Outcome: gives current, specific answers instead of a stale, pre-indexed approximation
Document and work instruction retrieval
An operator asks a question best answered by a document rather than a table, like a work instruction or a spec, and RAG retrieval finds and quotes the relevant passage, citing the source document and revision.
Touches: Work Instructions, Specifications, Document Control records
Outcome: surfaces the right passage from the right document instead of requiring a manual search
Role-scoped dashboard summaries
The same question, 'how are we doing on on-time delivery this month,' returns different underlying data to a plant manager, a regional VP, and a corporate executive, each scoped to what their ERP role already permits them to see.
Touches: Order/shipment history, role and permission tables
Outcome: one assistant serves every role correctly instead of building a separate report per audience
Hybrid RAG plus structured query for compound questions
A question like 'what does the spec say the tolerance should be, and is the last inspection result within it' requires pulling from a document (the spec) and a table (the inspection result) and combining them in one answer.
Touches: Specifications, Inspection/Quality records
Outcome: answers compound questions that neither RAG nor structured query alone could fully resolve
Confidence-aware refusal
When a question cannot be answered from data the system has access to, or when the governed query layer cannot produce a confident result, the assistant says so explicitly instead of generating a plausible-sounding guess.
Touches: Query execution layer, access control mapping
Outcome: reduces the risk of a confidently wrong answer being acted on
Shadow AI displacement
Once a sanctioned, private assistant reliably answers the questions employees were previously taking to a public chatbot, usage of unsanctioned tools with company data drops because there is a faster, safer, equally convenient option.
Touches: Usage data across whichever ERP objects the sanctioned assistant covers
Outcome: reduces the incentive to paste sensitive ERP data into an ungoverned public tool
Query audit and explainability
Every answer that used the governed structured-query layer can be traced back to the exact query executed, and every RAG answer can be traced to the specific source passages retrieved, giving IT a way to verify or correct any given answer.
Touches: Query logs, retrieval logs
Outcome: gives IT a concrete trail to investigate a questioned answer instead of a black box
Reference architecture
The architecture keeps RAG and structured querying as two distinct paths that the model chooses between (or combines) based on the question, with a permissions layer that governs both paths using the asking user's actual ERP role, not a shared service account.
- 1
ERP connectors
Read-only connections into the ERP's transactional tables (for structured queries) and its document/content stores (for RAG indexing), each scoped to specific modules rather than the whole database.
- 2
Data and semantic layer
A mapping from ERP-native table and field names to plain-language business concepts, plus a maintained index of documents for retrieval, kept current on a defined refresh schedule for structured data.
- 3
Model serving
An open-weight model (Llama, Qwen, Mistral, or similar) served with vLLM or Ollama on customer-owned or private-cloud GPUs, with no data sent to any external inference API.
- 4
Retrieval and query execution
A router that decides whether a question needs RAG retrieval, a governed structured query, or both, and executes each with the asking user's mapped ERP permissions applied.
- 5
Governance and audit
Every answer logs which query ran or which documents were retrieved, under whose permission scope, giving IT a concrete, reviewable trail for any answer.
Integration notes for your ERP team
- Structured queries against transactional tables (inventory, orders, scheduling) should hit current data at query time, not a periodic export, since these questions have a correct answer that changes hour to hour.
- RAG indexing of documents (specs, work instructions, prior tickets) runs on a refresh schedule appropriate to how often that content changes, typically daily or on a document-update trigger rather than real time.
- Permissions mapping requires translating the ERP's native role and security model (e.g. SAP authorization objects, Infor ION API roles, NetSuite role-based permissions) into an equivalent scope the query layer enforces, which is usually the most detailed part of implementation.
- Where a question needs both RAG and structured data (a spec plus an inspection result), the router combines both retrieval paths in one answer rather than forcing a second query.
- Text-to-SQL accuracy depends heavily on how well table and field names are mapped to business terms; a poorly documented ERP schema requires more upfront semantic-layer work than a well-documented one.
- A confidence threshold governs refusal: if the query layer cannot resolve a question to a specific, executable query, or if RAG retrieval returns nothing sufficiently relevant, the assistant says so rather than generating an answer anyway.
- Usage and query logs are retained according to the same policy the ERP itself uses for access logs, so AI-mediated access does not create a separate, shorter-lived, or less rigorous audit trail.
Deployment options
Air-gapped on-prem
Manufacturers with export-controlled or CUI data where no data can leave the network under any circumstance
Model serving, retrieval index, and structured-query layer all run inside the customer's network boundary, with no outbound API dependency, satisfying the strictest data residency and export-control requirements.
Private sovereign cloud
Organizations wanting managed infrastructure with data residency guarantees but without operating their own GPU hardware
A single-tenant deployment in a chosen region or jurisdiction, isolated from other customers, trading infrastructure ownership for reduced operational overhead while keeping data out of a shared multi-tenant service.
Hybrid
Organizations wanting to reserve on-prem capacity for sensitive queries while handling lower-sensitivity questions more cheaply
A policy layer routes questions touching sensitive tables or documents to the on-prem or private-cloud model, while clearly non-sensitive general questions may use a separate, less restricted path, based on explicit classification rather than the user's judgment alone.
Compliance and data control
How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.
GDPR / data minimization
The structured-query layer only retrieves the specific fields needed to answer a given question, and retrieval indexes are scoped to defined document sets, rather than the model having standing access to entire tables or file shares.
ITAR / EAR (export-controlled technical data)
On-prem or private-cloud deployment with no external inference API means export-controlled data referenced in a query or document never transits infrastructure outside the customer's control.
CMMC 2.0 / NIST SP 800-171 (CUI)
Because the model's query and retrieval layers enforce the same role-based permissions as the ERP itself, CUI access through the assistant is governed and auditable the same way direct ERP access already is.
SOC 2 / internal access control audits
Query and retrieval logs, tied to the asking user's identity and role, give an auditor the same evidence trail for AI-mediated access that they would expect for direct application access.
Where Netray fits
ERPray
ERPray is built specifically for grounded natural-language question answering over live ERP data, read-only by default, showing the underlying query, and respecting ERP roles, which is exactly the RAG-plus-structured-query-plus-permissions design described here, for NetSuite, Infor SyteLine, M3, and LN today.
DataRay
Where the question spans data sources beyond the ERP itself, file shares, databases, or document stores, DataRay extends the same on-prem, permission-aware approach to that broader mixed-source environment.
Custom build
For ERPs outside the current connector list, or where the permissions-mapping work is unusually complex, this is delivered as a scoped custom implementation using the same architecture.
How an engagement runs
Phase 1 . 2-3 weeks
Discovery
- -Inventory of the highest-value question types, split between structured (transactional) and unstructured (document) needs
- -Mapping of the ERP's existing role and permission model to what the query layer will need to enforce
- -A data handling and hosting decision based on data sensitivity and residency requirements
Phase 2 . 6-8 weeks
Pilot
- -Structured-query layer connected to a defined set of transactional tables
- -RAG index built over a defined document set
- -Permission mapping validated with test users across at least two different ERP roles
Phase 3 . ongoing
Production
- -Assistant rolled out to the full intended user base with role-based access enforced
- -Query and retrieval audit logging fully operational
- -A defined process for reviewing and correcting any answer flagged as wrong
Phase 4 . ongoing
Scale
- -Additional tables and document sets added based on observed question patterns
- -Confidence and refusal thresholds tuned based on real usage
- -Periodic re-validation that permission mapping still matches the current ERP role configuration after any ERP changes
Questions to ask any vendor, including us
A short list that separates real SAP S/4HANA AI work from a chatbot demo.
- Does the system distinguish between a question that needs a live query and one that can be answered from an indexed document, or does it treat everything as retrieval?
- How does the model's data access map to the ERP's own role and permission system: does a user get the same visibility through the assistant as they have logging in directly?
- What happens when the system cannot confidently answer: does it say so, or does it generate a plausible guess?
- Can we see the exact query or documents behind any given answer?
- Where does the model actually run, and does any query or document content leave our network during processing?
- How current is the data behind a transactional answer: real-time, hourly, or a nightly export?
- How was the permissions mapping tested, and can we validate it ourselves with real test accounts before go-live?
- What is the process for correcting the system when an answer turns out to be wrong?
Frequently asked questions
What is the difference between RAG and text-to-SQL for ERP data?
RAG retrieves and quotes relevant passages from documents and unstructured content, like a spec or a work instruction. Text-to-SQL, or governed structured querying, translates a question into an actual database query against transactional tables, like current inventory or order status. Most useful ERP assistants need both, chosen based on the type of question asked.
Why not just use RAG for everything, including live data?
RAG retrieves from an index, which is a snapshot, not live data. A question with a specific, current, changing answer, like on-hand inventory, needs an actual query at the moment it is asked; answering it from an indexed snapshot risks returning a plausible but outdated number.
How do you stop the AI from showing data a user is not supposed to see?
The query and retrieval layer executes with the asking user's actual ERP permissions, or an equivalent mapped role, rather than a single shared service account. This means the assistant's visibility mirrors what that user already sees when logged into the ERP directly, not a broader default.
Does a private LLM ever send our ERP data to a public AI company?
Not in the deployment model described here. The model is an open-weight model served on infrastructure the customer controls, on-prem or in a private single-tenant cloud, with no calls to a third-party inference API. Data stays inside the customer's own network boundary throughout.
What happens when the model does not know the answer?
A well-designed system says so explicitly rather than generating a confident but fabricated answer. This depends on the query layer being able to detect when it cannot resolve a question to an executable query or sufficiently relevant retrieved content, and returning a clear 'cannot answer this confidently' response instead.
How long does the permissions mapping take to get right?
It is typically the most detailed part of implementation, often several weeks depending on how complex the ERP's existing role structure is, and it is validated with real test accounts across multiple roles before go-live, not assumed correct from documentation alone.
Is this the same as just giving employees access to a public chatbot with our data pasted in?
No, and that comparison is exactly the risk this is designed to remove. A public chatbot has no awareness of ERP permissions, no live connection to transactional data, and sends whatever is pasted to a third party. A grounded, permission-aware, on-prem assistant answers from live or indexed ERP data without leaving the network and without exceeding what the user is already authorized to see.
Related guides
AI Agents for ERP, Running On-Prem
A practical guide to on-prem AI agents for ERP: what they can safely automate, where human approval belongs, and how to design the guardrails.
Ask your ERP anythingNatural Language Query for ERP Data: Ask SAP, Infor, or Oracle a Question in Plain English
See how natural language query over SAP, Infor, Oracle, and NetSuite data works: grounded text-to-SQL, role-based permissions, and a visible audit trail.
GDPR + AI on ERP dataBuilding GDPR-Compliant AI on Top of Your ERP
Design AI on ERP data that satisfies GDPR: lawful basis, DPIA, data minimisation, and an on-prem architecture that avoids the Schrems II transfer problem.
ERP AI readinessIs Your ERP Ready for AI? A Readiness Assessment Checklist
Is your ERP actually ready for AI? A practical checklist covering master data quality, access and permissions, GPU sizing, and governance before you fund a pilot.
SAP BTP + self-hosted AISAP BTP and a Private LLM: Where the Generative AI Hub Fits and Where It Doesn't
Where SAP BTP's Generative AI Hub fits and where a self-hosted LLM belongs instead. CAP, Integration Suite, and an architecture decision framework.
SyteLine + air-gapped AIAn Air-Gapped Private LLM for SyteLine, Built for Defense Suppliers
Deploy a private LLM on an air-gapped network alongside Infor SyteLine for defense suppliers: no internet egress, ITAR and CMMC-aware architecture.
Plan it with numbers
Private RAG Corpus Sizing Calculator
Estimate chunk counts, vector index storage, raw text volume, and embedding compute time before you build a private retrieval system over your document estate.
Free ToolRAG vs Fine-Tuning Decision Assessment
Answer 8 questions about knowledge volatility, citation needs, data availability, and team capability to find whether RAG, fine-tuning, or a hybrid fits your project.
Free ToolPrivate AI Total Cost of Ownership Calculator
Model the full 3-year cost of an on-prem AI deployment, including amortized hardware, power, staff time, and support, against comparable API spend.
GuideNatural Language ERP Query Interface
Query your ERP using natural language. Transform plain English questions into SQL/API calls with LLM-powered interfaces that democratize ERP data access.
GuideEnterprise RAG Architecture: The Full 2026 Blueprint
A practitioner's blueprint for enterprise RAG in 2026: ingestion, chunking, embedding, retrieval, rerank, generation, and the eval loop that keeps it honest.
GuideA Data Governance Framework for Private AI
A practical data governance framework for private AI: classification, access controls, retention, and audit patterns that satisfy CMMC 2.0 and ITAR.
Talk it through with an engineer who knows SAP S/4HANA
Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.