SAP BTP + self-hosted AI
SAP BTP and a Private LLM: Where the Generative AI Hub Fits and Where It Doesn't
Short answer
SAP BTP's Generative AI Hub gives enterprise architects a managed way to call various models through a consistent interface inside the SAP ecosystem, using CAP, Integration Suite, and SAP's extensibility model. A self-hosted LLM outside BTP is the right choice when the model, the data, or both need to stay entirely within customer-controlled infrastructure rather than SAP-managed cloud infrastructure, even a well-governed one.
- ERP
- SAP Business Technology Platform, SAP S/4HANA
- Industries
- Manufacturing, Aerospace, Defense, Electronics
- Written for
- Enterprise Architect
If you're the enterprise architect responsible for the AI layer across an SAP landscape, you're weighing BTP-native tooling, CAP extensions, Integration Suite, and the Generative AI Hub, against a fully self-hosted stack outside BTP entirely. Both are legitimate, and the right answer usually depends on where you draw the trust boundary, not on which is technically superior.
BTP offers a genuinely useful extensibility model: CAP-based services give you the standard SAP developer experience and authorization model, Integration Suite handles data movement to and from S/4HANA, and the Generative AI Hub and AI Core provide a managed way to call various models, including some open-weight ones, without standing up your own inference infrastructure.
Where it falls short for some organizations is the trust boundary itself: even when the Generative AI Hub is used to call an open-weight model, the inference runs on SAP-managed infrastructure. For air-gapped requirements, export-controlled technical data, or strict data-sovereignty mandates, that distinction is not a detail, it is the whole question.
This page lays out the architecture options and an honest decision framework: where BTP-native tooling is sufficient, where a self-hosted model belongs instead, and how the two combine in a hybrid design that most enterprise architects end up building.
What usually gets in the way
The problems we hear most from enterprise architect teams running SAP Business Technology Platform.
Bring your own model still runs on SAP infrastructure
BTP's Generative AI Hub is presented as bring your own model, but the inference still runs on SAP-managed infrastructure, a different trust boundary than fully self-hosted, even when the underlying model is open-weight.
No clear framework for when BTP is sufficient
Enterprise architects are expected to have an SAP AI strategy without a clear framework for when BTP-native tooling is sufficient and when it isn't.
Documentation thins out for self-hosted patterns
CAP-based extensions and Integration Suite flows are well documented for standard scenarios, but the guidance gets thin fast for a fully self-hosted model architecture.
Two different security reviews to reconcile
Procurement and security teams ask different questions about BTP, a known SAP product, than about a self-hosted stack, a newer pattern for most SAP shops, and architects get caught reconciling the two reviews.
The BTP boundary isn't always obvious
It's not always clear which parts of a proposed architecture are inside BTP and which are outside it, which slows down every security and cost review.
Where AI earns its place in SAP Business Technology Platform
Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.
CAP extension calling a self-hosted model
A Cloud Application Programming service on BTP orchestrates business logic and SAP data access, but the actual inference call goes to a self-hosted model outside BTP's managed AI services.
Touches: CAP service, OData/CDS views, custom BTP extension
Outcome: keeps the SAP-native development experience while inference stays on customer-controlled infrastructure
Integration Suite as the connector, not the model host
Integration Suite flows move data between S/4HANA and an external retrieval and agent layer without routing through BTP's AI services.
Touches: Integration Suite iFlows, OData services, IDocs
Outcome: reuses existing integration investment for the data path, independent of where inference happens
Document extraction alongside a private RAG layer
BTP's Document Information Extraction service handles structured extraction from certain document types, feeding a separately hosted retrieval index for broader question-answering.
Touches: Document Information Extraction service, attached PDFs, DMS documents
Outcome: combines a managed SAP service for one narrow task with a private layer for everything else
Fiori-native extension backed by a self-hosted model
A Fiori-embedded extension built on BTP's UI5/Fiori elements framework calls an API backed by a self-hosted model rather than AI Core.
Touches: Fiori elements extension, custom OData service, self-hosted model API
Outcome: gives users a native-feeling SAP experience without SAP-managed inference
Hybrid model routing by sensitivity
Routine, low-sensitivity queries go through BTP's Generative AI Hub for operational simplicity, while sensitive queries route to the self-hosted model.
Touches: query classification logic, BTP Generative AI Hub, self-hosted model endpoint
Outcome: balances operational simplicity against data-sensitivity requirements on a per-query basis
Event-driven agent triggers via BTP
SAP Event Mesh on BTP triggers an agent workflow hosted outside BTP when a business event occurs, such as a goods receipt or a new notification.
Touches: SAP Event Mesh, IDoc/event triggers, external agent service
Outcome: uses BTP's event backbone without requiring the agent's model to run there
Unified cost and governance view
A BTP-hosted extension surfaces usage and cost across both BTP-native AI services and the self-hosted layer in one place for the architecture team.
Touches: BTP usage/cost APIs, self-hosted model logging
Outcome: gives architects one view instead of reconciling two separate reporting systems
Reference architecture
The core design decision is where model inference runs: SAP-managed infrastructure through BTP's AI Core and Generative AI Hub, or customer-controlled infrastructure outside BTP entirely, with CAP and Integration Suite handling the SAP-side logic and data movement either way.
- 1
ERP connectors
OData v2/v4, CDS views, and Integration Suite iFlows, whether the model itself runs on BTP or outside it.
- 2
Data/semantic layer
A retrieval index hosted outside BTP's managed services when data sovereignty requires it, or within BTP's data services when that trust boundary is acceptable.
- 3
Model serving
The core decision point: SAP's Generative AI Hub or AI Core, running on SAP-managed infrastructure with various model choices, versus a self-hosted open-weight model on customer-controlled GPUs.
- 4
Retrieval/agents
CAP-based extensions or an external agent framework, calling either model path through a consistent interface so the choice remains swappable.
- 5
Governance/audit
BTP's own audit logging for anything that touches its services, plus a separate audit trail for the self-hosted layer, reconciled into one reporting view.
Integration notes for your ERP team
- CAP, the Cloud Application Programming model, is the standard way to build BTP-native extensions that call out to a model, whether that model is AI Core-hosted or an external self-hosted endpoint.
- Integration Suite iFlows handle SAP-to-external-system data movement regardless of where the model runs, so existing Integration Suite investment is not wasted by choosing a self-hosted model.
- SAP Event Mesh can trigger externally hosted agent workflows on business events without requiring the agent's inference to run inside BTP.
- Authentication between BTP extensions and a self-hosted model endpoint typically uses OAuth2 client credentials or mutual TLS, following the same pattern BTP uses for any external destination.
- Document Information Extraction and other narrow BTP AI services can be used for their specific task, such as structured document extraction, even when the broader retrieval and agent layer is self-hosted, since they serve different purposes.
- Cost from BTP-native AI services is metered separately from self-hosted infrastructure cost; architects should model both before committing to an all-BTP or all-self-hosted approach.
- Where a fully self-hosted architecture is required, Integration Suite and CAP are still useful as the SAP-side data and event layer; only the model-serving layer moves outside BTP.
Deployment options
Air-gapped on-prem (fully outside BTP)
defense and export-controlled environments where even SAP-managed cloud infrastructure is not acceptable for the model or the data used to ground it
Connectors may still use BTP Integration Suite for data movement between systems, but inference and retrieval run entirely on customer-owned infrastructure.
Private/sovereign cloud, BTP-adjacent
organizations comfortable with SAP-managed infrastructure for orchestration but wanting model choice and cost control
CAP extensions and Integration Suite handle SAP-side logic, while a self-hosted model runs in a private VPC that BTP calls into through an API.
Hybrid (BTP AI Hub + self-hosted)
large landscapes with mixed sensitivity, low-risk queries and high-sensitivity queries side by side
Query classification routes to BTP's Generative AI Hub for routine cases and to a self-hosted model for sensitive ones, under one governance framework.
Compliance and data control
How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.
Data residency (BTP region vs. customer infrastructure)
Confirm the specific BTP region and sub-service hosting for any BTP-native AI feature in scope; for data that cannot leave customer infrastructure under any circumstance, route inference to the self-hosted layer instead.
Export control (ITAR, EAR)
Technical data subject to export control restrictions should not transit BTP's managed AI services regardless of region, since the infrastructure remains SAP-operated; keep that data path entirely on customer-controlled infrastructure.
EU AI Act / NIS2
For EU-based landscapes, document which parts of the architecture are SAP-managed through BTP versus customer-managed and self-hosted, since risk classification and incident-reporting obligations can differ by who operates the system.
Cost and usage governance
BTP-native AI services bill on SAP's consumption model; the self-hosted layer bills on infrastructure cost; a combined governance dashboard avoids surprises from either side of the architecture.
Where Netray fits
ERPray
The retrieval and agent layer described here, whether self-hosted alongside BTP or fully outside it, is ERPray's core architecture: grounded question-answering with the underlying OData or CDS query shown.
Custom build
Hybrid routing between BTP's Generative AI Hub and a self-hosted model, or CAP-based extensions calling an external model, is architecture-specific work that typically needs a scoped build.
How an engagement runs
Phase 1 . 2-3 weeks
Discovery
- -Inventory of BTP services already in use, including Integration Suite, CAP extensions, and AI Core
- -Data classification of what can transit BTP-managed AI services versus what cannot
- -Target architecture diagram showing the BTP and self-hosted boundary
- -Cost model comparing BTP-native and self-hosted paths at expected volume
Phase 2 . 6-8 weeks
Pilot
- -A CAP extension or Integration Suite flow wired to a self-hosted model for one or two use cases
- -Side-by-side comparison against a BTP-native AI Hub call where applicable
- -Latency and cost measurement for both paths
- -Go/no-go on the target architecture
Phase 3 . 8-12 weeks
Production
- -Hardened self-hosted model serving with authenticated BTP integration
- -Combined governance and audit view across BTP and self-hosted components
- -Documented boundary for future security reviews
- -Handoff to the enterprise architecture team
Phase 4 . ongoing
Scale
- -Additional CAP extensions or agent workflows added on the established pattern
- -Periodic review of BTP's evolving AI Core and Generative AI Hub capabilities against the self-hosted boundary
- -Cost optimization across both paths
- -Model upgrades on the self-hosted side evaluated independently of SAP's release cycle
Questions to ask any vendor, including us
A short list that separates real SAP Business Technology Platform AI work from a chatbot demo.
- Which specific BTP region and sub-service would host any BTP-native AI feature we use, and what's SAP's data processing agreement for it?
- Can we mix BTP-native AI Hub calls and a self-hosted model in the same application, or does that add unsupported complexity?
- What does authentication look like between a BTP extension and infrastructure we host ourselves?
- How is cost metered differently between BTP's consumption-based AI pricing and our own GPU infrastructure?
- Does using Integration Suite for data movement create any dependency on BTP's AI services, or are they fully separable?
- Who audits the boundary between what's SAP-managed and what's customer-managed in this architecture?
- If we start BTP-native and later need to move inference off SAP-managed infrastructure, how much of the CAP and Integration Suite investment carries forward?
Frequently asked questions
Is SAP BTP's Generative AI Hub the same as an on-prem or private LLM deployment?
No. The Generative AI Hub is a managed service running on SAP-operated infrastructure within BTP, even when it is used to call an open-weight model. It gives you model choice and a consistent interface, but the inference itself is not on customer-controlled infrastructure the way a self-hosted deployment is. For strict data-sovereignty or export-control requirements, that distinction matters.
Can we use BTP's Integration Suite without using BTP's AI services?
Yes. Integration Suite iFlows move data between S/4HANA and external systems regardless of where any AI inference happens. It is common to use Integration Suite for the SAP-side data path while routing model inference to a self-hosted endpoint entirely outside BTP's AI services.
What's the advantage of a CAP extension over a fully external application?
A CAP-based extension gives you the standard SAP developer experience, authorization model, and deployment lifecycle inside BTP, while still letting the actual model call go to a self-hosted endpoint through a normal outbound destination. It keeps the SAP-native parts of the architecture SAP-native without forcing the inference itself onto SAP-managed infrastructure.
Should sensitive and non-sensitive queries use different model paths?
It's a reasonable architecture for large, mixed-sensitivity landscapes: route routine, low-risk queries to BTP's Generative AI Hub for operational simplicity, and route anything touching export-controlled or highly sensitive data to a self-hosted model. This requires a reliable query classification step and clear governance for what counts as sensitive, which should be defined before, not during, rollout.
Does choosing a self-hosted model waste our existing BTP investment?
Generally no. Integration Suite flows, CAP extensions, Event Mesh triggers, and Document Information Extraction all remain useful regardless of where the core model runs; only the AI Core or Generative AI Hub piece specifically is bypassed. Most of an enterprise architect's BTP investment sits in the integration and extensibility layer, not the AI service itself.
How do we explain this architecture to a security review that only knows is it SAP or not?
Draw the boundary explicitly: which components are SAP-operated, including BTP services and any AI Core or Generative AI Hub calls, and which are customer-operated, meaning the self-hosted model, its retrieval index, and its infrastructure. Reviewers generally accept a hybrid architecture once the boundary and data flow are documented clearly, rather than treated as an all-or-nothing SAP question.
Is this architecture more complex to maintain than an all-BTP or all-self-hosted approach?
It adds one more integration point, the connection between BTP and the self-hosted model, but often reduces overall complexity by letting each part of the landscape use the tool suited to it: BTP for SAP-native extensibility and integration, a self-hosted model for the inference that must stay off SAP-managed infrastructure. Most enterprise architects find the added connector simpler than forcing everything through one path that doesn't fit all their data.
Related guides
AI for SAP S/4HANA, Running On-Prem or in Your Private Cloud
Run AI on SAP S/4HANA without sending ERP data to a public API. On-prem and private-cloud architecture, CDS views, OData, and honest deployment trade-offs.
SAP Joule + on-prem alternativeSAP Joule or a Private LLM Beside SAP: An Honest Comparison
An honest comparison of SAP Joule and Business AI against a private, self-hosted LLM beside SAP: what each covers, where they overlap, and where CIOs run both.
SAP ECC 6.0 + AIGet AI Value From SAP ECC Now, and Use It to De-Risk the Move to S/4HANA
SAP ECC 6.0 mainstream maintenance ends in 2027. Add AI value on ECC now, using its existing BAPIs and IDocs, and use it to de-risk the S/4HANA migration.
Ask your ERP anythingNatural Language Query for ERP Data: Ask SAP, Infor, or Oracle a Question in Plain English
See how natural language query over SAP, Infor, Oracle, and NetSuite data works: grounded text-to-SQL, role-based permissions, and a visible audit trail.
RAG + SQL + permissionsA Private LLM Grounded on Your ERP Data
How a private LLM answers questions on your ERP data: RAG plus text-to-SQL, role-based permissions inherited from the ERP, and where each fits.
SAP Business One + AIAI for SAP Business One, Without Sending Your Books to a Public API
Add AI to SAP Business One without sending customer data or financials to a public API. Service Layer and DI API integration, sized for an SMB budget.
Plan it with numbers
Self-Hosted LLM Hardware Estimator
Estimate the VRAM footprint, GPU count, and hardware budget required to self-host an open-weight LLM with your concurrency and context needs.
Free ToolLLM API vs Self-Hosted Cost Calculator
Model your API bill from requests and token mix, compare it against an all-in self-hosted monthly cost, and see monthly and annual savings.
Free ToolRAG vs Fine-Tuning Decision Assessment
Answer 8 questions about knowledge volatility, citation needs, data availability, and team capability to find whether RAG, fine-tuning, or a hybrid fits your project.
GuideEnterprise RAG Architecture: The Full 2026 Blueprint
A practitioner's blueprint for enterprise RAG in 2026: ingestion, chunking, embedding, retrieval, rerank, generation, and the eval loop that keeps it honest.
GuideAn API-First ERP Integration Strategy
An API-first ERP integration strategy replaces nightly file drops with REST and event APIs. Learn contract design, versioning, throttling, and security.
GuideOn-Prem LLM Deployment Architecture: Reference Guide
Reference architecture for on-prem LLM deployment: inference servers, GPU sizing, RAG pipelines, and security zones for regulated manufacturers.
Talk it through with an engineer who knows SAP Business Technology Platform
Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.