Any ERPBuyer Guide

90-day ERP AI pilot

A 90-Day ERP AI Pilot Plan With Real Success Criteria

Short answer

A well-scoped ERP AI pilot fits in 90 days: two weeks to scope one use case against one ERP module, four to six weeks to build and ground it against real data, three to four weeks for real users to test it against defined accuracy and adoption targets, and a final decision week. The plan below gives week-by-week milestones and the specific success criteria, not vague ones, that turn a pilot into either a funded production rollout or an honest no-go.

ERP
SAP, Infor, Oracle, Microsoft Dynamics, Epicor, IFS, Deltek Costpoint
Industries
Manufacturing, Aerospace, Defense, Electronics
Written for
VP Operations

Most ERP AI pilots that fail do not fail because the technology did not work, they fail because nobody defined what success meant before starting, so the pilot drifts for six months with a shifting scope and ends in a shrug instead of a decision. A 90-day pilot with fixed dates and pre-agreed success criteria forces the opposite outcome: either the use case is proven and ready to fund for production, or it clearly is not, and everyone knows why.

Ninety days is enough time to properly ground an AI assistant or agent on one ERP module's real data, security-review it, and put it in front of real users for a few weeks, without letting scope creep turn it into a year-long program. It is not enough time to do that for five use cases at once, which is the most common way pilots blow their timeline: pick one use case, one ERP module, and one group of real users.

The plan below assumes a single, well-scoped use case, something like natural-language query over one module, PO exception follow-up, or NCR/CAPA drafting, run against a read-only connection to real (not synthetic) data, with a defined group of 10-30 users. It is written for a VP Operations sponsoring the pilot, coordinating with IT/security and a small group of end users, not for an enterprise-wide AI transformation program.

Each phase below has a duration and specific, checkable deliverables. The success criteria section further down gives concrete accuracy and adoption thresholds to agree with stakeholders before week one, so the go/no-go conversation in week 13 is a formality, not a debate.

What usually gets in the way

The problems we hear most from vp operations teams running SAP.

Pilots with no end date drift indefinitely

Without a fixed 90-day boundary and a scheduled decision meeting, pilots naturally extend as 'just one more feature', and the organization never gets a clean answer on whether the use case actually works.

Success criteria get defined after the fact

When nobody agreed in advance what 'good' looks like, the same result can be spun as a win by the sponsor and a failure by a skeptical stakeholder, and the disagreement itself becomes the reason nothing gets funded next.

Scope creep turns a pilot into a program

A pilot scoped to answer questions about open POs quietly grows to cover inventory, quality, and scheduling because stakeholders keep adding 'just this one more thing', and the 90-day window disappears.

The pilot environment does not resemble production

Testing against a sanitized sandbox or synthetic data set produces answers that look great in a demo and fall apart against the messy real data and edge cases of production.

Users disengage without visible weekly progress

If the pilot group does not see something new to try every one to two weeks, engagement drops, and the final adoption numbers understate what the tool could actually do.

Where AI earns its place in SAP

Each use case names the ERP objects it reads or writes, so your ERP team can judge the integration effort before anyone commits budget.

Natural-language query over one module

Users in one department, purchasing or finance, ask questions in plain language instead of running a report.

Touches: PO/PR tables and purchasing worklists, or GL/AP tables scoped to one module, with the corresponding existing reports as an accuracy baseline

Outcome: A well-scoped 90-day pilot can typically show whether question-answering accuracy meets the target against a fixed test set of 30-50 real questions collected from the department.

PO exception follow-up drafting

The assistant flags at-risk purchase orders and drafts a supplier follow-up email for a buyer to review and send.

Touches: Open PO lines, promise dates, ASN and receipt status, supplier contact records

Outcome: Success is measurable directly: number of exception emails drafted per week and buyer edit rate (how much a buyer had to change before sending).

NCR/CAPA drafting from inspection notes

A quality engineer's shorthand note or a photo caption gets turned into a structured NCR/CAPA draft record.

Touches: QM notifications or the equivalent NCR/CAPA module tables, prior similar-defect records

Outcome: Success is measured as time saved per record (compare manual entry time against draft-plus-review time) and quality-team acceptance rate of the draft.

Shop-floor traveler question answering

Operators ask about routing steps, work instructions, or the current status of a job without leaving the floor terminal.

Touches: Work order/traveler records, routing steps, linked document attachments

Outcome: Success is measured as reduction in interruptions to supervisors for routine status questions, tracked over the pilot's final three weeks.

AP invoice matching exception triage

The assistant flags 3-way match exceptions and suggests a likely cause (price variance, quantity variance, missing receipt).

Touches: AP invoice, PO, and receipt tables, existing exception queue

Outcome: Success is measured as reduction in average time-to-resolution for exceptions the assistant correctly categorized.

Engineering change impact summary

Given an ECO/ECN, the assistant summarizes which open orders, BOMs, and routings are affected.

Touches: ECO/ECN records, BOM and routing tables, open order lines referencing the affected item

Outcome: Success is measured as accuracy of the impact list against a manual review by an engineer, on a sample of real historical ECOs.

Demand forecast explanation

Given a forecast number, the assistant explains which factors (seasonality, recent order pattern, promotion) drove it.

Touches: Historical sales/demand tables, existing forecast output

Outcome: Success is measured as planner trust: whether planners report the explanation helped them decide to accept or override the forecast, tracked via a short weekly survey.

Reference architecture

A pilot's architecture should be a deliberately smaller version of the eventual production architecture, same layers, narrower scope, so that what you learn and build during the pilot carries forward rather than getting thrown away.

  1. 1

    ERP connectors

    Scope to a single module's API or read-replica tables; do not build a general-purpose connector framework during the pilot, that is production-phase work.

  2. 2

    Data and semantic layer

    Document only the fields and business terms the pilot's one use case actually touches; resist the urge to map the whole schema.

  3. 3

    Model serving

    Size for the pilot's 10-30 users, not for eventual production scale; a single mid-size GPU is often enough for a pilot's concurrent load.

  4. 4

    Retrieval and agents

    Keep the pilot read-only or draft-only (no autonomous writes) so the approval workflow question does not become a blocker before you have even proven the core answer quality.

  5. 5

    Governance and audit

    Log every query and draft output from day one; this log is also your evidence base for the accuracy and adoption metrics in the go/no-go decision.

Integration notes for your ERP team

  • Lock the use case and the ERP module scope in writing before week one; any request to expand scope during the 90 days gets logged as a 'phase 2 candidate', not folded into the current pilot.
  • Use a read-only database replica or the vendor's official API, never a direct write-capable connection to production, for the entire pilot duration.
  • Collect 30-50 real historical questions or exception cases from the target users in week one; this becomes your accuracy test set for the whole pilot, so build it before the tool exists to avoid biasing it toward what the tool happens to do well.
  • Schedule a working session with end users every one to two weeks, not just a final demo, so adoption and usefulness feedback shapes the build in real time rather than surfacing all at once in week twelve.
  • Track every AI answer against the test set with a simple correct/partially correct/incorrect rubric agreed with users in week one, so accuracy numbers are not argued about after the fact.
  • Keep write-back actions in draft-only mode (a human sends the email, approves the NCR, clicks submit) for the entire pilot; proving answer and draft quality first is a cleaner test than also proving an approval workflow in the same 90 days.
  • Book the go/no-go decision meeting on the calendar in week one, for week thirteen, with the sponsor and key stakeholders, so the decision cannot quietly slip.

Deployment options

Pilot on a production read-replica, on-prem

Regulated data, or any organization that wants the pilot environment to match production exactly

Slightly more setup time in week one to provision the replica, but avoids the risk of a pilot succeeding against sanitized data and failing against the real thing.

Pilot in a private cloud tenant

Organizations that want to avoid new on-prem hardware for a time-boxed pilot

Faster to stand up if GPU procurement would otherwise delay the start date; migrate to on-prem at the production phase if that is the target end state.

Pilot on a subset of data, hybrid

Multi-site organizations piloting at one site before deciding on the others

Keeps the pilot's blast radius small while still testing against real, not synthetic, data from the chosen site.

Compliance and data control

How the architecture supports your obligations. Certification and accountability stay with your organisation; the design keeps the evidence straightforward.

ITAR / EAR

If the pilot's data set could include controlled technical data, confirm the deployment (on-prem, no external calls) is compliant before week one; do not plan to 'clean the data later', scope the pilot to exclude controlled records from day one instead.

CMMC 2.0 / NIST SP 800-171

Run the pilot inside your existing enclave boundary if it will touch CUI, so you are not creating an unscoped asset that has to be retrofitted into your System Security Plan later.

Internal change control

Even a pilot should go through a lightweight change-control and security review before touching production-adjacent data; skipping this to save a week often costs more time later when security objects mid-pilot.

Data retention

Define and document how long pilot logs and any indexed documents are retained, and delete them per that policy if the pilot ends in a no-go.

How an engagement runs

Phase 1 . Weeks 1-2

Scope and set up

  • -Written scope: one use case, one module, named pilot users
  • -30-50 item accuracy test set collected from real users
  • -Read-only data connection provisioned and security-reviewed
  • -Go/no-go meeting booked for week 13

Phase 2 . Weeks 3-6

Build and ground

  • -Working prototype answering or drafting against the in-scope use case
  • -First accuracy pass against the test set with a documented score
  • -Weekly user working session started

Phase 3 . Weeks 7-10

User testing

  • -Tool in daily use by the pilot group
  • -Weekly accuracy and adoption tracking (queries per user, correct/partial/incorrect rate)
  • -User feedback log with issues triaged and fixed or deferred

Phase 4 . Weeks 11-13

Decision

  • -Final accuracy score against the agreed threshold
  • -Adoption report (active users, queries per week, time saved estimate)
  • -Documented go/no-go recommendation with production scope and cost estimate if go

Questions to ask any vendor, including us

A short list that separates real SAP AI work from a chatbot demo.

  1. What specific accuracy threshold, against what test set, defines success before we start?
  2. What adoption level (active users, queries per week) needs to be hit for this to be worth funding at production scale?
  3. Is the pilot running against real production-representative data, or a sanitized sample?
  4. What happens to the pilot infrastructure and code if the answer is no-go, is it disposable or does it seed production?
  5. Who has final authority to say go or no-go, and is that person committed to the week-13 decision date?
  6. What is scope creep going to cost us if we let it happen, in weeks and in dollars?
  7. What is the realistic production timeline and cost if the pilot succeeds, so the go decision is informed?

Frequently asked questions

Why 90 days specifically?

It is long enough to properly ground a use case against real data, security-review it, and get three to four weeks of genuine user testing, and short enough that scope stays disciplined and the organization gets a decision instead of an open-ended program.

Can we pilot more than one use case at once?

You can, but each additional use case roughly doubles the coordination overhead and dilutes attention on the test set and user feedback for each one. A single well-proven use case that leads to a confident go decision is usually more valuable than three shallow ones.

What if the pilot reveals the use case needs more than 90 days?

That is a legitimate and common outcome, and it is exactly what the structured plan is for: you get a clear read on what specifically is missing (data quality, a connector gap, an accuracy shortfall) rather than an ambiguous stall, which is a far better basis for deciding whether to extend.

Should the pilot include write-back actions?

Generally no, for a first 90-day pilot. Keeping write-back in draft-only or human-approved mode isolates whether the AI's answers and drafts are good, which is the harder and more important question, from the separate question of how to design an approval workflow.

How many users should be in the pilot group?

10 to 30 real users doing their actual jobs is usually the right range: enough to generate meaningful adoption data, small enough to manage feedback sessions and support closely.

What does a no-go outcome actually mean?

A no-go means this specific use case, against this data, in this timeframe, did not meet the agreed bar, not that AI on your ERP is a dead end generally. A good pilot report identifies exactly what would need to change (better data, a different use case, more time) for a future attempt.

Who should sponsor a 90-day ERP AI pilot?

A business sponsor (VP Operations, department head) who owns the use case and can commit named users' time, working alongside IT for the technical and security pieces, tends to produce better outcomes than an IT-only pilot with no business owner accountable for adoption.

Talk it through with an engineer who knows SAP

Bring one real question your team cannot answer from the ERP today. We will map the data path, the model, and where it runs, and tell you honestly if AI is the wrong tool for it.