How to Hire On-Prem AI Consultants: A Buyer's Guide
Hiring an on-prem AI consultant is a different exercise than hiring a general AI consultant, and treating it the same way is the single most common reason regulated manufacturers end up with a partner who cannot deliver. Most of the market built its portfolio on hosted APIs and cloud notebooks, which teaches almost nothing about GPU sizing, air-gapped networks, quantization tradeoffs, or integrating a model into a live SyteLine or Infor LN environment without breaking a nightly batch job. The skills that predict success on an on-prem engagement are specific, testable, and rarely visible on a slide deck. This guide covers what to test for, what to ask in an interview, and the red flags that show up before the contract is signed if you know where to look.
The Skills That Actually Matter for On-Prem AI Work
Screen for hands-on experience across four areas: GPU capacity planning (can they size a cluster for a target model and concurrency, not just quote a vendor's marketing spec sheet), model serving and quantization (have they run vLLM or SGLang in production, and can they explain when AWQ or FP8 quantization is the right call versus when it degrades accuracy unacceptably), network and data security for isolated environments (do they understand what changes when there is no outbound internet access), and integration engineering into systems of record. A consultant who is excellent at prompt engineering but has never sized a GPU server or touched an ERP API will struggle the moment the project leaves the sandbox.
- GPU sizing: can size a cluster for a named model, context length, and concurrent user count, not just quote list specs
- Serving stack depth: production experience with vLLM, SGLang, or TensorRT-LLM, not only local Ollama demos
- Quantization judgment: can explain the accuracy tradeoff of FP8, AWQ, or GGUF for a specific workload
- Systems integration: has shipped a working integration against an ERP, MES, or PLM API, not only a REST demo
Interview Questions That Separate Practitioners From Slide Decks
Ask questions that require a specific, defensible answer rather than a general one. A practitioner will answer with numbers, tradeoffs, and a story about something that went wrong. A generalist will answer in platitudes about transformation and innovation. Push on the failure question especially hard: everyone who has actually done this work has a project that underdelivered, and how they talk about it tells you more than any success story.
- Walk me through sizing a GPU cluster to serve a 70B parameter model to 50 concurrent users with a 2-second latency budget
- Describe a fine-tuning or RAG project that did not hit its accuracy target. What did you change, and what did not work?
- How would you architect this for an environment with no outbound internet access?
- What is your process for handing off a golden evaluation set and a runbook so we do not depend on you forever?
Red Flags in an On-Prem AI Consultant or Firm
Some warning signs are obvious once you know to look for them. A portfolio built entirely on cloud API integrations with no GPU hardware line item anywhere is a strong signal the firm has never actually deployed on-prem. Vague answers about data residency, an inability to name the compliance frameworks relevant to your industry such as CMMC or AS9100, and pricing that seems too good given the scope are all worth probing further. So is a firm that cannot produce a single reference from a regulated industry client when your entire reason for going on-prem is regulatory.
- No GPU hardware or on-prem deployment experience anywhere in the portfolio, only cloud API integrations
- Cannot explain the accuracy tradeoff of a specific quantization method when asked directly
- Vague or evasive answers about data residency and where model inference actually runs
- No reference available from a client in a regulated industry, despite regulation being your driver for going on-prem
Checking References the Right Way
Ask referees the questions they were not prepared for. Do not ask if they liked the consultant, ask whether the project hit its original timeline and budget, and if not, why not. Ask what broke after go-live and how the consultant responded. Ask whether the client's internal team could operate the system without the consultant six months later, or whether they are still dependent. Ask if they would hire the same individual again for a different project, not just the same firm, because firms rotate staff and the person who sold the engagement is not always the person who built it.
How Netray Approaches These Conversations
Netray works exclusively in on-prem and air-gapped environments for aerospace and defense, discrete manufacturing, and electronics clients, which means GPU sizing, quantization, and ERP integration against SyteLine and Infor LN are the daily work, not a side capability. We expect buyers to ask the questions above and we will connect you directly with references from regulated clients rather than curated case studies. If a use case does not need on-prem deployment, we say so during scoping rather than selling infrastructure the requirement does not justify.
Frequently Asked Questions
What certifications should an on-prem AI consultant have?
There is no single certification that reliably predicts competence in this field yet, so weight demonstrated project experience over credentials. Cloud provider AI certifications tell you almost nothing about on-prem work. What matters more is a track record of GPU cluster deployments, production model serving, and integration work in regulated environments, backed by references you can actually call and ask hard questions.
How much technical vetting is needed before hiring an AI consultant?
Plan for at least one technical interview run by someone on your own IT or engineering team, not just a business stakeholder conversation. Ask the specific sizing and failure-mode questions covered above, and request a short architecture write-up for your actual use case as part of the proposal. A consultant who cannot produce a credible architecture sketch before the contract is signed will not produce one after.
Should I hire an individual consultant or a firm for on-prem AI work?
Individual consultants can be excellent for a narrowly scoped build but create single-point-of-failure risk if they become unavailable mid-project, and rarely cover the full stack from GPU infrastructure through ERP integration to security review alone. Firms spread that risk and typically bring broader coverage, but verify that the specific people assigned to your project, not just the firm's marquee names, have the on-prem experience you are hiring for.
Key Takeaways
- 1The Skills That Actually Matter for On-Prem AI Work: Screen for hands-on experience across four areas: GPU capacity planning (can they size a cluster for a target model and concurrency, not just quote a vendor's marketing spec sheet), model serving and quantization (have they run vLLM or SGLang in production, and can they explain when AWQ or FP8 quantization is the right call versus when it degrades accuracy unacceptably), network and data security for isolated environments (do they understand what changes when there is no outbound internet access), and integration engineering into systems of record. A consultant who is excellent at prompt engineering but has never sized a GPU server or touched an ERP API will struggle the moment the project leaves the sandbox..
- 2Interview Questions That Separate Practitioners From Slide Decks: Ask questions that require a specific, defensible answer rather than a general one. A practitioner will answer with numbers, tradeoffs, and a story about something that went wrong.
- 3Red Flags in an On-Prem AI Consultant or Firm: Some warning signs are obvious once you know to look for them. A portfolio built entirely on cloud API integrations with no GPU hardware line item anywhere is a strong signal the firm has never actually deployed on-prem.
Put this into numbers
Free interactive tools for exactly this problem. No signup to use them.
DeepSeek R1 GPU Requirements Calculator
Estimate the GPU count and capital cost required to self-host DeepSeek R1, a 671B-parameter mixture-of-experts reasoning model with only 37B active per token.
Free ToolGLM-4.5 On-Prem Sizing Calculator
Size VRAM, GPU count, and capital cost for GLM-4.5 or the smaller GLM-4.5-Air, both mixture-of-experts models tuned for agentic and coding workloads.
Free ToolKimi K2 Deployment Cost Calculator
Estimate the multi-GPU cluster cost required to self-host Kimi K2, a roughly 1 trillion parameter mixture-of-experts model with only 32B active per token.
Terms used in this article
Vetting a shortlist of on-prem AI consultants? Netray will walk through our own architecture decisions and references so you have a real benchmark to compare against.
Related Resources
Selecting an AI Implementation Partner: Evaluation Criteria
Selecting an AI implementation partner: a weighted evaluation scorecard, reference-check questions, and why a pilot-first contract beats a big-bang one.
AI & AutomationOn-Prem AI Consulting Rates in 2026
On-prem AI consulting rates for 2026: realistic day rates by role and region, what actually drives the variance, and fixed-fee versus time-and-materials pricing.
AI & AutomationThe AI Statement of Work Checklist
The AI statement of work checklist: scope language that prevents creep, IP and model ownership clauses, measurable acceptance criteria, and payment milestones.