ITAR and CMMC: Handling Technical Data in AI Workloads
ITAR and CMMC obligations do not disappear because the system processing your technical data happens to be an AI model instead of a traditional application, and treating an AI pipeline as somehow exempt from export control or CUI handling requirements is a fast path to a finding. If an engineering drawing, a technical specification, or a defense article description is ITAR-controlled technical data before it reaches an AI system, it remains ITAR-controlled technical data when a model reads it, summarizes it, or answers a question about it. The same logic applies to Controlled Unclassified Information under CMMC 2.0: the assessed boundary that governs where CUI can be processed governs an AI workload exactly as it governs any other application, which is why most defense manufacturers land on on-premises deployment as the only architecture that reliably clears the requirement.
Where AI Workloads Intersect ITAR and CMMC Obligations
Any AI workload that ingests, processes, retrieves, or generates output derived from export-controlled technical data or CUI inherits the full obligations attached to that data, regardless of how the workload is architected. This includes obvious cases like an engineering drawing chatbot, but also less obvious ones: a general-purpose document summarization tool that happens to be pointed at a folder containing technical data, or a customer support agent whose retrieval corpus includes service records referencing controlled part numbers or specifications. Map every data source an AI system touches against your existing ITAR technical data and CUI inventories before deployment, not after, since the AI system's data path often crosses boundaries the original data classification exercise never anticipated.
Technical Data Boundaries: What Counts as Export-Controlled Input
ITAR technical data includes information required for the design, development, production, manufacture, assembly, operation, repair, testing, maintenance, or modification of a defense article, which in practice covers far more than finished drawings: manufacturing process parameters, test procedures, and even certain training or troubleshooting documentation can qualify depending on content. An AI system fine-tuned on service manuals or process documentation for a defense product needs the same technical data classification review the source documents would get on their own, and the resulting model checkpoint inherits that classification. When in doubt about whether a specific document or dataset qualifies, treat it as controlled and route it through your export control review process before it becomes AI training or retrieval input.
- Technical data classification applies to process parameters and procedures, not only finished drawings
- A model fine-tuned on controlled technical data inherits the classification of its training input
- RAG corpora need document-level review before indexing, not just after a problem is discovered
- Route ambiguous documents through existing export control review before AI ingestion, not after
CUI Handling and the CMMC 2.0 Assessed Boundary
CMMC 2.0 requires CUI to be processed within an assessed environment meeting the applicable NIST SP 800-171 controls, and an AI workload processing CUI must live inside that same assessed boundary just as any other information system handling CUI would. This means the model, the inference infrastructure, the retrieval corpus, and any logging or monitoring system touching that data all fall within scope of the assessment, not just the application layer users interact with. Organizations sometimes assume a commercial AI vendor's general security certifications substitute for CMMC assessment, but a SOC 2 report or a general enterprise security posture does not equal a CMMC-compliant assessed boundary, and conflating the two is a common and costly misunderstanding.
Why Most Commercial AI APIs Fail This Test
Commercial AI APIs, even those with strong general security postures, typically cannot provide the specific guarantees ITAR and CMMC require: assurance that no non-US person can access the processing environment, a fully documented and assessable system boundary, and contractual terms matching the specific control requirements your CMMC level demands. Some cloud providers now offer government-specific or FedRAMP-authorized AI offerings that narrow this gap for CUI at lower CMMC levels, but ITAR technical data processing in a shared multi-tenant cloud environment remains a difficult case to clear regardless of the vendor's certifications, because the non-US-person access question is structurally hard for a shared commercial service to answer definitively.
Building an Assessed On-Prem AI Enclave
On-premises deployment inside your existing CMMC-assessed boundary, using open-weight models like Llama 3.3, Qwen3, or gpt-oss running on GPUs physically located within your controlled facility, removes the non-US-person access question and the shared-tenancy boundary question entirely, since the data never leaves an environment you already control and have already assessed. This does mean the AI infrastructure itself, GPU servers, storage, and networking, needs to meet the same NIST SP 800-171 controls as the rest of your assessed environment: access control, audit logging, configuration management, and the rest of the control families apply to the AI stack exactly as they apply to your ERP or file servers.
- Open-weight models on customer-owned GPUs inside the existing assessed physical and network boundary
- AI infrastructure subject to the same NIST SP 800-171 control families as other assessed systems
- No third-party API in the data path for ITAR technical data or CUI processing
- System security plan updated to explicitly include the AI workload's data flows and access controls
How Netray Delivers ITAR/CMMC-Compliant AI Deployments
Netray defaults to on-premises deployment for defense manufacturing clients specifically to answer the ITAR and CMMC data residency question at the architecture level rather than through vendor promises, running open-weight models on customer-owned hardware inside your existing assessed boundary. We help map data sources against your existing technical data and CUI inventories before any AI system touches them, update system security plan documentation to reflect the AI workload's data flows, and build the access control and audit logging the AI stack needs to meet the same control families as the rest of your assessed environment, so the AI deployment strengthens your CMMC posture instead of introducing a new gap.
Frequently Asked Questions
Can defense contractors use commercial AI APIs for ITAR technical data?
Generally no. Most commercial AI APIs cannot provide assurance that no non-US person can access the processing environment, which is a structurally difficult guarantee for a shared multi-tenant service to make regardless of its general security certifications. The practical path for most defense manufacturers is on-premises deployment of open-weight models inside their existing assessed boundary, removing the question entirely.
Does a SOC 2 certification satisfy CMMC requirements for an AI vendor?
No. SOC 2 addresses general security controls and is not equivalent to a CMMC-assessed boundary meeting the specific NIST SP 800-171 control families CUI processing requires. A vendor's SOC 2 report may be useful supporting evidence, but it does not substitute for confirming the AI workload operates within an environment that meets your applicable CMMC level's specific requirements.
What counts as ITAR technical data in an AI training or retrieval pipeline?
Any information required for the design, development, production, operation, repair, or maintenance of a defense article, which extends beyond finished drawings to manufacturing process parameters, test procedures, and certain troubleshooting documentation. A model fine-tuned on or retrieving from such documents inherits the technical data classification, so route ambiguous documents through export control review before they become AI input.
Does the AI infrastructure itself need to be inside the CMMC assessed boundary?
Yes. The model, inference infrastructure, retrieval corpus, and any logging or monitoring system touching CUI all fall within the scope of assessment, not just the user-facing application. GPU servers, storage, and networking supporting the AI workload need to meet the same NIST SP 800-171 control families, access control, audit logging, and configuration management, as the rest of your assessed environment.
Key Takeaways
- 1Where AI Workloads Intersect ITAR and CMMC Obligations: Any AI workload that ingests, processes, retrieves, or generates output derived from export-controlled technical data or CUI inherits the full obligations attached to that data, regardless of how the workload is architected. This includes obvious cases like an engineering drawing chatbot, but also less obvious ones: a general-purpose document summarization tool that happens to be pointed at a folder containing technical data, or a customer support agent whose retrieval corpus includes service records referencing controlled part numbers or specifications.
- 2Technical Data Boundaries: What Counts as Export-Controlled Input: ITAR technical data includes information required for the design, development, production, manufacture, assembly, operation, repair, testing, maintenance, or modification of a defense article, which in practice covers far more than finished drawings: manufacturing process parameters, test procedures, and even certain training or troubleshooting documentation can qualify depending on content. An AI system fine-tuned on service manuals or process documentation for a defense product needs the same technical data classification review the source documents would get on their own, and the resulting model checkpoint inherits that classification.
- 3CUI Handling and the CMMC 2.0 Assessed Boundary: CMMC 2.0 requires CUI to be processed within an assessed environment meeting the applicable NIST SP 800-171 controls, and an AI workload processing CUI must live inside that same assessed boundary just as any other information system handling CUI would. This means the model, the inference infrastructure, the retrieval corpus, and any logging or monitoring system touching that data all fall within scope of the assessment, not just the application layer users interact with.
Put this into numbers
Free interactive tools for exactly this problem. No signup to use them.
ITAR AI Workload Compliance Assessment
Score your AI deployments across eight dimensions of ITAR exposure, from technical data classification and US persons access control to technology control plan coverage.
Free ToolSovereign AI Readiness Assessment
Score your organization across eleven dimensions of sovereign AI readiness, from data residency and model provenance to cleared personnel and air-gapped operations.
Free ToolOn-Prem AI Security Hardening Checklist
A practical control checklist for securing self-hosted language models, covering model provenance, network isolation, data governance, host hardening, and audit readiness.
Terms used in this article
Deploying AI against ITAR technical data or inside a CMMC-assessed boundary? Netray builds on-prem AI deployments designed around your existing assessment, not around a vendor's promises.
Related Resources
Securing Model Weights in the Enterprise
Secure model weights end to end: custody controls, encryption at rest, access policies, and exfiltration prevention for regulated AI deployments.
AI & AutomationAir-Gapped Model Updates: A Patching Guide
Air-gapped model updates for enterprise AI: secure transfer procedures, hash verification, and staged rollout so patches never introduce risk.
AI & AutomationAudit Trails for AI Decisions: A Compliance Guide
Build audit trails for AI decisions that satisfy internal and external auditors: what to log, how long to retain it, and how to prove provenance.