- A large majority of enterprise AI contracts implicitly allow vendors to train models on customer data.
- Nearly half of all employee AI interactions occur on personal accounts, bypassing corporate security controls.
- Enterprise data protections from major AI providers require companies to audit contracts and curb shadow AI use.
A large share of enterprise AI contracts quietly grant vendors the right to train models on customer data, and most companies signing them have no idea. A June 24 PYMNTS.com analysis drawing on Stanford Law School’s CodeX center found the practice is built into standard SaaS agreement language that predates generative AI. At the same time, nearly half of all employee AI interactions are happening on personal accounts, bypassing whatever enterprise protections companies think they have.
Hidden Contract Clauses Erode Enterprise Data Safeguards
Enterprise leaders typically assume that sensitive corporate data, source code, financial records, legal documents, stays proprietary when they deploy AI tools built for business use. The PYMNTS.com report, published June 24, 2026, challenges that assumption directly.
The analysis, drawing on data from TermScout and Stanford Law School’s CodeX center, found that a large majority of AI contracts assert data usage rights extending beyond what is strictly necessary for service delivery, well above the rate seen across SaaS agreements generally. These provisions were often drafted before generative AI existed, and they use broad language, “improve,” “build,” “enhance”, that can cover AI model training and fine-tuning unless customers negotiate explicit restrictions. The concern is not hypothetical: proprietary workflows and commercially valuable patterns absorbed during training could, in principle, surface in products sold to competitors.
The Federal Trade Commission warned in February 2024 that unilaterally revising privacy commitments to enable AI training could be considered unfair or deceptive under the FTC Act.
Real disputes have already emerged. Design software company Figma faced a proposed class-action lawsuit in November 2025 alleging unauthorised use of customer designs to train its generative AI tools, a claim Figma denied, stating explicit authorisation was required. Adobe clarified its policies after customer concerns, committing not to train AI systems on customer data, though its Firefly generative AI engines use a licensed image collection. The distinction that matters here is between conventional SaaS, where data is processed and returned, and AI engagements, where vendors may seek to use that data to improve models serving other clients.
Shadow AI Exposes Half of Enterprise Conversations to Training
Contractual risk is only part of the problem. What employees actually do with AI tools creates a second exposure that enterprise agreements cannot cover.
According to a LayerX report on enterprise AI usage, 47% of all enterprise AI conversations are conducted using personal identities rather than corporate accounts, placing nearly half of all AI-related work outside corporate security and governance controls entirely.
The LayerX report found sensitive data exposure rates vary considerably by platform, with different figures recorded for DeepSeek, ChatGPT, Copilot, Gemini Enterprise, and Copilot M365.
Leading Providers Detail Enterprise-Grade Data Privacy
The major cloud AI providers have been explicit about what their enterprise tiers do and do not do with customer data, and the commitments are substantive, provided customers are actually using those tiers.
Google Cloud states it does not use customer data to train models without prior permission, a commitment codified in the Training Restriction sections of its Service Specific Terms for Google Cloud Platform and Google Workspace. Fine-tuning data, the company states, is exclusively for the customer’s use. One caveat: the grounding with Google Search feature stores prompts and contextual information for 30 days for debugging, and prompt logging for abuse monitoring may occur unless a zero data retention exception is requested.
Microsoft‘s Azure OpenAI Service takes a comparable position. Customer prompts, completions, embeddings and uploaded training data are not used to train OpenAI’s or Microsoft’s models and are not made available to other customers. The default includes a 30-day abuse-monitoring window, with a “modified abuse monitoring” option that disables data storage for sensitive workloads. Fine-tuned models remain exclusively with the customer that created them, and data stays within the customer-specified geography.
Amazon Bedrock commits to not storing or using any data sent to or received from its foundation models for training. AWS states that customer data is not shared with model providers and is not used to improve base models. When a foundation model is tuned, it operates on a private copy for that customer’s exclusive use. For AWS-sold models on Bedrock, the model provider reportedly has no access to customer data, prompts or completions.
OpenAIfor its ChatGPT Business, Enterprise and Edu offerings, states it does not train models on business data by default. These policies, taken together, mark a clear operational distinction from consumer-grade AI services, but they only apply to customers actually using enterprise-tier accounts.
The Persistent Gap Between Policy and Practice
The commitments from major providers are real. The problem is that many organisations are not actually operating under them.
The report recommends explicit contractual restrictions on data usage for model training, clear definitions of what constitutes training data, and notification requirements for model changes.
The LayerX data makes the operational picture harder to ignore. With nearly half of enterprise AI interactions occurring outside sanctioned corporate channels, enterprise-grade privacy assurances from providers are effectively bypassed for a large portion of actual usage. The gap is most acute at smaller companies without dedicated AI security teams, where neither the legal review capacity nor the technical monitoring infrastructure is typically in place. That is where shadow AI proliferates most freely, and where the distance between policy and practice tends to be widest. For broader context on how enterprises are managing AI adoption risks, the recent data on enterprise AI agent rollbacks points to similar governance gaps playing out in deployment decisions.
Strengthening Enterprise AI Governance to Close the Loop
Closing the gap between perceived and actual data privacy requires action across legal, IT and operational teams simultaneously, there is no single fix.
Legal and procurement teams need to review AI-related clauses in all SaaS contracts with precision. Generic “improve product” language is not sufficient protection. Contracts should include explicit “no training on customer data” commitments, clear definitions of what constitutes training data, and clarity on ownership of any insights derived from models. This kind of clause-level review requires investment, but it is far less costly than discovering post-hoc that proprietary data has been absorbed into a vendor’s model.
IT and security departments need to address shadow AI directly, using discovery and monitoring tools to identify unsanctioned usage before it compounds. The LayerX report recommends a layered approach focused on identities, data exposure and real-time monitoring. Clear acceptable use policies, paired with mandatory employee training, can reduce the volume of sensitive data reaching consumer AI tools through personal accounts. Neither alone is sufficient; the combination matters.
Significant governance and operational readiness gaps persist even as AI adoption accelerates, though the scale of those gaps varies considerably by organisation size and sector.



