OpenAI and Google Reject Publisher Royalty Deal Leaving AI Risk Open

Licensed Data vs Fair Use
Key Takeaways

  • OpenAI’s updated Enterprise Agreement caps legal indemnity at $1 million through revised “Usage Responsibility” clauses, shifting copyright liability directly onto enterprise customers.
  • IBM Granite and Adobe Firefly are winning Fortune 500 procurement deals by offering uncapped legal indemnification backed by fully licensed training sets, a measurable cost advantage over Fair Use-dependent alternatives once insurance premiums are included.

Copyright liability is now a line item in enterprise AI budgets. Collective bargaining between publishers and AI providers collapsed without an industry-wide licensing standard, leaving companies to choose between paying a premium for legally clean models or carrying the litigation risk that comes with the Fair Use defence.

The breakdown is concrete. The Global Press Alliance pushed for a recurring per-token royalty rate covering premium financial and legal data. OpenAI and Google countered with a flat-fee structure the alliance considered insufficient to offset projected traffic losses. No deal was reached.

Criteria for Evaluating AI Data Sourcing

Models trained on fully licensed or public domain data carry a price premium over alternatives that rely on the Fair Use defence. Lower upfront cost is the appeal of the Fair Use path. The tail risk is real, though: if future court rulings require deletion of model weights trained on unauthorised data, enterprises that built on those models face a rebuild, not just a legal bill.

Cost-per-token is no longer a purely technical metric. It now includes what might be called a compliance surcharge that varies by provider. Enterprises building customer-facing agents that synthesise current events cannot rely on static datasets, they need models with live access to news feeds. Without a collective agreement in place, companies must negotiate individual licensing arrangements with publishers directly, a process that adds months to deployment timelines and can run to millions in legal costs.

The Licensed Data Path: OpenAI and News Corp

OpenAI has moved from general web scraping toward high-value bilateral agreements. Its multi-year deal with News Corp is the clearest public example of that shift.

Morgan Stanley’s deployment of a custom GPT-4 instance for its financial advisors relies on licensed internal research and verified external news feeds, according to reports. Operating within a licensed data environment, the firm benefits from the legal guarantees that come with verified training data provenance.

The Fair Use Defence: Perplexity and Open Source

Perplexity’s Pro and Enterprise tiers offer real-time web indexing without individual publisher payments, with paid tiers reported to start at roughly $40 per user per month.

The fragility of this approach is already visible in court. In a recent filing in the Southern District of New York, publishers argued that AI-generated summaries directly substitute for the original work, a factor that weighs against a Fair Use finding. If a court rules those summaries infringing, enterprises using these tools for internal knowledge management could face secondary infringement liability. Professional liability insurers are already responding: policies are adding exclusions for platforms that cannot provide a verified data manifest.

Startups and mid-market firms often favour the Fair Use path for speed and lower cost. A marketing agency using Perplexity to generate competitive intelligence can scale without the overhead of dozens of individual publisher subscriptions. The risk is that output which mirrors source material too closely could trigger a copyright claim, and unlike a large enterprise, smaller firms have limited capacity to absorb that exposure.

Insurance and Indemnity as Differentiators

Legal indemnity has become one of the most scrutinised clauses in AI service-level agreements. Microsoft and Google have updated their terms to offer copyright commitments, promising to defend customers sued over AI-generated output, but those protections are conditional. They require the customer to have used available content filters and guardrails. Disable a filter to get more flexible output, and the indemnity is voided.

Because IBM owns or has explicitly licensed every element of the Granite training set, it offers uncapped indemnity. That is a meaningful contrast to the absence of any indemnity in most open-source Llama 3 implementations.

Policies covering companies using fully licensed models are currently priced at a lower rate per $1 million in coverage. For companies using models with unknown or disputed training data, that figure rises considerably, where coverage is available at all.

Why Training Data Quality Is Diverging

Models using licensed publisher APIs demonstrate improved factual recall compared to those relying on general web search data. For a legal department automating contract review, even a modest accuracy gap separates a viable tool from a liability.

The problem compounds over time. Scraped models are increasingly trained on content generated by other AI systems, creating a feedback loop that degrades factual grounding. Licensed, human-curated data sources sidestep that loop entirely. How far this divergence will widen as litigation narrows the Fair Use path is not yet clear from public evidence.

Comparing Strategy Costs and Risks

A licensed-first deployment using OpenAI Enterprise or Google Vertex AI with licensed add-ons for 500 users can involve substantial annual costs, including seat licenses and API surcharges. A Fair Use strategy using open-source models or Perplexity may look cheaper at baseline. Once insurance premium increases and a legal reserve for copyright disputes are factored in, that cost difference narrows considerably, and that calculation does not include the cost of rebuilding an integration if a court orders a provider to stop serving a specific model.

There is also a market access dimension. European clients operating under the EU AI Act’s transparency requirements are declining to interact with systems that cannot provide a detailed training data manifest. For companies with significant European operations or those selling into the public sector, the Fair Use path may be commercially viable domestically but effectively closed in key international markets. The EU AI Act’s expanding reach is already reshaping these procurement decisions.

The shift from collective negotiation to fragmented litigation requires a concrete change in how enterprises approach AI procurement. A big-tech logo on a model is not a shield against copyright claims. The practical starting point for any automation project is to request a data provenance report from the vendor. A refusal to provide one is itself informative and should be logged in the corporate risk register.

For mission-critical applications where factual accuracy and legal durability matter, healthcare, law, financial services, the licensed model path is the defensible choice. The premium over Fair Use alternatives is, in effect, an insurance payment against a court-ordered stop-work notice. IBM, Adobe and the licensed tiers of OpenAI and Google are the current reference points for that tier.

The Fair Use path has a legitimate role for internal tools, creative work and non-customer-facing applications where copyright exposure is manageable and perfect factual accuracy is not required. Open-source models remain attractive for those use cases on cost and speed grounds. Even there, developers should implement source attribution controls to prevent verbatim reproduction of copyrighted text, the most common trigger for litigation. If you’re building agentic workflows where data provenance matters, the RAG consistency patterns covered here are worth reviewing alongside your licensing strategy. The question for enterprise leaders is no longer whether to treat AI compliance as a discipline, but how much of that cost to carry as fixed overhead. For more on AI agents and automation tools, visit our AI Agents section.

Riley Cross
Riley Cross

Riley covers AI agents, workflow automation, and the tools building the autonomous future of work. With a focus on practical deployment, Riley helps builders and operators understand which agentic frameworks and platforms are actually worth using.

📰 Journalists welcome — cite Auton AI News with attribution. Press & Media → | press@autonainews.com