Copyright Wars: How News Publishers Are Licensing Content to AI Companies
Some publishers signed. Some sued. Both are negotiating strategies for the same problem: an industry that spent twenty years giving content away for traffic is trying to price it for a buyer that does not send traffic back.
For twenty years the publishing industry gave its content to platforms for free in exchange for traffic. The bargain was never good, but it was legible. Generative systems broke it in a specific way: they consume the content and return an answer, not a visit. That is why licensing negotiations between publishers and AI developers are less about copyright doctrine than about pricing a customer who does not restock the shelf.
What is actually being sold
Two distinct products get bundled under the word "licence", and conflating them is the most expensive mistake a publisher can make at the table.
Training corpus. Historical archive used to build model capability. It is a one-time-ish asset with diminishing marginal value — the tenth newspaper archive teaches a model far less than the first. Publishers negotiating on archive alone are selling a depreciating good.
Live retrieval. Ongoing access to current reporting for grounding answers, with citation and linking. This is the recurring product. It is worth more, it renews, and it is the only part of the relationship where the publisher retains leverage after signing, because tomorrow's news does not exist in yesterday's corpus.
A deal that prices the archive generously and the live feed cheaply looks attractive in year one and terrible in year three.
Why the European position is stronger
European publishers negotiate on firmer legal ground than their American counterparts, for two reasons rooted in EU copyright law.
The first is the press publishers' right introduced by the Digital Single Market Directive, which gives publishers a specific neighbouring right over online use of press content by information society service providers. It was designed for news aggregation, and its application to generative retrieval is contested, but it exists and it has to be negotiated around.
The second is the text and data mining regime. EU law permits mining of lawfully accessible content for commercial purposes unless the rightholder has expressly reserved it in a machine-readable form. That reservation is the practical lever: a publisher that has properly and demonstrably opted out has something to trade. A publisher that has not has already given it away by default. Getting the technical reservation right — not merely a line in the terms of service — is the precondition for any serious negotiation.
Sue, sign, or both
Publishers have split into visible camps, and the split is less philosophical than it appears. Some signed multi-year agreements covering archive and live access. Others filed suit alleging unlicensed use of their work. A third group did both — litigating over past use while negotiating future terms.
The third position is the coherent one. Litigation establishes that the content had a price; licensing collects it. Publishers who ruled out legal action early found their licence terms drifting downwards, because the counterparty had no reason to fear the alternative.
| Deal term | Publisher-favourable | Warning sign |
|---|---|---|
| Scope | Named models and products | "Current and future services" |
| Duration | Two to three years with renegotiation | Long term with fixed fee |
| Attribution | Visible citation and working link | Attribution "where feasible" |
| Reporting | Query volume and citation counts | No usage reporting at all |
| Exclusivity | Non-exclusive | Exclusive or right of first refusal |
| Downstream rights | No sublicensing | Broad sublicensing to affiliates |
| Termination | Deletion or retraining obligation | Perpetual rights to trained weights |
The substitution problem nobody has solved
Every licence signed so far shares one unresolved weakness. If an assistant answers a reader's question using a publisher's reporting, the reader has no reason to visit the publisher. The licence fee replaces advertising and subscription revenue that the visit would have generated — and it is being priced today, against traffic patterns that are still changing.
This is why per-query economics matter more than the headline number. A fixed annual fee is a bet that referral traffic will not fall much further. Usage-based terms, citation guarantees and reporting obligations are a hedge. Publishers with strong direct subscription businesses can afford the bet; publishers dependent on search-driven advertising cannot.
What to do before negotiating
Three preparations change the outcome materially. Implement a machine-readable rights reservation properly, and document when it took effect — this determines what you are entitled to charge for retrospectively. Separate archive and live access into distinct commercial lines so they can be priced and renewed independently. And define internally what your content is worth per thousand queries, because the counterparty will arrive with that number and you should not hear it for the first time in the room.
Note on data. Deal terms described here reflect patterns reported publicly across multiple agreements; specific commercial terms are generally confidential and figures circulating in trade press are frequently estimates. Legal points summarise the EU framework for orientation and are not legal advice — national implementation of the directive varies and litigation is unresolved in several jurisdictions.
Sources
The claims in this article rest on the documents below. Each is linked to what it establishes, so you can check any statement against its origin rather than taking ours for it.
- OpenAI, partnership with Axel Springer (December 2023) — the first publisher deal, announced by OpenAI itself — training, summarisation and attribution, including paywalled titles
- OpenAI, global news partnerships with Le Monde and Prisa Media (March 2024) — the French and Spanish agreements, again in OpenAI's own words rather than a press report
- Columbia Journalism Review, on the split between licensing and litigating — the counter-position: why some publishers signed and others sued instead
Frequently asked questions
Why are news publishers licensing content to AI companies?
Because generative systems consume reporting without sending readers back, breaking the traffic-for-content bargain that funded digital publishing. Licensing converts an uncompensated use into a revenue line, and for publishers with strong archives it is currently one of the few genuinely new sources of income.
What does the EU text and data mining rule mean for publishers?
EU law permits commercial mining of lawfully accessible content unless the rightholder has expressly reserved that use in a machine-readable form. Practically: if you have implemented a proper technical reservation, you have something to license. If you have not, the default position works against you.
Is suing better than licensing?
They are complements rather than alternatives. Litigation establishes that unlicensed use carries a cost; licensing monetises the position it creates. Publishers who publicly ruled out legal action generally negotiated weaker terms, because the counterparty faced no downside from waiting.
What should a publisher never concede in an AI licence?
Perpetual rights to content already absorbed into trained model weights, broad sublicensing to unnamed affiliates, and scope language covering "future products". Each converts a two-year commercial arrangement into an unbounded one at a fixed price.
Do AI licences replace lost search traffic?
Not yet, and not reliably. Reported fees are meaningful for large archives but small relative to the advertising and subscription revenue historically generated by referral visits. Usage-based terms with citation and linking obligations hedge that gap better than a flat annual fee.