AI Companies Are Paying Two Different Prices for the Same Books: License or Settlement
The content licensing market now has enough disclosed deals to compare, and the comparison produces an odd result. There are two prices for training material, they differ by an order of magnitude, and which one you pay depends less on what you used than on when you decided to ask.
The license price is visible. OpenAI has signed roughly two dozen publisher and data agreements, the largest reported at two hundred and fifty million dollars over five years with News Corp. Reddit disclosed two hundred and three million dollars in aggregate data licensing contract value in its listing documents. Perplexity took a different route entirely, running a revenue-share program for publishers whose material it cites, backed by a reported pool in the tens of millions.
The settlement price is also visible now. Anthropic agreed to a proposed one and a half billion dollar class settlement over books obtained from shadow libraries, an amount described as among the largest copyright settlements in American history, covering past conduct rather than granting any forward license. Music publishers have filed a separate action seeking more than three billion.
So one company can license a major news archive for a few hundred million across five years while another faces a billion-plus for material acquired without asking. The gap is not a valuation of the content. It is a valuation of the acquisition method, and the courts are increasingly signaling that this is the axis that matters: not whether training on lawfully obtained material is transformative, but whether the material was lawfully obtained at all.
A German court has already gone further, holding that memorization inside a model and reproduction of lyrics in outputs both infringe, and rejecting the text-and-data-mining exception as a defense. That ruling is under appeal. The first American jury trial on AI copyright is calendared for September, and the European transparency mandate requiring disclosure of training data sources reaches full enforcement this month.
For publishers, the strategic reading is that the licensing window is a function of litigation risk, and litigation risk is currently rising. Deals get signed when the alternative looks expensive. That has been true for the last two years and the settlements are making it more true.
For everyone downstream, the thing to watch is who does not appear on either list. Companies with no licenses and no settlements are not necessarily in the clear. They are unresolved.