News content licensing for AI

News content licensing for AI can be a tricky proposition. It’s undoubtedly a big business–according to Rob Kelly, there were 48 licensing deals including news/journalism as of June 2026, making it by far the largest category of deals. But not all news content is created equal. In a world where a single week can often feel like it contains a year’s worth of content, a dated archive can be nearly worthless. A news archive that ends on December 31, 2025, for example, will say nothing about the military conflict between the US and Iran even though that has consistently dominated the headlines for months throughout 2026.
Training rights and grounding rights are not the same purchase
Licensing news content for AI tends to cover one or both of the following:
- Training rights are used for the kind of foundational work that goes into creating a model.
- Real-time display or grounding rights enable a model to deploy live material inside a product (think Google’s AI summaries).
Many data-purveyors deliberately restrict access to their most current material. For example, Thomson Reuters only licenses the text of its archive while specifically excluding live coverage and other up-to-the-minute material.
Many deals bundle both. Consideration isn’t just monetary: Publishers including the Associated Press, The Atlantic, Vox Media, and Axios have all received access to OpenAI's technology and product expertise as part of their agreements, using it to build their own tools and features.
That sits on the same side of the ledger as the payment, not alongside the rights being granted.
The grounding side is where the growth is. Kelly's tracking shows deals covering attribution and live access rising from 2 in 2023 to 11 in 2024 and 18 in 2025, with roughly 34 projected for 2026. While this is a projection rather than a census, trends do suggest that buyers increasingly want an ongoing feed rather than a one-time archive dump.
Suppliers have also noticed and many segment their catalogs accordingly. Thomson Reuters CEO Steve Hasker described Reuters' approach at a London industry summit in May 2026. Their agreements focus on the text archive, price it as high as the market will bear, and are kept deliberately short in duration so the terms quickly come back up for renegotiation. Hasker was blunt about the gap that leaves. Reuters breaking news still surfaces in chatbots without the company being paid for it. The lesson for buyers is that licensing content doesn’t always guarantee freshness.
What to check before signing
As always, doing your due diligence is key.
Archive or live feed? Confirm this explicitly, in the contract. This is the single most common source of mismatch between what a buyer thought they bought and what they actually receive.
Freshness cadence. How quickly does new content reach you? Is it minutes, hours, or days? For grounding use cases this determines whether the license is usable at all.
Corrections and retractions. When a publisher corrects or retracts a story, does that propagate automatically through the feed, or do you find out some other way? This is close to unique to news, and it is rarely asked about. A system that keeps serving a retracted claim is a liability that no amount of licensing cures.
Attribution obligations. What exactly are you required to display, and where? Deal terms increasingly specify this.
Territoriality and exclusivity. Does the grant cover the markets you operate in?
Format and metadata. Structured, machine-readable delivery via API or JSON, versus files you have to normalize yourself.
Why news carries unusual litigation exposure
News is the most contested category in AI content, and the volume is substantial. As of March 2026, the running count of US copyright suits against AI companies stood at 91. The consolidated multidistrict litigation against OpenAI in the Southern District of New York, before Judge Sidney Stein, folds in more than a dozen publisher suits led by The New York Times. Press Gazette's tracker also covers CNN, the Chicago Tribune, Yomiuri Shimbun, and a coalition of more than 30 US local publishers owning nearly 400 titles between them.
Two developments should concern buyers specifically.
The disputes now reach past training into live retrieval. When Encyclopedia Britannica and Merriam-Webster sued OpenAI in March 2026, the complaint covered not only mass copying for training but also retrieval-augmented generation. It added Lanham Act trademark claims over ChatGPT generating invented answers and attributing them to Britannica. These are reference publishers rather than news organizations, but it shows how the grounding layer is being litigated on its own terms.
Intermediaries are in scope too. News Corp filed copyright counterclaims against Brave in July 2026, inside a declaratory-judgment case Brave had brought preemptively in March 2025. News Corp alleges Brave masked its crawlers so publishers could not detect or block them, and resold News Corp content to AI companies competing with News Corp's own licensing programs. If that theory holds, sourcing news content through a data middleman does not insulate the buyer downstream.
Unlicensed access to news carries materially higher exposure than other content categories, and provenance documentation now matters as much as the license itself.
Sourcing news content for AI
For teams sourcing news and editorial content for AI systems, Newstex supplies rights-cleared, provenance-documented content from a network of vetted publishers, delivered as structured feeds with the metadata and documentation your legal and compliance reviews will ask for. Request a demo to see the catalog, delivery cadence, and rights structure.


