AI Copyright Lawsuits More Than Doubled: The Free Data Era Is Ending

AI Copyright Lawsuits Tripled: The Free Data Era Is Ending

Three years ago, AI companies read the internet for almost nothing. Today, lawsuits and licensing contracts arrive on the same desk.


Key Takeaways

Copyright infringement cases filed against AI companies more than doubled in 2025 — from around 30 at the end of 2024 to over 70.

70+ — AI copyright cases filed in 2025, up from ~30 in 2024

Meta signed News Corp for up to $50 million a year, and Google pays Reddit around $60 million a year for training data.

Cloudflare will block mixed-use AI crawlers by default on ad-supported pages starting September 15, 2026.


What Happened

CompanyLicensing Commitment
MetaUp to $50M/year (News Corp)
Google~$60M/year (Reddit)

The legal pressure has compounded quickly.

Following The New York Times’ suit against OpenAI and Microsoft, publishers including Japan’s Nikkei and Asahi, CNN, Britannica and Merriam-Webster have all filed.

In June 2026, a coalition of 35 publishers representing nearly 400 local newspaper brands sued OpenAI (its 26th case) and Microsoft (its 11th).

Their complaint noted that the AI boom generated hundreds of billions in market value, and that none of it reached the publishers whose work made it possible.

Korea joined in February 2026, when broadcasters KBS, MBC and SBS filed a copyright suit against OpenAI in Seoul Central District Court — arguing that OpenAI had signed paid licenses with News Corp and other outlets globally while refusing to negotiate with them at all.

Meanwhile the licensing market kept growing.

Meta reversed years of distancing itself from news and signed News Corp for up to $50 million a year in March. Google pays Reddit around $60 million annually for access to its data.

Then infrastructure entered the fight. On July 1, Cloudflare — which sits in front of a large share of all web traffic — announced that from September 15, crawlers mixing search and AI-training purposes will be blocked by default on ad-supported pages.

It is also replacing “Pay Per Crawl” with “Pay Per Use,” which compensates publishers only when their content is actually cited in an AI answer.

The gap between 30 and 70-plus lawsuits

Doubling from around 30 cases to more than 70 in a single year isn’t just a bigger number — it reflects publishers of very different sizes deciding litigation was worth the cost at the same time.

The June 2026 coalition suit is the clearest example. Thirty-five publishers representing nearly 400 local newspaper brands filed together, making it the 26th case against OpenAI and the 11th against Microsoft.

That suggests individual local papers concluded they couldn’t afford to sue OpenAI alone but could pool resources to make the filing happen collectively.

Korea’s broadcasters took a different route in February, filing directly in Seoul Central District Court rather than joining a U.S. case.

They explicitly cited OpenAI’s paid deals with News Corp and other global outlets as the reason they expected — and were refused — similar treatment.

The pattern across both filings is the same: publishers without a licensing deal are the ones suing, while publishers with one, like News Corp and Reddit, are conspicuously absent from the plaintiff lists.

None of this settles whether the underlying legal claim is strong. It shows which publishers judged the fight worth the cost, and which ones apparently already had.


The Two Lenses

AI crawler share of all crawler requests

What the crawl-to-click numbers are actually measuring

Anthropic’s roughly 38,000-to-1 crawl-to-referral ratio and OpenAI’s 1,091-to-1 aren’t the same kind of number, and the gap matters.

Anthropic’s crawler is reportedly pulling far more content per visitor it sends back, which is part of why publishers frame this as extraction rather than a fair trade.

The 39.8% drop in organic clicks from Google’s AI Overviews, found by Agarwal and Sen, is a different measurement again — it’s about what happens after a search, not during a crawl.

Put together, the two data points describe the same publisher complaint from opposite ends of the pipeline: less traffic in, less content credited out.

OpenAI’s fair-use defense hasn’t been tested against either data point in court yet. Both numbers describe traffic and compensation, not the legal question courts haven’t settled.

Cloudflare’s own crawler-share numbers — AI training requests rising from 22% to 52% between spring 2025 and June 2026 — are the infrastructure-side evidence behind both ratios.

Lens one — the exchange was never balanced.

Cloudflare reports AI training crawlers rose from 22% of crawler requests in spring 2025 to 52% by June 2026. The payback was negligible: as of July 2025, Anthropic’s crawler made roughly 38,000 crawls for every visitor it referred back; OpenAI’s ratio was 1,091-to-1.

A field experiment by researchers Saharsh Agarwal and Ananya Sen found Google’s AI Overviews cut organic clicks by 39.8% — and the clicks it removed showed no measurable difference in bounce rate, time on site or return-to-search behaviour. From this view, the correction is overdue.

Lens two — who actually gets paid?

News Corp negotiates eight-figure contracts. A family-owned local paper in Ohio has litigation and little else. The end of free scraping does not automatically mean fair compensation; it may simply mean compensation concentrated among those who already had leverage. OpenAI’s position remains that training on publicly available data qualifies as fair use — and courts have not settled the question.


Why It Matters

Spotify once argued that exposure was its own reward, and only introduced royalty payments after sustained legal and regulatory pressure. The same argument, and the same trajectory, is now playing out between AI labs and publishers.

What makes September 15 notable is that it moves the dispute out of courtrooms and into infrastructure. A lawsuit takes years; a default setting takes effect on a date.

If Cloudflare’s approach holds — and if OpenAI, Google and Microsoft choose to pay rather than route around it — the economics of training a model, and the incentive to publish original work at all, could shift within twelve months.

That “if” is doing considerable work. Worth watching closely through September.


FAQ

Q. Does this mean ChatGPT or Gemini will get worse?

A. Not immediately. But if quality content carries a price long-term, it could reshape which sources AI companies prioritize in training data.

Q. Do licensing deals stop publishers from suing?

A. No. Several publishers have signed with one AI company while litigating against another, treating deals and lawsuits as parallel strategies rather than alternatives.

What would change our view

My read is that this correction favors publishers with leverage — News Corp, Reddit — over the local papers that just sued as a group of 35.

That would change if licensing terms comparable to News Corp’s start reaching smaller, independently owned outlets rather than staying concentrated among the largest names.

It would also change if Cloudflare’s September 15 default gets widely bypassed or renegotiated by major AI labs choosing to route around it rather than pay.

Sources

Related from 2mind

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *