The Copyright Docket
The foundational input of every frontier model is text and imagery taken at scale — much of it without licence. By mid-2026 that practice had produced the largest copyright settlement in American history, a second nine-figure settlement built on an admission of downloading pirated books, and a still-unresolved trial that could decide whether training itself is infringement.
What's at stake in the docket
Solid bars: amounts actually paid. Hatched bars: claims still unresolved — the amounts shown are what plaintiffs seek or statutory maxima, not findings.
Oct 2023 → Jun 2025
$1.5 billion — and a licensing future
Music publishers alleged Anthropic's Claude reproduced copyrighted lyrics verbatim in outputs. On the eve of trial, Anthropic agreed to pay roughly $3,000 per work across hundreds of thousands of works — the largest copyright settlement ever — plus forward-looking licences. Substantiated Court docket, Jun 2025
Class action · Settled Sep 2025
The pirated-books admission
In court filings, Anthropic admitted downloading roughly seven million pirated books from shadow libraries after commercial sources proved too expensive, then deleting the copies. The class settled for $1.5 billion — again about $3,000 a book. The admission matters more than the money: it is a frontier lab acknowledging, on the record, that part of its training corpus was built on knowingly pirated material. Substantiated Court filings, N.D. Cal.
Filed Dec 2023 · Pending
The case that could decide everything
The Times alleges wholesale copying of its journalism, removal of copyright management information, and ChatGPT "hallucinating" false articles attributed to the Times. In February 2025, Judge Stein let the core infringement claims proceed, rejecting OpenAI's fair-use arguments at the motion stage. Statutory damages could reach $150,000 per work. A trial verdict here — not the settlements — will set the actual precedent for whether training on copyrighted text is infringement. Unresolved S.D.N.Y. docket
Filed Jan 2023
Two courts, two directions
In November 2025 the UK High Court dismissed Getty's main claims, finding the UK training acts and outputs did not infringe on the facts pleaded — a major win for model developers in Europe's largest common-law jurisdiction. The parallel US case continues, alongside suits from artists and a September 2025 class action by the Authors Guild against OpenAI. The law is not converging; it is forking by jurisdiction. Partly substantiated UK High Court, Nov 2025
A settlement is a price, not a verdict. Concord and Bartz resolved liability without any court ruling that training itself is illegal. The fair-use question at the heart of the NYT case remains legally unresolved as of our research cutoff. Anyone telling you the law is settled — in either direction — is selling something.
Extraction Without Consent
The lawsuits are the visible layer. Beneath them sits a broader practice: the default posture of the industry has been to take first and negotiate if caught. Crawlers ignored robots.txt directives (a core allegation in the NYT complaint); opt-outs required creators to submit forms to each lab individually; and licensing markets were built after the data was already in the weights. The Bartz admission shows the extreme end — shadow libraries deliberately chosen because paying was inconvenient.
The counter-argument is real: proponents note that human learners also absorb copyrighted work without paying, that licences for "everything ever written" were practically impossible to clear in advance, and that the settlements now being struck are creating a licensing market that didn't exist in 2022. The EU AI Act's Article 53 — requiring providers to publish training-data summaries and respect opt-outs — is the first major attempt to flip the default. Partly substantiated
Concentration: Who Controls the Compute
Power in the AI era is measured in floating-point operations, and the distribution is more lopsided than almost any market in modern history:
| Layer | Concentration | Basis |
|---|---|---|
| AI accelerator chips | NVIDIA holds an estimated ~90% data-centre GPU share | Analyst estimates |
| Frontier training compute | ~5–7 firms (OpenAI, Anthropic, Google, Meta, Microsoft, xAI) run essentially all frontier-scale training | Epoch AI / analyst estimates |
| Advanced packaging (CoWoS) | TSMC produces virtually all leading-edge AI packaging | Substantiated |
| Capital | ~$725B hyperscaler AI capex planned for 2026 — affordable by perhaps six balance sheets on Earth | Company guidance, Thread 02 |
The bottleneck stack — chips, packaging, power, capital — means that even a perfectly competitive market for models would sit atop an oligopoly of means. Export controls on advanced chips extend that concentration into geopolitics: who computes is now a national-security decision made in Washington. See Thread 04 · Chips
Open Weights vs. Closed Frontier
The sharpest ideological fight inside the industry is whether frontier model weights should be released openly. It pits the two most valuable strategies in AI against each other — and by July 2026 it had become a public clash between Anthropic on one side and Meta and Google on the other.
The closed-frontier case
Anthropic · Dario Amodei, "On the Nature of Frontiers Models", Jun 2025
- Frontier models are dual-use: the same weights that do biology do bioweapons
- Open release is irreversible — you cannot recall a capability
- Safety evaluation and export control are impossible once weights are public
- Democratic oversight requires a chokepoint that can be governed
The open-weights case
Meta (Llama) · Google (Gemma) · White House AI Action Plan, Jul 2025
- Security through obscurity fails; open scrutiny finds flaws faster
- Closed weights concentrate power in exactly the firms already dominant
- Allies need open alternatives to closed US and state-controlled Chinese models
- The US open-source ecosystem is a strategic asset, not a liability
Through 2025–26 the gap between open and closed models narrowed with each release cycle, sharpening the stakes: if open weights reach frontier parity, the closed-lab position rests entirely on the dual-use argument; if a catastrophic misuse traces to leaked weights, the open position may not survive the aftermath. The July 2026 clash — with Anthropic publicly opposing further open frontier releases while Meta and Google shipped open weights — left the industry's governance question explicitly unresolved. Unresolved
Lobbying & Regulatory Capture
Every major AI lab now maintains a federal lobbying operation in Washington — most did not in 2022. Disclosed AI-related lobbying spend has surged to record levels, and the revolving door turns in both directions: senior White House, National Security Council, and congressional staff move into lab policy roles, while lab executives take advisory government posts. Substantiated LDA disclosures, OpenSecrets
The capture question is not bribery; it is epistemic dependence. When the only people who understand frontier systems work for the companies building them, regulators inevitably see the world through company briefings. The late-2025 federal push to preempt state AI laws — Justice Department scrutiny of state statutes and an executive order directing a unified federal framework — was cheered by the industry and attacked by state attorneys general of both parties. Whether that is rational harmonisation or capture depends on which side of the preemption line you stand. Unresolved
Military & Surveillance Use
The industry's founding self-image — consumer chatbots, helpful assistants — has quietly been revised. The major labs have all moved toward defence and intelligence work:
Usage policies revised; the company later acknowledged ongoing conversations with US national-security agencies.
Policy revised to permit military and intelligence applications, with stated red lines on lethal autonomous weapons and mass domestic surveillance.
Joint work on counter-drone systems; part of a broader wave of lab–defence-startup integrations.
Google's 2017 Project Maven sparked employee revolt and withdrawal; the work continued under Palantir, Scale, and Anduril — and is now mainstream across the sector.
The pattern is consistent: red lines are announced, then narrowed, then crossed under the framing of "responsible participation." Each lab argues that democratic AI powers must field the technology or adversaries will. Critics note the same argument was made about every previous weapons system, and that the red lines (no lethal autonomy, no mass surveillance) depend entirely on definitions the labs themselves control. Partly substantiated
On the surveillance side, the through-line is data: models trained on scraped faces and text now power identification and analysis systems sold to governments — the extraction-without-consent problem, scaled up to state power. See also: Thread 07 · Labour
Two Regulatory Worlds
European Union
The AI Act — the world's first comprehensive AI statute — entered into force August 2024 and applies in phases:
- Feb 2025: prohibited practices (social scoring, manipulative AI)
- Aug 2025: general-purpose model obligations, incl. training-data summaries
- Aug 2026–27: high-risk system requirements (timeline under review via the "Digital Omnibus")
- Fines to 7% of global turnover
United States — Federal
No federal AI statute. The 2023 Biden executive order was revoked in January 2025; the current posture is explicitly accelerationist:
- Jul 2025: AI Action Plan — pro-buildout, pro-open-models
- Late 2025: federal moves to preempt state AI laws
- Export controls on chips remain — the one hard lever, aimed at China
United States — States
In the federal vacuum, states act:
- California SB 53 (2025): frontier-model transparency duties
- Colorado AI Act: algorithmic-discrimination duties, effective 2026
- Dozens of state AI bills per session since 2024
- Now the battleground for federal preemption
The result: a company building a frontier model in 2026 faces a comprehensive statute in Brussels, a permissive federal regime in Washington, a patchwork of state laws it may or may not have to follow, and export controls that treat its chips as munitions. Compliance strategy has become a form of geopolitical arbitrage — and the arbitrage currently favours speed over caution. Substantiated
Confidence Assessment
| Claim | Status | Confidence |
|---|---|---|
| Training used unlicensed copyrighted material at scale | Admitted in court (Bartz); settled twice for $1.5B each | Substantiated |
| Training itself is copyright infringement | NYT case pending; UK Getty claims dismissed; no final precedent | Unresolved |
| Frontier compute is oligopolistic | Chip, packaging, and capex concentration documented | Substantiated |
| Industry has captured regulators | Lobbying and revolving door documented; "capture" is interpretation | Opinion |
| Labs are moving into military/intelligence work | Policy changes and partnerships publicly announced | Substantiated |
| Open weights will reach frontier parity | Gap narrowing each cycle; parity undemonstrated | Unresolved |