1 · The Source Hierarchy
Not all sources are created equal, and pretending otherwise is how bad investigations happen. Every factual claim in this site was assigned to the highest tier its evidence supports. A claim cannot be promoted by repetition, rhetoric, or how convenient it would be for our narrative.
A tier-4 claim that later appears in an SEC filing gets re-tiered and re-labelled. A tier-6 opinion never becomes a fact no matter how many people repeat it. The Source Explorer shows every registered source with its tier and date.
2 · Confidence Labels, Explained
Tiers say where evidence comes from; confidence labels say how far it reaches. Five labels are used across the site, always inline with the claim they govern:
A final judicial or official finding, a primary document, or multiple independent confirmations of a specific fact. We state these as facts.
Example: "Anthropic admitted downloading roughly seven million pirated books" — admitted in N.D. Cal. filings.
The core fact is established but its scale, scope, or interpretation is not. We state the core as fact and flag the uncertain portion explicitly.
Example: "circular financing exists" (substantiated) + "$1.65T total scale" (one analysis, methodology-dependent).
A credible basis exists but no final determination has been made — typically live litigation or genuinely open empirical questions. We present both sides and refuse to pick one.
Example: "training on copyrighted text is infringement" — NYT case pending; UK Getty claims dismissed. The law has not converged.
An assertion by a company about itself, reported accurately but not independently verified. Treated as data about what the company says, not about what is true.
Example: any lab's self-reported safety-evaluation results or projected revenue milestones.
Analysis, interpretation, or analogy — including our own editorial framing. Opinions are allowed to be strong; they are never allowed to be unlabelled.
Example: "the railway-mania analogy transfers to AI" — argued in the strongest-cases page, labelled as analogy.
3 · Research Cutoff & Update Pipeline
Cutoff: 1 August 2026, 00:00 UTC. Nothing observed after that moment is reflected in Edition 01. Dates on the site mean "known as of cutoff," not "still true today."
Updates flow through a fixed pipeline so that new editions are reproducible rather than improvised:
research/.
The transcript pipeline lives in pipeline/fetch_transcripts.py and writes to research/transcripts/. The corpus, its fetch statuses, and its limitations are documented honestly in the next section.
4 · A Known Limitation, Documented: YouTube Transcripts
Video testimony — podcast interviews, congressional hearings, conference talks — is a primary source in this domain. Several key statements (Amodei's circular-deal defence, Huang's OpenAI prediction) originate on video. Our pipeline attempted to pull transcripts for a nine-video research corpus automatically, from our cloud build environment:
✗ eBpTG53xUog — no_transcript (YouTube blocks datacenter IPs)
✗ kNjdwBHU1CE — no_transcript
✗ yprtkF9SeXc — no_transcript
✗ f23hv1zXN8I — no_transcript
✗ t-8TDOFqkQA — no_transcript
✗ WcckBmkauBQ — no_transcript
✗ SX63mOl-RgI — no_transcript
✗ FMMpUO1uAYk — no_transcript
✗ 7xPlZUzJbJc — no_transcript
# 0 of 9 transcripts retrieved · 9 video IDs retained in corpus for retry
YouTube's anti-bot systems refuse transcript access from cloud/datacenter IP addresses; the fetch succeeds only from residential connections. Rather than silently dropping video-sourced claims or paraphrasing from memory, we adopted a documented-corpus approach:
- The nine video IDs are retained in
research/transcripts/with their failed fetch status — the corpus is visible and auditable, not discarded. - Claims that originate in those videos are, in Edition 01, cited only where a tier-3-or-better secondary source (press quote, company transcript, written statement) independently carries the same statement — and labelled accordingly.
- Transcripts will be re-fetched from a residential connection and added in the next edition; any claim that then changes or loses its secondary support will be corrected under the policy below.
Any quote attributed to a video source in Edition 01 rests on secondary citation, not on our own transcript verification. We say so wherever it matters. If you have access to the original video and find a discrepancy, that is exactly the kind of report our corrections policy exists for.
5 · Known Gaps — What We Could and Could Not Verify
No access to private deal terms
The most consequential numbers in this story — actual prices, warrants, and conditions inside the NVIDIA–OpenAI, hyperscaler–lab, and CoreWeave deals — live in contracts we cannot see. We cite the disclosed outlines (SEC filings, official announcements) and press reporting on the rest, never the rumour layer.
Every deal figure is traceable to a registered source; undisclosed terms are stated as undisclosed, not estimated.
Some sources paywalled
Key FT, WSJ, Bloomberg, and Information reporting sits behind paywalls. We cite these as tier-4 confirmations with headline-level detail and rely on primary documents wherever the underlying fact can be reached directly.
Where a paywalled claim has a free primary equivalent (filing, docket), the primary source is cited instead.
YouTube transcripts unavailable from cloud IP
As documented above: 0 of 9 corpus transcripts retrieved from our build environment. Video-originated claims rest on secondary citation in Edition 01.
Corpus retained with fetch statuses; residential re-fetch scheduled for the next edition.
Non-US evidence is thinner
SEC filings and US dockets are uniquely transparent. Comparable primary data for Chinese labs, sovereign compute projects, and private non-US companies is sparse, so the geopolitics thread carries proportionally more tier-4/5 sourcing.
Confidence labels are applied more conservatively in that thread, and gaps are flagged inline.
The honest ledger
What we could verify
- The circular-deal structure, from filings and on-record statements
- Both $1.5B copyright settlements and the pirated-books admission
- Hyperscaler capex guidance and backlog figures from earnings
- Job-cut counts and energy-demand growth from named datasets
- Published research on the adoption gap (from the labs themselves)
- The status of every major lawsuit as of cutoff, from dockets
What we could not verify
- Private contract terms behind the disclosed deal outlines
- Whether training is legally infringement — courts haven't decided
- The exact scale of hidden debt (single-analysis estimate)
- Video statements against original transcripts (cloud IP block)
- Chinese and private-company compute figures at primary-source quality
- Anything about the future — which is why the Verdict Lab is a simulator, not a forecast
6 · Corrections Policy
This policy is binding on the publisher, not aspirational. It applies to every page of every edition.
- Report. Send the disputed claim, the page it appears on, and your evidence. Anonymous reports are accepted; the evidence standard is the same either way.
- Triage within 72 hours. The claim is checked against the source register. If the register itself is wrong, that is a pipeline bug and is treated as urgent.
- Classify. Correction (factual error — fixed in place, logged, timestamped). Clarification (true but misleading framing — reworded, logged). Disagreement (competing interpretation — not a correction; published as a labelled counterpoint if it meets the evidence standard).
- Fix and log. Corrections are applied with a visible note on the affected page and an entry in the public corrections log. We do not silently edit.
- Cascade. If a corrected claim feeds a chart, dataset, or the Verdict Lab's inputs, the downstream artifact is rebuilt in the same update — numbers are never corrected in prose while charts keep the old figure.
Corrections that weaken our preferred narrative get the same treatment as corrections that strengthen it. The investigation's conclusion is an output of the evidence register, not an input to it. If the register changes the conclusion, the conclusion changes.