Thread 14 · Transparency · Protocol v1.0

Methodology & Limitations

An investigation is only as trustworthy as its audit trail. This page documents exactly how claims were sourced, weighted, and labelled — and, with equal care, everything we could not verify. If you find an error, the corrections policy below is binding on us.

Research cutoff 1 August 2026, 00:00 UTC Edition 01 Sources registered see Source Explorer Confidence labels 5 tiers Corrections log public

1 · The Source Hierarchy

Not all sources are created equal, and pretending otherwise is how bad investigations happen. Every factual claim in this site was assigned to the highest tier its evidence supports. A claim cannot be promoted by repetition, rhetoric, or how convenient it would be for our narrative.

1
Court judgments & official admissionsFinal judicial findings, settled dockets with admissions, regulator enforcement actions. The gold standard: someone was cross-examined or penalised.
e.g. Concord v. Anthropic settlement; Bartz pirated-books admission; UK Getty ruling
2
SEC filings & regulator findings10-Ks, 10-Qs, S-1s, proxy statements, official regulator reports. Signed under legal liability.
e.g. OpenAI–Amazon $138B AWS commitment (SEC filing); hyperscaler capex guidance
3
Company earnings transcripts & official statementsManagement's own words on the record — reliable as to what was said, not necessarily as to what is true. Company claims are labelled as such.
e.g. Anthropic $30–40B run-rate (company statement, press-corroborated)
4
Multiple independent investigationsTwo or more credible outlets (Bloomberg, FT, WSJ, Reuters…) independently confirming the same fact. Strong, but one step below primary documents.
e.g. $1.65T hidden-debt analysis (Prof G Media / Bloomberg, Jul 2026)
5
Analyst estimatesModelled numbers from named analysts or research firms. Useful for scale and direction; always labelled "estimate" and never presented as measured fact.
e.g. NVIDIA ~90% data-centre GPU share; $5T infrastructure spend by 2030 (HSBC/Goldman)
6
Opinion & commentaryAnalysis, interpretation, analogy. Included because arguments matter — but always visually and verbally labelled as opinion, never smuggled in as fact.
e.g. "this is a bubble" / "this is an industrial revolution" — both positions appear, both labelled
The promotion rule

A tier-4 claim that later appears in an SEC filing gets re-tiered and re-labelled. A tier-6 opinion never becomes a fact no matter how many people repeat it. The Source Explorer shows every registered source with its tier and date.

2 · Confidence Labels, Explained

Tiers say where evidence comes from; confidence labels say how far it reaches. Five labels are used across the site, always inline with the claim they govern:

Substantiated

A final judicial or official finding, a primary document, or multiple independent confirmations of a specific fact. We state these as facts.

Example: "Anthropic admitted downloading roughly seven million pirated books" — admitted in N.D. Cal. filings.

Partly substantiated

The core fact is established but its scale, scope, or interpretation is not. We state the core as fact and flag the uncertain portion explicitly.

Example: "circular financing exists" (substantiated) + "$1.65T total scale" (one analysis, methodology-dependent).

Unresolved

A credible basis exists but no final determination has been made — typically live litigation or genuinely open empirical questions. We present both sides and refuse to pick one.

Example: "training on copyrighted text is infringement" — NYT case pending; UK Getty claims dismissed. The law has not converged.

Company claim

An assertion by a company about itself, reported accurately but not independently verified. Treated as data about what the company says, not about what is true.

Example: any lab's self-reported safety-evaluation results or projected revenue milestones.

Opinion

Analysis, interpretation, or analogy — including our own editorial framing. Opinions are allowed to be strong; they are never allowed to be unlabelled.

Example: "the railway-mania analogy transfers to AI" — argued in the strongest-cases page, labelled as analogy.

3 · Research Cutoff & Update Pipeline

Cutoff: 1 August 2026, 00:00 UTC. Nothing observed after that moment is reflected in Edition 01. Dates on the site mean "known as of cutoff," not "still true today."

Updates flow through a fixed pipeline so that new editions are reproducible rather than improvised:

STEP 01 Ingest Transcript fetcher, filings feeds, and press monitoring pull candidate sources into research/.
STEP 02 Register Each source gets metadata: tier, date, entity, URL snapshot. Unregistered sources cannot be cited.
STEP 03 Claim-check Claims are matched against the register; confidence labels assigned or re-tiered per the promotion rule.
STEP 04 Datasets Chart-ready datasets rebuilt from registered sources only; charts never cite numbers absent from the register.
STEP 05 Edition Material changes ship as a numbered edition with a changelog; minor fixes go to the corrections log.

The transcript pipeline lives in pipeline/fetch_transcripts.py and writes to research/transcripts/. The corpus, its fetch statuses, and its limitations are documented honestly in the next section.

4 · A Known Limitation, Documented: YouTube Transcripts

Video testimony — podcast interviews, congressional hearings, conference talks — is a primary source in this domain. Several key statements (Amodei's circular-deal defence, Huang's OpenAI prediction) originate on video. Our pipeline attempted to pull transcripts for a nine-video research corpus automatically, from our cloud build environment:

$ python pipeline/fetch_transcripts.py # youtube_transcript_api 1.2.4, cloud IP
✗ eBpTG53xUog — no_transcript (YouTube blocks datacenter IPs)
✗ kNjdwBHU1CE — no_transcript
✗ yprtkF9SeXc — no_transcript
✗ f23hv1zXN8I — no_transcript
✗ t-8TDOFqkQA — no_transcript
✗ WcckBmkauBQ — no_transcript
✗ SX63mOl-RgI — no_transcript
✗ FMMpUO1uAYk — no_transcript
✗ 7xPlZUzJbJc — no_transcript
# 0 of 9 transcripts retrieved · 9 video IDs retained in corpus for retry

YouTube's anti-bot systems refuse transcript access from cloud/datacenter IP addresses; the fetch succeeds only from residential connections. Rather than silently dropping video-sourced claims or paraphrasing from memory, we adopted a documented-corpus approach:

What this means for you as a reader

Any quote attributed to a video source in Edition 01 rests on secondary citation, not on our own transcript verification. We say so wherever it matters. If you have access to the original video and find a discrepancy, that is exactly the kind of report our corrections policy exists for.

5 · Known Gaps — What We Could and Could Not Verify

No access to private deal terms

The most consequential numbers in this story — actual prices, warrants, and conditions inside the NVIDIA–OpenAI, hyperscaler–lab, and CoreWeave deals — live in contracts we cannot see. We cite the disclosed outlines (SEC filings, official announcements) and press reporting on the rest, never the rumour layer.

Every deal figure is traceable to a registered source; undisclosed terms are stated as undisclosed, not estimated.

Some sources paywalled

Key FT, WSJ, Bloomberg, and Information reporting sits behind paywalls. We cite these as tier-4 confirmations with headline-level detail and rely on primary documents wherever the underlying fact can be reached directly.

Where a paywalled claim has a free primary equivalent (filing, docket), the primary source is cited instead.

YouTube transcripts unavailable from cloud IP

As documented above: 0 of 9 corpus transcripts retrieved from our build environment. Video-originated claims rest on secondary citation in Edition 01.

Corpus retained with fetch statuses; residential re-fetch scheduled for the next edition.

Non-US evidence is thinner

SEC filings and US dockets are uniquely transparent. Comparable primary data for Chinese labs, sovereign compute projects, and private non-US companies is sparse, so the geopolitics thread carries proportionally more tier-4/5 sourcing.

Confidence labels are applied more conservatively in that thread, and gaps are flagged inline.

The honest ledger

What we could verify

  • The circular-deal structure, from filings and on-record statements
  • Both $1.5B copyright settlements and the pirated-books admission
  • Hyperscaler capex guidance and backlog figures from earnings
  • Job-cut counts and energy-demand growth from named datasets
  • Published research on the adoption gap (from the labs themselves)
  • The status of every major lawsuit as of cutoff, from dockets

What we could not verify

  • Private contract terms behind the disclosed deal outlines
  • Whether training is legally infringement — courts haven't decided
  • The exact scale of hidden debt (single-analysis estimate)
  • Video statements against original transcripts (cloud IP block)
  • Chinese and private-company compute figures at primary-source quality
  • Anything about the future — which is why the Verdict Lab is a simulator, not a forecast

6 · Corrections Policy

This policy is binding on the publisher, not aspirational. It applies to every page of every edition.

  1. Report. Send the disputed claim, the page it appears on, and your evidence. Anonymous reports are accepted; the evidence standard is the same either way.
  2. Triage within 72 hours. The claim is checked against the source register. If the register itself is wrong, that is a pipeline bug and is treated as urgent.
  3. Classify. Correction (factual error — fixed in place, logged, timestamped). Clarification (true but misleading framing — reworded, logged). Disagreement (competing interpretation — not a correction; published as a labelled counterpoint if it meets the evidence standard).
  4. Fix and log. Corrections are applied with a visible note on the affected page and an entry in the public corrections log. We do not silently edit.
  5. Cascade. If a corrected claim feeds a chart, dataset, or the Verdict Lab's inputs, the downstream artifact is rebuilt in the same update — numbers are never corrected in prose while charts keep the old figure.
The asymmetry commitment

Corrections that weaken our preferred narrative get the same treatment as corrections that strengthen it. The investigation's conclusion is an output of the evidence register, not an input to it. If the register changes the conclusion, the conclusion changes.