Methodology

What this actually is, precisely

The single rule this project is built around: never claim more precision than the evidence supports. This page exists so that rule is checkable, not just asserted.

1. The corpus

52,408 trials registered on ClinicalTrials.gov as terminated, withdrawn, or suspended were machine-classified into a 19-code taxonomy describing why each one actually stopped — separating scientific failure from operational reasons (funding, staffing, business strategy, external events) and outright non-failures. The corpus is complete: zero missing, zero duplicates, independently re-verified.

2. No human ground truth exists

No human-labeled ground truth exists yet. Every figure above is machine classification, cross-model agreement, or model-adjudicated — never validated accuracy. This site does not round that off.

What stands in for it: cross-model agreement, model self-consistency (91.6% code-level agreement with itself, n=1400), and — the strongest signal — two independent, non-overlapping external checks that converge on the same number. Two independent methods — CT.gov posted-results data and blind Opus adjudication — converge on ~0.2% error for non-contested records.

3. Contested records are excluded from headline claims

Records where the classifier's own secondary code disagrees with its primary code run a 44% false-positive rate under blind adjudication. Non-contested records run 0.2%. Every headline number on this site excludes contested records for exactly this reason.

4. The 10-case pilot is the only web-verified sample

The corpus-level numbers above are machine classification. The 10 Evidence Briefs are a different, stronger kind of claim: each one was checked by hand against live web search and primary sources — PubMed, journal publishers, the ClinicalTrials.gov API directly, and news archives — not just re-read from the registry's own text. That distinction matters and this site does not blur it: corpus-wide numbers are classification; the 10 briefs are verification.

5. What this project got wrong, and kept

The classification pipeline has been through 7 audit rounds. Several were rounds of fixing real bugs the project introduced and then caught itself — a false-positive regex that flipped real records the wrong way, a re-classification prompt that led its own witness and overstated an error rate by ~25%, a cost estimate off by 3x from unverified hardcoded pricing. 22 corrections were later found wrong and retracted. All of that is tracked in the source repository's commit history, on the theory that a verification project that hides its own errors isn't one worth trusting.

6. What “confidence: high / medium” means here

It is a two-level qualitative judgment made by the person who did the verification, not a statistical estimate and not a percentage. Where a numeric probability would imply false precision, this site deliberately does not compute one.