The Longevity Table
I have put a table on this site. It grades 152 things people are sold to live longer — supplements, drugs, procedures, diagnostic tests and behaviours — against whatever human evidence exists for each one.
You can read it here.
Unproven is not the same as ineffective
Unproven means nobody ran the study that could answer the question. The claim is unsupported, not refuted. Ineffective means somebody did run it and the answer was no.
131 of the 152 rows are unproven. 7 are ineffective. That ratio is the finding: the longevity industry is not mainly selling things that failed, it is mainly selling things nobody tested. The two get collapsed into “no evidence” constantly, and they call for opposite responses.
Every row sits on the rung of its best human evidence, and the rung is a ceiling. A striking biomarker result cannot produce a strong verdict, because moving a marker has repeatedly failed to move the disease. Population is a ceiling too: a drug that saves lives in people who already have the disease is making a different claim in a healthy person. Metformin therefore occupies two rows — possibly effective for type 2 diabetes, unproven for longevity, on evidence three rungs apart.
9 rows reach likely effective for lifespan. Four are behaviours available to anyone, five are drugs for people who already have the condition they treat, and none is a supplement.
How it was built
The audit is machine-run and adversarial. One model reads the literature and grades each claim; a second model attacks the result — wrong rung, stale evidence, contradicted by the table’s own rules — and the first either concedes with a citation or refuses with one. That runs for round after round until neither side is finding anything new.
The arguing is the point, because it is checkable. Every claim either resolves to a paper or it does not. One model asserted a randomised fisetin trial with a sample size and three null endpoints; eight searches could not find it, and it turned out to be a conference abstract described as a journal article. Another round produced two real trials the table had missed, both of which made a grade worse. A third insisted a mortality figure was unciteable; it was, in a follow-up paper — all-cause mortality 0.81, though most of the benefit was fewer deaths from infection rather than from heart disease, which is not how the drug is sold.
That is the useful property. Two models with different training and no institutional stake will not make the same mistake in the same direction, and every disagreement between them ends at a primary source. What survives is not consensus about the evidence. It is the part neither model could argue away.