What the longevity industry sells you — and the handful of it that holds up.
Read the two grades separately. An intervention can be likely effective for healthspan and unproven for lifespan at the same time — those are different claims, tested by different trials, and something can have good evidence that it improves how you function while never having been tested against death at all. The same molecule also appears twice wherever the claim for people with a disease and the claim for healthy people diverge. Not medical advice.
Lifespan and healthspan carry the same five verdicts, judged separately. Risk is a different axis and uses its own three: Minor, Serious, Unknown.
| Intervention | Lifespan | Healthspan | Risk | Type | Rung | Promoted by |
|---|
999 citations retained from the consensus pipeline that produced the previous version of this table. They are keyed by that pipeline’s own row ids and are preserved verbatim.
Every row sits on the rung of its best human evidence, and the rung is a ceiling: it caps how strong a verdict the row can carry, however striking the underlying result.
| Rung | What sits here | Ceiling it imposes |
|---|---|---|
| 6 | ≥2 independent human all-cause-mortality RCTs, or 1 very large multicentre RCT + confirmatory data | Only rung that can produce Likely effective for lifespan (plus the behaviour exception) |
| 5 | 1 good human all-cause-mortality RCT, or high-quality target-trial observational all-cause mortality | Possibly effective for lifespan at best |
| 4 | Human hard healthspan / multimorbidity-composite RCTs | Can be Likely for healthspan; Possibly for lifespan if the composite is dominated by lethal events |
| 3 | Human RCTs on validated function / disease incidence / cause-specific death with all-cause not going the wrong way; good conventional cohorts; best cis-MR for a product | Possibly for healthspan |
| 2 | Clock and biomarker RCTs, short metabolic RCTs, ordinary longevity cohorts, ordinary biomarker MR | Unproven for lifespan or healthspan |
| 1 | Mammalian lifespan (NIA ITP mice), primate biomarker studies | Hypothesis generator only |
| 0 | Worms, flies, cells, mechanism slides, n=1 self-experiment screenshots | Unproven |
Three exceptions. Foundational behaviours — stopping smoking, stopping heavy drinking, exercising rather than not, eating a recognised pattern rather than a poor one — can reach Likely without the top rung, because no ethics committee will ever randomise them. Treating sleep apnoea does not qualify: it has been randomised, and the event trial was null. Replacement medicine — levothyroxine for a failed thyroid — is graded Likely for healthspan on the same never-to-be-randomised basis. Disease drugs are Likely effective in the population their trials enrolled; the same molecule taken by a healthy person is a different claim and drops rungs.
The same eight steps run for every row, in order. Most rows stop at step 7.
Who is taking it, what they are taking, what they are compared against, and which outcome. Change any one of those and the grade can flip — which is why one molecule can occupy two rows with opposite verdicts.
Animal work never sets a grade. It appears in a row only where there is no human evidence at all, and it caps the row near the bottom of the ladder.
A trial that counted deaths sits high whether or not it found a benefit. This is the step people skip: a spectacular biomarker result is still a biomarker result.
The rung caps the verdict. Biomarker evidence cannot produce Likely effective however striking it looks, because moving a marker has repeatedly failed to move the disease.
A drug that saves lives in people who already have the disease is making a different claim in healthy people. That claim drops rungs, and usually several.
Against an active comparator, no difference is not no effect. Against a placebo nobody took, a difference is not necessarily the drug.
If an adequately powered trial measured the outcome and found nothing, the verdict is Ineffective. If no such trial exists, it is Unproven — unsupported, not refuted. Most of this table is Unproven.
They are different endpoints tested by different trials. Something can improve function convincingly and never have been tested against death at all.