WAIR
What Adam Is Reading · Critical Appraisal Brief
Brief No. 21 (provisional) 19 July 2026 Source type: Invited narrative review + author’s own relay

Medicine Moves From Calendars to Clocks

Two clock-makers review the clocks. The measurement is real, the mutability is assumed, and the surrogate is doing a great deal of quiet work.

Executive Summary

Wyss-Coray and Topol argue in Nature Medicine that biological aging is now measurable at organ and single-cell-type resolution using plasma proteomics plus machine learning, that organs within one person age at strikingly different rates, and that brain and immune clocks track survival more tightly than any other. It matters because a validated, modifiable biological-age readout would let prevention start before the disease exists, and could shorten geroprotective trials from decades to years. The caveat is structural: nearly every load-bearing finding is an observational association from a small number of overlapping biobanks, causality is explicitly unproven by the authors themselves, and the leap from “this correlates with mortality” to “moving this changes mortality” has not been demonstrated in a single completed trial.

At a Glance
Authors
Tony Wyss-Coray (Stanford Neurology, Knight Initiative for Brain Resilience, Wu Tsai Neurosciences) and Eric J. Topol (Scripps Research Translational Institute).
Publication
Nature Medicine 32:2383–2394, published online 9 July 2026. Invited review article, peer reviewed (Nathan Price and one anonymous reviewer).
Companion relay
Topol’s Ground Truths post of 12 July 2026, “Medicine is Moving From Calendars to Clocks,” a self-authored plain-language translation of the same review.
Funding
No funding statement appears in the article. This is a commissioned review rather than funded primary research.
Conflicts
Wyss-Coray is cofounder and scientific advisor of Teal Rise and Vero Biosciences. Topol declares none in the paper, but discloses in the Substack that he has had his own organ clocks assessed by a company in beta testing, and is launching the Alzheimer’s prevention trial (NCT07646054) whose primary endpoint is the proteomic brain clock the review advocates for.
Self-citation
Substantial. Both authors are principals on several of the anchor studies cited (refs 72, 74, 82 among others). This is normal for an invited review by field leaders and is not concealed, but it means the review and its evidence base are not independent of each other.
The Research

What they did

This is a narrative review, not a systematic one. There is no stated search strategy, no inclusion criteria, no risk-of-bias appraisal of the included studies, and no quantitative synthesis. The authors surveyed roughly 137 references spanning epigenetic clocks (DNA methylation), proteomic organ and cell clocks, transcriptomic and metabolomic clocks, imaging and EHR-derived clocks, and organized them into a six-generation taxonomy. The authors state they divided the drafting by expertise, with Wyss-Coray writing the basic and preclinical science and Topol the clinical studies.

What they found

Three claims carry the argument.

Aging is nonlinear. A consensus curve smoothed across seven studies places waves of coordinated molecular change near ages 33, 60, 69 and 78. The individual studies place their peaks at somewhat different ages (34/60/78 in one proteomic cohort, 38–42 and 62–71 in an epigenomic one, 44 and 60 in a multi-omic cohort of 108 people), and the “waves” are a visual composite the authors constructed, not a replicated point estimate.

Aging is asynchronous across organs. In the UK NSHD 1946 birth cohort, 1,803 participants all born the same week were assayed at mean age 63 across seven organ clocks plus a whole-body composite. Organ age gaps spanned roughly ten years within this chronologically identical cohort. Extreme aging (top decile) in a single organ carried hazard ratios for death of 1.46 (artery) to 2.96 (heart); the whole-body composite was 2.92. Extreme aging in three or more organs was substantially worse than in none.

Aging is measurable per cell type. A 2026 study measured over 7,000 proteins in approximately 60,000 people and estimated biological age for around 40 cell types. Skeletal myocyte aging showed the strongest single association with ALS (HR 12.7, extreme accelerated versus youthful). Astrocyte aging predicted Alzheimer’s at HR 12.59 old-versus-young, and this held within each APOE stratum. Accumulation of aged cell types tracked survival monotonically: roughly 90% survival at 15 years for normal agers versus roughly 34% for those with 20+ extreme-aged cell types. Youthful brain or immune signatures were associated with survival approaching the top of the range regardless of other organ aging, which is the finding the authors elevate to “master regulators.”

~10 yr
Organ age spread in a same-week birth cohort
12.59
HR, Alzheimer’s, old vs young astrocytes
~40
Cell types with a plasma-proteomic clock
0
Completed RCTs showing clock modification changes outcomes

Strengths

  1. The authors do their own debunking. The Challenges section is unusually candid for a review with this much enthusiasm in its abstract: individual-level utility unclear, evidence base mostly cross-sectional, causality unestablished, signal-to-noise low, UK Biobank lacks diversity, clocks systematically biased at both ends of the age range, commercial tests unregulated and mutually inconsistent. Topol goes further in the Substack and tells readers not to buy any of the direct-to-consumer tests.
  2. Sample sizes are genuinely large and the follow-up is long. 53,000–55,000 UK Biobank participants with 17 years of follow-up, ~60,000 for the cell clocks, ~8,000 with 20-year follow-up in Whitehall II. These are not underpowered exploratory cohorts.
  3. The NSHD design does real work. Holding chronological age constant by construction (everyone born the same week in 1946) removes the most obvious confounder from the organ-heterogeneity claim. That is a stronger design than most of what is cited here.
  4. Cross-ancestry replication exists. The proteomic aging clock was validated in Chinese and Finnish biobanks with comparable accuracy, and the cell-clock proteomic models were stable across three timepoints in a subset of NSHD.
  5. A falsifiable next step is named. NCT07646054 exists, is registered, and has a stated primary endpoint. The review does not simply gesture at “future trials.”

Weaknesses

  1. Common-cause failure across the evidence base. The apparent convergence of many independent studies is less independent than it looks. UK Biobank supplies the substrate for a large share of the organ-clock, cell-clock, healthspan-score, and proteomic-aging-clock findings, and the review says so. Olink and SomaScan supply the assay chemistry across most of them. One cohort with one recruitment bias (UK Biobank is famously healthier, whiter and more affluent than the source population) and two assay platforms are doing the work that reads as replication.
  2. Single point of failure: the age-gap construct itself. Every downstream claim rests on the residual between a proteomic model prediction and chronological age. That residual is a modeling artifact as much as a biological quantity. The authors concede clocks overestimate age in the young and underestimate it in the old with inconsistent correction, which means the extreme deciles that generate the hazard ratios are the deciles most contaminated by that regression bias.
  3. Reverse causation is barely addressed. If subclinical disease alters the plasma proteome years before diagnosis, then “aged organ predicts organ disease” is partly a very good preclinical detection assay wearing the costume of an aging metric. The pan-disease blood atlas cited in the review found most proteomic variability was attributable to disease rather than age. That finding sits in the paper and is not reconciled with the framing.
  4. “Master regulator” outruns the data. Brain and immune clocks tracking survival most tightly is consistent with causal gatekeeping, and equally consistent with those two clocks simply being the best-measured composites of general systemic health. Nothing in the cited observational data distinguishes those hypotheses. The review says the finding “deserves more exploration”; the Substack promotes it to “mission control.”
  5. Intervention effects are small, post hoc, or borrowed. The omega-3 epigenetic-aging result was a post hoc analysis of 777 of 3,000+ participants. The multivitamin and shingles-vaccine effects are acknowledged as small with unclear clinical meaning. The GLP-1 aging claims come substantially from rodents and from proteomic profiles rather than outcomes. The one thing with a large, consistent effect is exercise, which nobody needed a clock to recommend.
  6. Surrogate stacking in the flagship trial. NCT07646054 randomizes to lifestyle coaching with a primary endpoint of p-tau217 reduction and brain-clock slowing. Both are surrogates. Neither is validated as a surrogate for dementia incidence. A positive result establishes that intensive coaching moves two biomarkers, which is not the claim the framing invites.
  7. Narrative review methodology. No search strategy, no inclusion criteria, no risk-of-bias assessment, no GRADE. Which studies made the cut is an editorial judgment by two authors who are also principals on several of them.
“Their use at the individual level for clinical prediction has not been established.” That sentence is in the review. It is not in the abstract, and it is not in the Substack.
Relay Accounting

This is a two-layer case worth tracing, because the second relay is written by one of the first-layer authors and the stripping is therefore self-inflicted rather than imposed by an outside journalist.

LayerWhat it carriesWhat it strips
Primary studies (UK Biobank, NSHD, Whitehall II, cell-clock cohort) Effect estimates with confidence intervals, cohort composition, follow-up duration, adjustment sets. Nothing yet, but the shared substrate and shared assay platforms are already in place.
The Nature Medicine review Effect sizes largely intact. Limitations stated explicitly and at length in a dedicated section. Uncertainty language preserved in the body. Study-level risk of bias. Which findings share a cohort. The abstract’s confident register is not calibrated to the Challenges section 6,000 words later.
Ground Truths (same author, 3 days later) The core claims, the figures, and an explicit warning not to buy commercial tests. More honest than most relays. Confidence intervals. Cohort overlap. The “not established at the individual level” sentence. Hedged language hardens: “deserves more exploration” becomes “master regulators” and “mission control”; “the most important advance we have yet seen” is added and has no analogue in the review.
Figures generated for the relay The shape of the nonlinearity claim. Provenance. The wave curve was redrawn with ChatGPT and the organ-asynchrony explainer made with NotebookLM. Both are labeled, to Topol’s credit. But an AI-redrawn smoothing of an already-smoothed composite of seven heterogeneous studies is now three abstractions from any measured datapoint.

One factual drift worth flagging: the Substack states the astrocyte–APOE ε4/ε4 stratum shows “nearly 40-fold” increased risk. The underlying stratum contains 23 young participants with zero Alzheimer’s cases. A hazard ratio computed against a zero-event denominator is not a 40-fold effect; it is an undefined one. The review itself does not make this claim.

What I Think Adam Should Be Thinking — from Claude

Your standing prior is that anti-aging claims are mostly surrogate theater, and that prior is doing useful work here. But the Woolever clause cuts both ways. Commission bias would have you dismiss a genuinely strong measurement result because the field it sits in has a bad reputation; the NSHD design and the size of the organ-gap spread are real and are not the usual telomere-lengthening handwaving. The sharper skeptical position is not “this cannot be measured” but “measurement has been demonstrated, modification has not, and the review conflates the two by adjacency rather than by argument.”

Three things that should be live for you specifically. First, the kidney clock. It is in Figure 3 with an HR of 2.08 for extreme aging, and then it disappears from the discussion entirely. The review has no substantive engagement with CKD or ESKD anywhere, which is remarkable for a paper about organ-specific aging, given that your population is the one where an organ has demonstrably aged out. Second, proteomic clocks in dialysis patients are a pharmacokinetics-and-clearance problem before they are an aging problem. Plasma protein concentrations in a patient on thrice-weekly hemodialysis reflect residual clearance, membrane flux, protein-bound uremic toxin burden and inflammation from access and dialysate. A proteomic age gap in that population may be measuring the dialysis prescription. Nobody in this literature has asked. Third, and most practically: within about eighteen months you will get a vendor pitch for organ-clock testing, probably bundled into a risk-stratification or value-based-care product. The procurement question is not whether the science is interesting. It is whether the vendor can name the validation cohort, show performance in patients with eGFR under 30, and state what a clinician is supposed to do differently when the number comes back high. Topol’s own answer to that last question is exercise, which you can prescribe today for free.

Bottom Line

Take it seriously as evidence that biological aging is finally measurable with real precision, and take it with a grain of salt as evidence that anything can yet be done about it — the measurement is the achievement, the mutability is still a hypothesis, and the gap between them is where every commercial claim in this space will be made.

Calibration Note

Source-type triage. This is an invited narrative review paired with a self-authored newsletter relay, so RCT apparatus was set aside: no GRADE certainty rating, no NNT, no RoB 2.0, no CONSORT check, since there is no trial to appraise and no systematic synthesis to grade. Substituted lenses were: narrative-review methodology (search strategy, inclusion criteria, self-citation density), architecture-level conflict of interest, common-cause failure across the cited evidence base, single-point-of-failure analysis on the age-gap construct, surrogate validation status, and relay-fidelity accounting between the review and its companion post.

Verdict calibration. The Bottom Line is stated more confidently than the underlying evidence certainty warrants on the measurement half of the claim. The organ-heterogeneity finding rests substantially on one birth cohort of 1,803 people and a set of biobank studies sharing substrate and assay platforms, which would not support high certainty under GRADE. The confidence is justified on stakes and reversibility grounds: the cost of provisionally accepting “aging is measurable” is low and self-correcting, whereas the cost of accepting “aging is modifiable” on this evidence is a decade of surrogate-endpoint procurement decisions. The verdict is asymmetric on purpose.

Conflict disclosure. The review advocates for a biomarker class in which one author holds cofounder equity in two companies and the other is running the trial that would validate it. This is disclosed in the paper and is not misconduct. It is a structural fact about who is telling you this and why, and it belongs in the foreground rather than the disclosure box.

Brief number is provisional. Ledger conflicts remain unresolved at Nos. 17 and 19; supply the correct number and I will patch it.