What Adam Is Reading
The Number We Built the Field On
A foundation model trained on ten thousand sleep studies found five risk groups and a clean mortality gradient. The apnea-hypopnea index, which is how sleep medicine diagnoses, bills, and reasons, predicted nothing.
Single source review · Nature Communications · Open access · August 4, 2026

Sleep medicine has a number. It is called the apnea-hypopnea index, it counts how many times an hour you stop breathing, and essentially everything downstream of a sleep study runs on it. Diagnosis runs on it. Severity categories run on it. Insurance coverage for a CPAP machine runs on it. It is the number.

A team from IBM Research and the Cleveland Clinic trained a foundation model on more than ten thousand clinical sleep recordings and let it find its own structure. It found five patient groups with a clean mortality gradient across them. The apnea-hypopnea index showed no significant association with mortality at all.

That is the paper. Everything else is detail.


What they did

A polysomnogram is an extraordinarily rich object. Overnight EEG, ECG, airflow, oxygen saturation, effort belts, leg movement, all sampled continuously for hours. It is one of the densest physiologic recordings routinely collected in medicine.

We then take that recording and compress it, mostly, into one integer.

The team trained a foundation model to learn representations directly from the raw multichannel signal rather than from the scored summary. They then clustered patients in that learned embedding space, without reference to any outcome, and asked afterward what happened to the people in each cluster.

Five groups fell out. Ordered from lowest to highest, they show monotonic increases in cardiovascular events, neurologic outcomes, and death. Group five carried roughly a 2.71 fold mortality hazard against group one, along with elevated risk of heart failure, atrial fibrillation, and cognitive impairment.

Then the comparison that gives the paper its teeth. Ranked against those same outcomes, conventional apnea-hypopnea severity categories did not significantly predict mortality.

Why the external validation matters more than usual. The findings replicated in the Sleep Heart Health Study, an independent cohort with older and lower resolution recordings. A learned representation that only works on the equipment it was trained on is a lab curiosity. One that survives a transfer to worse data is describing something about the patient rather than something about the machine.

It is not just the breathing

The obvious objection is that the model quietly rediscovered a better breathing metric and the whole thing is apnea severity with extra steps. The authors ran sensitivity analyses against exactly this and report that the model integrates EEG and ECG signal alongside respiratory channels rather than leaning on breathing alone.

If that holds, the finding is considerably more interesting than a better apnea score. It suggests that what a sleep study captures about your risk of dying is a whole body phenomenon, distributed across brain activity and cardiac rhythm and respiration together, and that we have spent forty years reading one channel of a multichannel signal and discarding the rest.

Which, to be fair, is what you do when you have to score studies by hand. The apnea-hypopnea index is not a stupid number. It is a number designed for a technician with a pen, and it has been kept long past the arrival of machines that do not need the simplification.


The part that will not move

Suppose all of this is exactly right. Consider what would have to happen next.

The apnea-hypopnea index is not merely a clinical convention. It is a coverage criterion. Thresholds in that index determine who qualifies for a device. It is embedded in accreditation standards, in the scoring manuals, in the training of every technologist, in decades of trial inclusion criteria, and in the entire commercial structure of sleep testing. A learned embedding with no threshold, no units, and no interpretable definition cannot simply be substituted into a coverage policy.

This is the same wall that every good clinical AI result hits, and it is not primarily a scientific wall. A model can tell you the metric is wrong. It cannot tell you what to write in the local coverage determination.

I would expect the near term path to look like risk stratification layered on top of existing testing rather than replacing the index. Same study, same billing, an additional output that flags the group five patients for closer cardiovascular and cognitive follow up. Less satisfying than the headline. Considerably more likely to actually reach a patient.

So What

A model that learned sleep from scratch found a mortality gradient the field's central metric cannot see. The interesting claim is not that the model is clever. It is that the number we built the specialty on was never measuring the thing we thought it measured.

Expect the finding to survive and the metric to survive alongside it, because one of them is a scientific claim and the other is a payment rule.

Confidence: worked from the published abstract and results summary of an open access paper. The full methods, cluster characterization, and confounder adjustment have not been read in detail, and the strength of the AHI null result depends on how the comparison was specified. Treat the direction as solid and the magnitude as provisional pending a full read.

Sources

Primary: Bilal E, Araujo MLD, Beck KL, Heinzinger CM, Ghosn S, Saab CY, Foldvary-Schaefer N, Rogers JL, Mehra R. "A foundation model for sleep-based risk stratification and clinical outcomes." Nature Communications, 2026. nature.com

Validation cohort: Sleep Heart Health Study, National Sleep Research Resource. sleepdata.org