Sleep medicine has a number. It is called the apnea-hypopnea index, it counts how many times an hour you stop breathing, and essentially everything downstream of a sleep study runs on it. Diagnosis runs on it. Severity categories run on it. Insurance coverage for a CPAP machine runs on it. It is the number.
A team from IBM Research and the Cleveland Clinic trained a foundation model on more than ten thousand clinical sleep recordings and let it find its own structure. It found five patient groups with a clean mortality gradient across them. The apnea-hypopnea index showed no significant association with mortality at all.
That is the paper. Everything else is detail.
A polysomnogram is an extraordinarily rich object. Overnight EEG, ECG, airflow, oxygen saturation, effort belts, leg movement, all sampled continuously for hours. It is one of the densest physiologic recordings routinely collected in medicine.
We then take that recording and compress it, mostly, into one integer.
The team trained a foundation model to learn representations directly from the raw multichannel signal rather than from the scored summary. They then clustered patients in that learned embedding space, without reference to any outcome, and asked afterward what happened to the people in each cluster.
Five groups fell out. Ordered from lowest to highest, they show monotonic increases in cardiovascular events, neurologic outcomes, and death. Group five carried roughly a 2.71 fold mortality hazard against group one, along with elevated risk of heart failure, atrial fibrillation, and cognitive impairment.
Then the comparison that gives the paper its teeth. Ranked against those same outcomes, conventional apnea-hypopnea severity categories did not significantly predict mortality.
The obvious objection is that the model quietly rediscovered a better breathing metric and the whole thing is apnea severity with extra steps. The authors ran sensitivity analyses against exactly this and report that the model integrates EEG and ECG signal alongside respiratory channels rather than leaning on breathing alone.
If that holds, the finding is considerably more interesting than a better apnea score. It suggests that what a sleep study captures about your risk of dying is a whole body phenomenon, distributed across brain activity and cardiac rhythm and respiration together, and that we have spent forty years reading one channel of a multichannel signal and discarding the rest.
Which, to be fair, is what you do when you have to score studies by hand. The apnea-hypopnea index is not a stupid number. It is a number designed for a technician with a pen, and it has been kept long past the arrival of machines that do not need the simplification.
Suppose all of this is exactly right. Consider what would have to happen next.
The apnea-hypopnea index is not merely a clinical convention. It is a coverage criterion. Thresholds in that index determine who qualifies for a device. It is embedded in accreditation standards, in the scoring manuals, in the training of every technologist, in decades of trial inclusion criteria, and in the entire commercial structure of sleep testing. A learned embedding with no threshold, no units, and no interpretable definition cannot simply be substituted into a coverage policy.
This is the same wall that every good clinical AI result hits, and it is not primarily a scientific wall. A model can tell you the metric is wrong. It cannot tell you what to write in the local coverage determination.
I would expect the near term path to look like risk stratification layered on top of existing testing rather than replacing the index. Same study, same billing, an additional output that flags the group five patients for closer cardiovascular and cognitive follow up. Less satisfying than the headline. Considerably more likely to actually reach a patient.
A model that learned sleep from scratch found a mortality gradient the field's central metric cannot see. The interesting claim is not that the model is clever. It is that the number we built the specialty on was never measuring the thing we thought it measured.
Expect the finding to survive and the metric to survive alongside it, because one of them is a scientific claim and the other is a payment rule.
Sources
Primary: Bilal E, Araujo MLD, Beck KL, Heinzinger CM, Ghosn S, Saab CY, Foldvary-Schaefer N, Rogers JL, Mehra R. "A foundation model for sleep-based risk stratification and clinical outcomes." Nature Communications, 2026. nature.com
Validation cohort: Sleep Heart Health Study, National Sleep Research Resource. sleepdata.org