Home / Insights / Technology and power
AI Is Learning a New Language: The Human Body
Sleep may be the first proof that foundation models are moving beyond text, images and code into human physiology.
A model at Stanford can now predict dementia risk from a single night of sleep. Concordance index 0.85, where 0.5 is a coin flip.
Not a genetic test. Not a scan. Not a decade of follow-up. Eight hours of physiology, read by a machine.
Most of the coverage filed this under sleep.
That is the wrong file. This is not a story about sleep. It is the first clear signal that foundation models are leaving the digital world entirely.
The languages AI has already learned
For the past five years, AI has learned the languages we produce.
Text. Code. Images. Speech. Video.
Every foundation model you have heard of was trained on the exhaust of human thought. The words we typed. The pictures we posted. The code we shipped. Large language models did not learn to think. They learned the statistical structure of what we chose to externalize.
That was phase one. It is close to saturated. The public internet is a finite corpus, and the industry has already scraped most of what is worth scraping. The current scramble for proprietary data, licensing deals and synthetic generation is a symptom of that ceiling, not a strategy for breaking it.
So the interesting question is not how to squeeze more out of text. It is where the next genuinely untapped corpus lives.
The next frontier is not another language we produce with our minds. It is the language our bodies produce without us deciding to.
Brain signals. Heart activity. Respiration. Muscle activity. Sleep cycles.
We have been generating this data our entire lives. We simply lacked models capable of reading it.
What Stanford actually built
On 6 January 2026, Nature Medicine published “A multimodal sleep foundation model for disease prediction.” The work was led by Rahul Thapa, with senior authors Emmanuel Mignot and James Zou at Stanford.
The model is called SleepFM. Here is what it demonstrably is, before I add any interpretation.
It is a foundation model trained on more than 585,000 hours of polysomnography from roughly 65,000 participants. Polysomnography is the full wired-up sleep study run in a hospital lab. It captures brain activity, heart activity, muscle tone, eye movement and respiration at the same time, all night. It is the clinical gold standard, and it is expensive, which is precisely why nobody had assembled a corpus at this scale before.
SleepFM integrates four signal modalities: brain signals including EEG and EOG, electrocardiography, electromyography, and respiratory signals. The researchers built a new contrastive learning approach so the model could handle recordings from different labs with different sensor configurations. That is the unglamorous engineering problem that usually kills this kind of work.
The detail that matters most is how they tokenized the input. Each recording was broken into short segments, roughly five seconds each. The researchers compared this directly to how words are used to train a language model.
They took a night of human physiology and turned it into a vocabulary.
James Zou described the outcome in one sentence. SleepFM, he said, is essentially learning the language of sleep.
That is not a metaphor I invented for a headline. That is the researcher’s own framing of his own model.
The results, stated precisely
I want to be careful here, because factual authority is the entire point of writing this.
SleepFM does two things.
First, it performs standard sleep tasks such as sleep staging and scoring apnea severity as well as or better than current state of the art models. Useful. Also incremental, and not why this paper matters.
Second, it predicts future disease risk from a single night of sleep. Across more than 1,000 candidate disease groupings, it identified 130 conditions predictable with a concordance index of at least 0.75, Bonferroni corrected. That list includes:
- All-cause mortality: 0.84
- Dementia: 0.85
- Myocardial infarction: 0.81
- Heart failure: 0.80
- Chronic kidney disease: 0.79
- Stroke: 0.78
- Atrial fibrillation: 0.78
For anyone unfamiliar with the metric: a concordance index of 0.5 is a coin flip. 0.85 for dementia, from eight hours of sleep signals, is not a coin flip.
One night. Not a genetic test. Not a scan. Not a decade of longitudinal follow-up.
Now the caveats, because they are part of the authority, not a threat to it. This is a retrospective study on clinical cohorts, which means the population skews toward people who already had a reason to enter a sleep lab. Prediction is association, not causation, and the model does not tell you why a given night forecasts dementia. Nothing here has been validated prospectively at population scale. A concordance index is a ranking statistic, not a diagnosis for an individual.
That is the demonstrated result, honestly bounded. Everything past this line is my thesis, and I will label it as such.
The thesis: foundation models are escaping the digital world
Here is what I think this actually means.
LLMs learned from the exhaust of human thought. The next generation of AI may learn from the exhaust of human biology.
Consider what a body produces. Every heartbeat is data. Every breath is data. Every brain wave is data. Every night of sleep is an eight-hour multimodal dataset generated whether you consent to think about it or not.
You cannot run out of it. Roughly a third of your life is spent producing a clean, dense, continuous training signal that until now was thrown away.
SleepFM is one instance of a broader pattern. Foundation models are already being built across biological domains: protein language models that learned the grammar of amino acid sequences, genomic models trained on DNA, single-cell models trained on gene expression. What SleepFM adds is the whole-organism, real-time, sensor-level layer. Not the blueprint of a body, but the live telemetry of one.
I am not claiming to have coined a scientific category. I am offering a lens: the biological foundation model. A model whose training corpus is not the internet, but the signals of a living organism.
If phase one was AI reading what we wrote, phase two is AI reading what we are.
Where this leaves the sleep-tracking industry
Most people already track their sleep. A ring, a watch, a mattress sensor. A number in the morning that decides whether you are allowed to feel tired.
Be honest about what that is. Consumer sleep scores are estimates derived largely from movement and heart rate, not clinical measurement, and they are not standardized across devices. Two devices on the same body will disagree. Chasing a perfect number can backfire into anxiety about sleep itself, a pattern clinicians have named orthosomnia.
That is measurement without understanding. It is the fitness-tracker paradigm applied to a system nobody involved actually models.
SleepFM points at something structurally different. It does not count minutes in REM. It learns the relationships between physiological signals, and what those relationships predict over years. That is a shift from observation to interpretation. It is the same shift we watched when language models stopped counting keywords and started modeling meaning.
The question nobody is asking
If your body is a training corpus, someone owns it.
That is not a rhetorical flourish. It is the operational question that follows directly from the thesis.
The clinical data behind SleepFM sits inside research cohorts with consent frameworks, ethics review and publication scrutiny. The data on your wrist tonight does not. It sits with a consumer hardware company, governed by a terms-of-service document you did not read, in a category where the regulatory line between wellness gadget and medical device is deliberately blurry.
The value of a biological foundation model scales with continuous, longitudinal, individual-level physiology. The entity that assembles that corpus first will hold something more predictive about you than your own doctor does, and will hold it without a duty of care.
I ship AI systems for a living. The pattern is always the same. The capability arrives before the governance, and whoever moves early sets the defaults for everyone else.
This one is worth setting deliberately.
The direction of travel
I have to be exact here, because this part is forward-looking and not proven by SleepFM itself.
Today, sleep advice is population-level. Avoid caffeine after 2 p.m. Keep the room cool. True on average, useless in particular.
The direction this research points toward is personal. Not “humans should avoid coffee in the afternoon,” but eventually something closer to: based on your own longitudinal physiology, caffeine after a specific hour appears associated with measurable deterioration in your sleep.
SleepFM does not do that yet. It runs on clinical lab data, not your wrist. The signal quality gap between a polysomnography rig and a consumer sensor is real and large. But the researchers themselves note that as wearable technology advances, models of this kind may open the door to noninvasive, continuous health monitoring.
The bridge from lab to living room is an engineering problem now, not a conceptual one. Engineering problems get solved. That is the part worth planning around.
The real conclusion
The future of AI may not be a machine that knows more about the world.
It may be a machine that identifies patterns in your own biology that you cannot see yourself.
Sleep is simply one of the first places where we can watch that future take shape. We spent a decade teaching machines to read our words. We are about to spend the next one teaching them to read us.
Most people who read about SleepFM will file it under sleep. Under wellness. Under one more thing to optimize before bed.
That is the wrong file. The machine reading tonight may know more about your next ten years than you do.
Sources
Thapa, R., Mignot, E., Zou, J., et al. “A multimodal sleep foundation model for disease prediction.” Nature Medicine, 6 January 2026. https://www.nature.com/articles/s41591-025-04133-4
Stanford Report. “AI model predicts disease risk while you sleep.” January 2026. https://news.stanford.edu/stories/2026/01/ai-model-sleep-disease-risk-research-sleepfm
First published in the OG Approved newsletter on 22/07/2026. Read it on Substack or subscribe to get the next one.


