Heliox: Where Evidence Meets Empathy π¨π¦β¬
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Heliox: Where Evidence Meets Empathy π¨π¦β¬
The Jagged White Line
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
π Read: https://helioxpodcast.substack.com/p/the-jagged-white-line
How Bayesian Math and Cross-Attention AI Are Learning to Read Your Medical Future From Your Medical Past
Not that a machine can see your future. Itβs that, for the first time, the data weβve been generating about ourselves for thirty years might finally be organized well enough to whisper something useful back β quietly, individually, and early enough to matter.
A Bayesian framework for longitudinal EHR and genetic discovery, and nine other references
This is Heliox: Where Evidence Meets Empathy
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Disclosure: This podcast uses AI-generated synthetic voices for a material portion of the audio content, in line with Apple Podcasts guidelines.
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Spoken word, short and sweet, with rhythm and a catchy beat.
http://tinyurl.com/stonefolksongs
You know, usually when we talk about a medical diagnosis, there is this fundamental expectation of precision. It's almost like engineering. Right. Like if you fall and break your arm, you go to the hospital and the x-ray shows that jagged white line fracturing the radius. And the doctor just points at the illuminated screen and says, there it is. That is the problem. Yeah, it's a beautifully binary system. You know, you are either broken or you're not broken. You have a pathogen or you don't. And there is something deeply comforting about that, even when the news is bad. Absolutely. We desperately want our biology to be visible, to be categorized neatly into a single, you know, spreadsheet column. Because it implies a straightforward fix. Exactly. But then you step out of the emergency room and into the world of chronic illness. into the realities of aging and into the slow compounding ways our bodies change over decades. Which is where things get messy. Really messy. Suddenly that x-ray machine is useless. We are looking at a diagnostic landscape that isn't a clear photograph, but a remarkably murky, sprawling, chaotic ecosystem. And that is incredibly frustrating for both the patient and the physician. Because chronic disease doesn't just snap into existence the way a bone breaks. It creeps. It totally creeps. So I want you to imagine your health not as a series of random isolated snapshots like a surprise diabetes diagnosis at 50, an unexpected heart attack at 55, a sudden autoimmune flare up at 60. Right. But instead, imagine it as a continuous, highly predictable, continuously running movie What if the ailments you develop at age 60 were actually quietly, invisibly telegraphed by the seemingly minor disconnected doctor's visits you had at age 30? It's a huge what if. What if the data was always there, but we just didn't have the math to read the script? Well, that is the exact paradigm shift we are unpacking today. We're looking at the death of reactive medicine, waiting for that jagged white line to appear, and the birth of continuous, proactive algorithmic prediction. Okay, let's unpack this. The mission for today's Deep Dive is to explore a monumental structural shift in precision medicine. It really is monumental. We have a formidable stack of research papers today, centered primarily around a breakthrough called a Bayesian framework for longitudinal EHR and genetic discovery. We are going to explore how a multidisciplinary team of scientists is taking decades of incredibly messy longitudinal hospital data, your electronic health records, and mathematically fusing them with your underlying genetic code. To create dynamic time traveling predictions of your future health. Exactly. It really operates like a biological time machine. We are looking at computational frameworks that don't just, you know, guess your risk for next year, but actually project your entire biological trajectory decades into the future. And it dynamically updates every single time you interact with the health. care system. Precisely. So here is our roadmap. We are going to trace the historical evolution of this quest, wade into the absolute nightmare of harmonizing raw hospital data, and then introduce this groundbreaking Bayesian model which has the truly magical name, Alodi Nulai. Such a great name. We will look at exactly how it mathematically outmaneuvers our current gold standard clinical risk scores. And because we cannot talk about the frontier of predictive data without talking about Silicon Valley, we will examine how big tech is deploying massive, large language models to solve this exact same problem through a completely different architectural lens. Two very different paths to the same destination. Right. And finally, we'll figure out what this all means for you, the listener, and your future relationship with the health care system. So, to truly appreciate the magnitude of what models like Aledunio are doing, we have to recognize that this is the culmination of a quest that is literally thousands of years old. The desire to predict disease trajectory isn't a modern invention at all. Right, it's not just a Silicon Valley thing. No, not at all. If we look at the evolution and trajectory of personalized and precision medicine paper in our sources... It actually anchors this pursuit all the way back in antiquity. Around 1550 BC, you had Hippocrates and the humoral theory.
Okay, right. The idea that human health was dictated by the balance of four humors:blood, yellow bile, black bile, and phlegm. Exactly. And while their biological understanding was obviously, well, flawed by today's standards, their underlying epistemology was incredibly prescient. How so? Even in those ancient Greek and Egyptian treatises, there was this fundamental driving hypothesis that unique patient characteristics exist. dictate both disease susceptibility and the trajectory of recovery. Oh, I see. They were actively trying to individualize care based on the data points they could observe. The intuition was flawless, but their data resolution was, well, checking if you had too much phlegm. Exactly. They were operating at an incredibly low resolution. Then you jump forward to the mid-20th century, 1953, the elucidation of the DNA double helix. Suddenly we unlock the foundational molecular code governing heredity. Huge leap. Massive. Fast forward to 2003, the completion of the Human Genome Project, which provided the first reference genome to map human variation. And then the 1010s hit, and we enter the next generation sequencing, or NGS era. Where things get cheap. Right. The cost of sequencing a genome plummets from billions of dollars to a few hundred bucks, making it suddenly feasible to actually integrate dense genomic data into everyday clinical research. But here is the critical failure point of modern medicine and the exact problem these new papers are trying to solve. Even with all that massive genomic knowledge, the way we actually practice medicine today still relies heavily on what we can call the snapshot problem. The cross-sectional illusion. Right. Like if I go to my cardiologist tomorrow, they are going to calculate my risk using a static formula. They might use the pooled cohort equation, the PCE, to figure out my 10-year cardiovascular risk. Or if we were looking at oncology, they might use the JL model for breast cancer risk. risk. But these models are fundamentally static. They are completely anchored in the present tense. Yeah. The PCE takes a cross-section of you at one specific moment in time. It asks, "What is your age today? What is your systolic blood pressure today? What is your total cholesterol today? Do you smoke today?" And then it spits out a percentage. It's like trying to predict the final score of a baseball game and exactly who is going to hit a home run in the ninth inning by only looking at a single static photograph taken during the second. That is a perfect visualization of the mathematical flaw. You are completely ignoring the momentum of the game, how the pitcher's mechanics are degrading over time, the weather front rolling in. You miss the entire kinetic trajectory. Traditional models treat diseases and your biology in strict isolation. They treat your body as a series of disconnected static events rather than a continuously evolving, highly dependent ecosystem. Right. If you calculate your PCE at age 40 and it says you have a low risk, but then at age 42 you develop an autoimmune condition that causes systemic inflammation. Your whole profile changes. Completely. The PCE from age 40 is now clinically useless. It cannot update dynamically based on the introduction of a new competing variable. But let me push back here because this is the part that always drives me crazy when we talk about medical tech. We have had electronic health records, EHRs for decades. Hospitals have been typing my blood pressure, my prescriptions, and my weird symptoms into servers since the late 90s. Oh yeah. Tons of data. Right. So if we possess all this longitudinal data and we know that static models like the PCE are flawed, why hasn't a computer already connected these dots for us? Why are we only having this deep dive today? That is the structural bottleneck of the entire field. And the answer is found in the harsh reality of clinical data. It is an absolute unmitigated disaster. Disgusting mess. The paper from KU Leuven on modeling time-dependent patient trajectories, along with the PEHRT framework paper, which stands for Pipeline for Harmonizing EHR Data, they outline exactly why this is so mathematically punishing. Let's get into the mud here. Why is the data so bad first? You have the problem of irregular sampling You only go to the doctor when you are sick or for a random annual checkup that you might skip for three years
Right. I don't plug myself into a diagnostic bay every Tuesday at 9:00 a.m. To provide a perfectly spaced statistically clean data point exactly there are massive gaps of years where I am perfectly healthy and Then suddenly I have three hospital visits in a single month because I caught a weird respiratory virus And standard machine learning models despise irregular intervals. A standard recurrent neural network, for example, gets back a nice, evenly spaced time series. When it sees a three-year gap, it struggles to understand if nothing happened, or if the data is just missing. Which brings up the second problem, actual missing data. Yes, the missingness is rarely at random. Maybe a doctor was rushed and didn't record a secondary symptom. Or maybe you went to an out-of-network clinic while on vacation so that data never made it into your primary EHR. But the third problem seems to be the most intractable one, the coding systems. The heterogeneous ontologies. This is where data scientists literally lose their minds. Different hospitals and even different departments within the same hospital use entirely different coding systems to describe the exact same biological event. Yes. One system might use ICD-9 codes. Another uses ICD-10. Another uses SNUM-DCT for clinical notes. Yes. And another uses RxNorm for medications. So it's like trying to read a diary written by a doctor who is sprinting down a hospital hallway, constantly switching between five different languages mid-sentence, occasionally using outdated slang, and leaving random pages entirely blank. That is precisely what raw EHR data looks like to a computer. So before you can even dream of doing any time-traveling predictions, you have to build a universal semantic translator. That is what the P8 pipeline is designed to achieve. How does PHR actually untangle that? Because just mapping one code to another sounds like an endless manual nightmare of spreadsheets. It would be, which is why PHR doesn't rely on manual mapping. It leverages representation learning and pre-trained language models. Okay. Instead of looking at a medical code as just a discrete, isolated identifier, it treats the codes like words in a complex language. Like training an AI to understand the context of a sentence. Exactly. It uses these models to generate high-dimensional embeddings. It maps all this structured and unstructured data into a vector space. A vector space, meaning? In this mathematical case, codes that appear in similar clinical contexts are placed closer together. So an outdated ICD-9 code for a specific type of hypertension and a new SNOMED note about elevated blood pressure might look totally different in a spreadsheet. But in the vector space, their embeddings are nearly identical. Because the model has learned the semantic relationship, it knows they mean the same thing based on the surrounding context of the patient's journey. Precisely. It acts as an ontological harmonizer. It effectively takes the messy, multilingual, irregularly sampled diary and translates it into a single, cohesive, machine-readable, mathematically continuous timeline. So they are literally building the foundational track before they can even attempt to run the predictive train. You cannot do any of the advanced multimodal Bayesian modeling we're about to discuss without that standardized track. The data harmonization is the invisible heavy lifting of precision medicine. Which perfectly sets the stage for the cast of characters who decided to actually build the train. Our primary source today is this monumental paper,"A Bayesian Framework for Longitudinal EHR and Genetic Discovery." And the team behind this is a multidisciplinary powerhouse. They really are. You have Dr. Sarah Erbit, who is a physician scientist. She is a working cardiologist, meaning she is on the front lines looking at these static PCE scores and watching them fail real human beings. Right. But she also has incredibly deep computational expertise. She teamed up with Alexander Gusev, Pradeep Natarajan, and Giovanni Parmigiani, drawing resources across institutions like Mass General Brigham, Harvard, and the Broad Institute. This is a vital point about the sociology of science. This isn't just a group of computer scientists playing with health data in a vacuum, and it isn't just a group of doctors writing opinion pieces about algorithms. Yeah. This is a team that intimately understands both the brutal, messy reality of clinical pathophysiology and the rigorous mathematics required to model it without overfitting. And when this team looked at the landscape of predictive medicine, they saw a field that was repeatedly slamming into a series of deeply frustrating dead ends. Understanding these dead ends is crucial because it explains why their new approach is so radical. It sets the baseline. Let's walk through them. Dead end number one, which we briefly touched on, is treating diseases as isolated silos. Right. If you use a traditional Cox proportional hazards model, which is a very standard statistical tool in medical research, it might look at your risk of developing heart disease and then completely separately look at your risk of developing chronic kidney disease. But biology is not compartmentalized. Conditions co-evolve and compound. They cascade. Exactly. A metabolic issue, like insulin resistance, directly drives an inflammatory cascade, which damages the vascular endothelium, which leads to cardiovascular disease, which stresses the kidneys. It's a domino effect. If your statistical model does not natively capture the dependent interplay between these conditions, If it assumes they are independent variables, it is mathematically guaranteed to miss the plot. Which leads to dead end number two. Some researchers realized this, so they tried to group diseases together using something called unsupervised clustering, or topic modeling. Topic modeling is borrowed from natural language processing. Imagine you feed an algorithm a million news articles, and it automatically groups them into sports, politics, and finance based on word frequencies. Researchers tried doing this with EHR data. The algorithm looks at a massive pile of data. patient records and groups diseases together based on how frequently they co-occur in the same patients. That sounds like a smart way to find the cascades you just mentioned. Why is it a dead end? Because it is fundamentally retrospective. These models are great at explaining the past. They can look at a 70-year-old patient and say, ah, yes, you belong to the metabolic cardiovascular topic cluster. Right. But they assign you to that topic only after the events have already occurred. They struggle immensely to predict what will happen next when you're only 45 and only have one minor symptom. They don't model the dynamic tempo of how your risk changes as you age. They are descriptive, not generative. Exactly. And then there is dead end number three, which is frankly the most seductive trap in modern science right now. Black box, deep learning. Deep neural networks are unparalleled at finding hidden nonlinear patterns in massive high dimensional beta sets. But they are notoriously stubbornly opaque. Let's ground this in a clinical reality. Imagine Dr. Urbut is sitting across from a patient. She looks at a standard black box deep learning model on her screen and she says the computer says you have an 85 percent chance of a massive myocardial infarction in the next five years. We need to start you on aggressive lipid lowering therapies and beta blockers immediately. Right. The patient naturally is going to be a patient. going to panic and ask why. What is driving this risk? And if she's using a black box, her only honest answer is, I have no idea. The computer multiplied millions of weights and biases in hidden layers, and it output this probability. The math says so. Which is clinically and perhaps legally unacceptable. It is entirely unacceptable. A physician requires interpretability. They need to understand the underlying biological mechanism driving the risk in order to tailor the intervention. That makes sense. If the risk is driven by an inflammatory pathway, you treat it differently than if it's driven by a lipid pathway. Deep learning models hide the pathway. Furthermore, these black box models often fail catastrophically when dealing with censored data, which is a huge issue in longitudinal health tracking. Yes, censored data is a fundamental statistical hurdle. Let's say you are tracking a cohort of patients to see who develops diabetes. Patient A is currently 45 years old and does not have diabetes. But the study ends today. What happens to patient A when they are 50? Or 60? We don't know. Their future is right censored from our current view. They didn't stop existing, they just aged out of our observation window. Right. And many standard machine learning models don't know how to handle that. They often either drop the patient from the training set, which loses valuable data, Or they assume that because the patient didn't get diabetes during the observation window, they never get it. Which totally skews the probabilities. Exactly. A robust generative model must mathematically account for the fact that patients simply haven't lived the rest of their lives yet. So Dr. Erbet and her multi-institutional team looked at these dead ends... They realized they needed a model that could gracefully handle write-sensor data. They needed a model that was totally transparent and interpretable, outputting biological mechanisms, not just percentages. And crucially, they needed a model that could explicitly synthesize The static, unchanging nature of your DNA with the dynamic, constantly evolving nature of your aging body and your hospital visits. It is a phenomenally complex mathematical and biological ask. They needed to fuse a static baseline with a dynamic timeline. Yes. But they actually built it. And they gave it a name that I am absolutely obsessed with. It sounds like something out of an arcane grimoire. A Lady Nuly. It is quite the portmanteau. It's brilliant branding. It blends Aladdin, evoking the magic of a wish-granting genie, peering into the future, illuminating the unseen with Jacob Bernoulli, the 17th century Swiss pioneer of probability theory. Magic and math. What is truly fascinating here is how they structured the mathematics to avoid the dead ends we just discussed. A lot of newly is a Bayesian generative framework and the core architectural innovation and we need to be precise here because this is where the magic happens is that they formulated the model as a mixture of probabilities rather than a probability of a mixture. Okay, stop right there. I know our listeners are smart, but mixture of probabilities versus probability of a mixture sounds like an agonizingly subtle semantic difference. Why is this distinction the linchpin of the entire model? It changes everything about how the algorithm views a human being. Let's break it down. In traditional clustering models, the probability of a mixture approach the algorithm is fundamentally trying to force a patient into a single, mutually exclusive bucket. It assumes competitive allocation. Let me try an analogy. It's like looking at a complex recipe, and the algorithm is rigidly programmed to say this dish must be either a cake or are a stew. It cannot be both. It calculates the probability that you belong to cluster A versus the probability you belong to cluster B and then assigns you to the domino one. Exactly. Yeah. It assumes that if you are a cardiovascular patient, that label dominates and defines your trajectory. But human health doesn't work like that. My biology isn't exclusive. Allergy Lee flips the map. By using a mixture of probabilities, it treats your health like a massive dynamic buffet. Yes. It says you do not have to be just one thing. It calculates the weighted combination of different underlying signatures happening simultaneously. That is exactly it. It assumes that multiple latent biological processes are running in parallel inside you. You can have a moderate baseline of a metabolic signature, a heavy escalating helping of a cardiovascular signature, and a sudden dash of an inflammatory signature all operating at the exact same time at different weights. So my risk for a specific disease tomorrow isn't based on what bucket I'm in today. It's the weighted sum of all these different signature probabilities churning away underneath the surface. Yes. Your overall hazard function is a composite. This is how they built a model that actually reflects clinical reality, where a 65-year-old patient isn't just dealing with one primary issue, but is managing multiple simultaneous chronic conditions that interact and persist over a lifetime. To make this incredibly complex buffet work computationally, Allodinouli is built upon three distinct pillars. Pillar 1, population-level signatures. The model ingested all this longitudinal data and discovered 21 distinct replicable disease signatures. And we must emphasize that these 21 signatures are not arbitrary statistical groupings. They map to deep, recognizable biology. Like what? They found a distinct metabolic signature, a neuropsychiatric signature, a malignancy signature, an inflammatory signature. They capture how hundreds of specific diseases tend to travel together in packs over the course of a human lifetime. And these signatures are alive. They are time-dependent. Yes. The model learns the natural tempo of these processes across the population. It learns, for instance, that within the non-ischemic cardiovascular signature, the mathematical probability of developing atrial fibrillation rises steadily and specifically after age 55. or within a malignancy signature, the hazard for metastatic disease spikes sharply between the ages of 60 and 75. Wow. The model understands the chronobiology of disease. So that's the population level. Pillar 2 is where the model zooms in on you specifically. This is the concept of individual signature load. This represents how heavily you, the individual patient, map to each of those 21 signatures. And crucially, this loading is not a static number written in stone. As you age and as you unfortunately accumulate new symptoms or diagnoses in your EHR, your loading dynamically shifts. Right. If you get diagnosed with type 2 diabetes at age 50, your loading on the metabolic signature shoots up, which automatically cascades to increase your future probability of renal failure or neuropathy. It is a continuously updating Bayesian posterior. Every new piece of data refines the model's belief about your true trajectory. Which brings us to Pillar 3, the absolute game changer and the hardest part to engineer, the genetic integration. They incorporated 36 different polygenic risk scores, PRSs, alongside biological sex and 10 genetic principal components to rigorously account for ancestry and population stratification. This is where we fuse the static with the dynamic. Let's clarify what a polygenic risk score is first. a moment. We aren't talking about single gene mutations here, like the BRCA mutations for breast cancer. No, we're talking about complex traits. Most chronic diseases, heart disease, diabetes, schizophrenia, are not caused by a single broken gene. They are polygenic. Yeah. A polygenic risk score looks at millions of tiny single letter variations of the gene. across your entire genome. Single nucleotide polymorphisms or SNPs. Each variant might only increase your risk by a fraction of a percent, but when you aggregate millions of them, you get a powerful singular baseline score representing your innate genetic susceptibility for a trait. So, Aladdinoli takes these static polygenic risk scores, your genetic blueprint, which is identical on the day you are born and the day you die, and uses them as the anchor. They act as the gravitational pull, driving how heavily you load into those dynamic, time-varying disease signatures over the course of your life. It is an incredibly elegant architecture. The genetics drive the signatures, and the signatures drive the specific diseases. And by structuring the architecture this way, the researchers managed to solve one of the most maddening headaches in medical statistics, which you alluded to. to earlier, the competing risk problem. If we connect this to the bigger clinical picture, the competing risk problem is the bane of longitudinal studies. In a standard model, if a patient is diagnosed with one major systemic disease, say severe heart failure, The model often breaks down when trying to figure out their risk for everything else. Because the heart failure fundamentally alters their survival probability. Yes, and their interaction with the health care system entirely. But in Aladinuli, because the model is inherently generative and because it is calculating the weighted probabilities for 348 different diseases simultaneously through those 21 signatures, you don't just fall out of the model when you get sick. Exactly. You remain at risk for the remaining 347 conditions, your overall hazard profile just elegantly mathematically updates as you absorb that new diagnosis. It recalculates the entire landscape of your future. It is like a GPS for your biology. If you take a wrong turn or in this case develop a new condition, it doesn't just crash. It says recalculating and maps out the new most probable path of your biological decline. That is exactly what it does. So the Bayesian math is elegant. The theoretical architecture is sound. But Dr. Uba and her team didn't just build this in a theoretical sandbox and publish a paper about the math. They put El Adanuli to the ultimate grueling test. They deployed it across three massive, completely independent biobanks to see if it actually worked on real human lives. The scale of the validation phase here is staggering. And it's what makes this paper so authoritative. They didn't just use one local hospital data set. They utilized the UK Biobank, the Mass General Brigham Clinical Cohort, and the All of Us Research Program. We are talking about over 683,000 distinct patients. With up to 52 years of continuous longitudinal medical follow-up data, They were tracking the intertwined probabilities of 348 different diseases over half a century. This is where we get into the head-to-head battle. They pitted their magical Bayesian engine against the reigning, unquestioned champions of clinical risk scores. And they didn't cheat. This is a vital point about methodology. They evaluated the model dynamically, simulating strict clinical reality. In data science, it is very easy to accidentally allow information leakage, where the model sneaks a peek at the future data to make its prediction look better. But they block that completely. Entirely. They set up strict temporal boundaries. They asked Aladinoli, based only on the genetic data and the EHR data available for this specific patient up to year 5, predict their disease trajectory in year 6. The model generates its probabilities. Then and only then, they roll the clock forward to year 6, check reality, and grade the model. Then they ask it to predict year 7 based on data up to year 6. It's a rolling dynamic evaluation. And in that strict, unforgiving, predictive environment, Elad-Nilouli absolutely dominated. It outperformed the pooled cohort equation and the newer PREVENT model for 10-year cardiovascular predictions. It also outperformed the GL model for one-year breast cancer horizons. Across the board, they reported it achieved a median dynamic AUC area under the curve of 0.85. For listeners who might not breathe statistics daily, the AUC is a metric that measures a classifier's ability to distinguish between classes. An AUC of 0.5 means the model is basically just flipping a coin. An AUC of 1.0 is absolute perfection. Right. In the incredibly noisy, chaotic world of predicting complex, multisystemic human health over decades... Achieving a median dynamic AUC of 0.85 across 348 diseases is a monumental achievement. It means the model is highly robust. But we have to address a massive elephant in the room regarding the data. No real-world data set is perfect. And the team had to pull off a deeply impressive, rigorous mathematical triumph to make sure their stunning results weren't just a mirage caused by biased data. This raises a critical issue about the fidelity of biobank data, particularly the UK biobank. It is a phenomenal resource, one of the best in the world, but it suffers from a well-documented epidemiological phenomena known as participation bias. The healthy volunteer effect. Exactly. The people who have the time, resources, and inclination to volunteer for long-term intensive health tracking studies tend to be, on average, significantly healthier, wealthier, and more health conscious than the general population. So if you train your AI exclusively on a population of super healthy, health conscious volunteers, the model is going to learn artificially low baseline hazards. It's going to be fundamentally miscalibrated when you try to deploy it in a real world general public clinic. Precisely. the baseline probabilities would be skewed. This is a violation of the missing-at-random assumption in statistics. If they just ran the data raw, Aladinoli would learn a distorted version of human biology. Which defeats the purpose. Right. But because Aladinoli is a generative model with an explicit mathematically defined likelihood formulation, They were able to implement a rigorous statistical correction called inverse probability weighting, or IPW. Let's break down exactly how IPW rescues the math here. IPW modifies the actual likelihood function during the model's training phase. It assigns an observation weight to every single patient in the training set. If a patient belongs to a demographic that we know is overrepresented in the biobank, say, wealthy, healthy individuals, IPW mathematically downweights their contribution to the final probability curve. Conversely, if a patient belongs to a demographic that is underrepresented, their data is upweighted. So by mathematically adjusting the weights and the likelihood function... they force the model to learn a probability distribution that accurately reflects the true underlying biological signals of the broader population, rather than the skewed demographics of the biobank. Exactly. It prevents the model from suffering asymptotic bias. It is a rigorous, necessary step to ensure that when Al Adinley says you have a 40 percent risk, that number is actually calibrated to reality, not to a fantasy world of perfectly healthy volunteers. Okay, so they built a mathematically rigorous, unbiased model that predicts the future better than the clinical gold standard. standards. That alone is worthy of a deep dive. But here is the massive "aha" moment for me in this paper. Elginouli isn't just a crystal ball for doctors, it is a microscope for biologists. This is where the dual nature of the Bayesian framework truly shines. It predicts the future, yes, but because it explicitly models those 21 latent signatures, it also maps the underlying biology. It becomes an engine for discovery. Because it maps these underlying biological signatures, the team realized they could run a new kind of genetic study. Instead of just looking at how genetics map to a single disease, which is how standard genome-wide association studies, or GLEDs, work, they looked at how genetics mapped to the signatures. They ran a signature-based G-violence, and the results were incredibly revealing. They discovered 151 genome-wide significant LOSAR. These are 151 specific locations on the human genome that are deeply associated with disease risk. But here's the crazy part. Traditional single trait analyses completely missed some of these leci. Scientists had been staring right at them, but because they were only looking for a direct line between a gene and one specific disease, like DNA versus heart attack, they couldn't see the signal. But because Aladinoli zooms out and looks at the entire signature of the neighborhood of diseases, it spotted the genetic culprits hiding in plain sight. It proves a concept in genetics called pleiotropy. Pleiotropy meaning that a single genetic variant exerts effects on multiple different traits or diseases simultaneously. Exactly. If a specific genetic variant subtly increases your inflammation, it might slightly increase your risk for arthritis, and slightly increase your risk for heart disease, and slightly increase your risk for psoriasis. Oh, I see. If you look at any one of those diseases in isolation, the genetic signal is too weak to cross the threshold of statistical significance. It looks up noise. But aladynulaya aggregates those subtle signals across the entire inflammatory signature. And suddenly the signal becomes massive and undeniable. Alladynly doesn't just catch the final disease. It catches the mechanism. And that leads to this incredible paradigm shattering concept of biological subtypes or heterogeneity within disease categories. What is truly fascinating here is that the model proves mathematically that our current diagnostic labels are often woefully inadequate. How so? Let's ground this in a hypothetical scenario that happens every day. You have two different patients, patient A and patient B. Both arrive at the emergency room and both suffer a myocardial infarction, a heart attack. On their medical chart, the diagnosis is identical. The ICD-10 code is identical. But Alati Newley reveals that their biological journey to that exact same hospital bed could be entirely fundamentally different. different. Entirely different. Because Aladinili maps their signature loadings over time, we can see the historical engine of their disease. The model shows that patient A might have arrived at that heart attack through a heavily metabolic pathway. Right. For 20 years, their risk was driven by escalating insulin resistance, obesity, and diabetic complications that slowly destroyed their vascular system. While patient B might have arrived at the exact same heart attack through a pathway, maybe they have zero metabolic issues, perfect blood sugar, but they have a massive chronic inflammatory loading driven by a specific polygenic risk score. And the model can separate these patients with astonishing clarity. The paper noted that the statistical effect sizes when comparing these patient clusters were massive. They cited a Cohen's D of up to 4.25. To put that in perspective for the listener in statistics, a Cohen's D measures the standardized difference between two means. A score of a 0.8 is generally considered a large effect size. It's very large. A 4.25 means the biological signatures of these patients, despite having the exact same clinical diagnosis, are separated by vast, undeniable canyons of biological difference. The researchers noted that over 95% of their comparisons between these subtypes yielded highly significant P-virus. values is a profound realization yeah you have two people with the same disease yeah but they likely require completely different preventative strategies and treatments because the biological engine driving their disease is different wow if you give patient B a drug designed to fix patient A's metabolic pathway it might do absolutely nothing Aleta Neely maps the engine it shifts the entire paradigm from treating a generalized label to treating a specific patient level biological process Yes, it really does. Now, to show that this isn't just one brilliant team in Boston having a eureka moment with Bayesian math, we have to talk about the tech giants. Because while Dr. Erbet's team was perfecting Eleni Newley, Silicon Valley was looking at the exact same massive problem merging messy EHRs with static genetics and deciding to throw overwhelming brute force computing power at it. Right. If we shift our focus to the paper titled Integrating Genomics into Multimodal EHR Foundation Model, which was authored by a massive collaboration between Verily, Google, and NVIDIA. They tackled the exact same biological problem, but they took a fundamentally different architectural approach. They didn't use Bayesian mixtures. They used large language models, LLMs. They took the exact same underlying transformer architecture that powers tools like ChatGPT, but instead of predicting the next word in a sentence, they are trying to predict the next disease in your medical chart. They treat the medical record literally as a language. A diagnosis is a word. A hospital visit is a sentence. Your entire medical history is a document. It's a crazy way to think about it. It is, but it works. They pre-train a massive foundation model on millions of these clinical trajectories, essentially teaching the neural network the complex grammar and syntax of human health. But the challenge they faced was immense. If you are building an LLM that reads a timeline, a sequence of events, How do you inject static, unchanging DNA into it? If I hand an LLM a book to read chapter by chapter, how do I simultaneously force it to inject a permanent overarching theme or vibe into its understanding of every single page? It is a profound engineering challenge. You have a dynamic time series modality on one hand and a static high-dimensional genomic modality on the other. How do you fuse them? The Verily team solved this using sophisticated neural network techniques, specifically adapter modules and cross-attention mechanisms. Let's focus on the cross-attention because the mechanics of how this works are wild. How does the model actually see the DNA while reading the hospital record? Cross-attention is a mechanism borrowed from encoder-decoder paradigms in machine translation. First, they take the polygenic risk scores, the static genetic data, and they project them through a dedicated neural network layer. This projection creates what they call soft tokens. Soft tokens. What does that mean in the context of an LLM? Think of the model's vocabulary. It has hard tokens for actual medical concepts, like an ICD code for asthma. A soft token is essentially a synthetic continuous vector in the model's embedding space. Okay, I follow. It doesn't correspond to a specific word in the medical dictionary. Instead, it is a dense mathematical essence of your genetic risk. So they turn your DNA into a mathematical vibe that the language model can understand. Exactly. Now, as the transformer model reads through your clinical history chronologically, step by step, the cross-attention mechanism kicks in. It uses a query key value system. Like a database search. Basically. The patient's current clinical state, where they are in their medical history right now, acts as the query. The genetic soft tokens act as the keys and values. So at every single step of your timeline, as the model reads the story of your life, the cross-attention mechanism mathematically forces the model to constantly attend to or reference those genetic soft tokens to see how the story should unfold next. It's a continuous dynamic fusion. The model asks, "Given that the patient just developed hypertension in the query, and given their specific underlying genetic risk profile for cardiovascular disease, the key value what is the most mathematically probable next event?" And did this wildly complex cross-attention mechanism actually work? Phenomenally well. They evaluated this multimodal architecture on the All of Us dataset when they explicitly injected that genetic data into the foundation model via cross-attention, They saw a massive 15.5% relative increase in the AUPRC, the area, under the precision recall curve, specifically for predicting the onset of type 2 diabetes within a 10-year window. A 15.5% jump in predictive precision just by giving the LLM the genetic content. It is a massive leap. It proves definitively that relying on EHR data alone, even with a massive foundation model, leaves critical predictive power on the table. Combining longitudinal clinical data with static genomic data is unequivocally the holy grail of predictive modeling. Here's where it gets really interesting for me regarding the Verily paper. It wasn't just about combining the data. They also invented a breakthrough in how they actually calculate the final risk percentage. They call it path computing probabilities. Yes, this is a crucial evaluation from how early, naive EHR foundation models function. Previously, if you wanted to predict a future risk, you would use something called Monte Carlo sampling. Which means rolling the dice. over and over again. Right. You have the patient's history up to today. You ask the LLM to generate, say, 100 possible synthetic futures for that patient based on the learned probabilities. Right. If the patient develops diabetes in 20 out of those 100 simulated futures, the model assigns a 20% risk score. But that creates really clunky, noisy, bucketed risk scores. Because it's discrete sampling, you get 20% or 21%, but it's not a smooth mathematical truth. It's an approximation based on how many times you rolled the dice. Exactly. It is computationally expensive and difficult to use for fine-grained clinical risk stratification. Path computing probabilities completely bypass the Monte Carlo dice rules. Instead of generating distinct futures, the algorithm uses exact Bayesian update rules at every single token generation step along the path. So it mathematically calculates the exact probability of every specific incremental step happening, multiplying the sequential probabilities together. Yes. It calculates the cumulative exact probability of a target token, like a diabetes diagnosis, occurring at any point in the future sequence. This generates a highly nuanced, completely continuous, exact mathematical curve of unique risk for every single individual. It extracts so much more meaningful, high-resolution signal from the data. It allows a hospital system to look at a patient and set incredibly precise, individualized thresholds for intervention, rather than relying on clunky 10% buckets. It represents a massive leap in the mathematical maturity of foundation models in health care. So we have to pull all these threads together. What does this all actually mean? We have Dr. Erbit's team building a Latinoi, identifying hidden biological signatures and disease engines with elegant Bayesian math. Yes. And we have Google and Verily using massive large language models with cross attention to predict the next clinical word in your chart with astonishing accuracy. What does this convergence mean for you, the listener, sitting in a doctor's office five or ten years from now? If we connect all of this to the bigger picture, it signals the definitive end of reactive acute intervention as the primary mode of health care. We are moving irreversibly toward continuous proactive health management. The synthesis of these technologies effectively allows the health care system to create a predictive digital twin of you. A digital twin. That sounds like sci-fi, but it's mathematically real. It is. A virtual, high-dimensional representation of your biology that is constantly, dynamically updated with every single lab test, every new prescription, every minor doctor's visit, all permanently grounded by your unique genetic code. Your physician won't just be looking at the physical use sitting on the exam table. They will be consulting your digital twin to see exactly where your current biological trajectory is mathematically leading. And think about how this completely overhauls the pharmaceutical industry and how we run clinical trials. Right now, drug development is incredibly inefficient. If I have a promising new experimental drug for breast cancer, I design a trial and look for a bunch of patients who have the broad label breast cancer on their chart. But as we discussed deeply with the myocardial infarction subtypes, that broad clinical label masks massive underlying biological heterogeneity. Right. If you test a drug on a heterogeneous group, the signal of efficacy often gets drowned out by the noise of patients who have a different underlying disease engine. The trial fails. But with L. DeNoli's continuous signature loadings, we don't have to rely on broad labels anymore. We can do signature-based patient matching. I can match you with 500 other patients who share your exact, specific, underlying biological engine. Say, a highly specific intersection of the metabolic and inflammatory signatures driven by a specific polygenic risk score. I can build a clinical trial based on the pathway, not the final endpoint. It could dramatically increase the success rate of clinical trials because the pharmaceutical interventions would be hyper-targeted at the actual verified mechanism of disease for that specific mathematically defined subgroup. We would waste less time, spend less money, and bring effective targeted cures to patients significantly faster. But I have to ask one final critical question regarding the data foundation of all this. Both the Bayesian models and the LOMs rely on massive amounts of data to learn these probabilities. Extensive medical histories, detailed genetic sequencing, decades of follow-up. How do we ensure that these models are actually accurate for everyone? That's a vital question. If a model is trained primarily on data from populations who have excellent healthcare access and frequent doctor visits, aren't we just building a mathematical reality that only works for them? That is the most critical hurdle for the clinical deployment of these models. It is a matter of pure data fidelity and calibration. If a model, particularly one relying on polygenic risk scores, is trained exclusively on genomic data from individuals of European descent, it will suffer from severe population stratification errors. The model will simply be mathematically blind to the genetic variations that drive disease in other populations. Precisely. The baseline probabilities will be fundamentally miscalibrated for individuals of non-European ancestries, rendering the predictions inaccurate and clinically useless for those patients. This is not just a theoretical concern. It is a known flaw in early genomic research. So how do the researchers solve this calibration problem? You solve it by fixing the training data itself. This is why both of our primary sources explicitly highlight their reliance on datasets like the All of Us Research Program and the Errigen Network. The All of Us cohort was central to validating the Verily Foundation model, right? Yes, and the Al Genoey team utilized it as well. The All of Us program is a massive nationwide initiative specifically engineered to capture comprehensive data from historically underrepresented populations. They intentionally bypass the standard homogenous biobank demographics. They recruit broadly to build a truly diverse multi-ancestry cohort. And the Emerge consortium serves a similar structural purpose. Exactly. Emerge focuses heavily on the real-world clinical implementation of genomic data across diverse patient populations. By actively choosing to train and calibrate L.D. Nolley and these massive foundation models on diverse, highly representative data sets, the researchers can mathematically ensure that the predictive probabilities are robust, accurate, and scientifically valid across a wider spectrum of human biology. Okay, let's take a deep breath and recap the incredible intellectual journey we have been on today. today. We started in a frustrating world of messy, isolated snapshots. Doctors looking at a single blood test and trying to guess the future using static equations like the PCE that mathematically failed the moment a patient's life changes. And we explore the immense, unglamorous difficulty of cleaning that clinical data, relying on sophisticated vector spaces and pipelines like PERT to harmonize the linguistic chaos into a structure that machines can actually comprehend. Then we met the team behind Alade Nulli, blending magic and math, using a rigorous Bayesian framework to treat your health like a dynamic buffet of probabilities, uncovering hidden biological signatures, proving the existence of distinct disease subtypes, and mathematically outperforming the clinical gold standard. And we saw how the tech giants are pushing the envelope from a different angle, using large language models and intricate cross-attention mechanisms to seamlessly inject static DNA into the dynamic, continuous timeline of your life, proving definitively that multimodal integration is the future of precision medicine. All of this research is driving toward a reality where you possess a predictive digital twin, and medicine evolves to treat the underlying engine of your disease long before the ultimate clinical label is ever applied. It is a profound, irreversible structural transition in our fundamental understanding of human health. Which leaves me with a final lingering thought for you to ponder as we wrap up this deep dive. If we are entering an era where a machine can seamlessly integrate your static DNA and the semantic context of every single doctor's visit you've ever had, in order to predict the exact continuous trajectory of your biological decline with unprecedented mathematical accuracy, Does knowing that trajectory rob you of the illusion of control? Or does seeing the jagged white line years before the bone even breaks give you the ultimate cheat code to actively, aggressively rewrite your own fate? That is the ultimate question at the frontier of predictive medicine. Thank you so much for joining us on this deep dive. The Future of Medicine is being written today, and we loved unpacking the incredible mechanisms behind it with you.
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Hidden Brain
Hidden Brain, Shankar VedantamAll In The Mind
ABC Australia
What Now? with Trevor Noah
Trevor Noah
No Stupid Questions
Freakonomics Radio + Stitcher
Entrepreneurial Thought Leaders (ETL)
Stanford eCorner
This Is That
CBCFuture Tense
ABC Australia
The Naked Scientists Podcast
The Naked Scientists
Naked Neuroscience, from the Naked Scientists
James Tytko
The TED AI Show
TED
Ologies with Alie Ward
Alie Ward
The Daily
The New York Times
Savage Lovecast
Dan Savage
Huberman Lab
Scicomm Media
Freakonomics Radio
Freakonomics Radio + Stitcher
Ideas
CBCLadies, We Need To Talk
ABC Australia