Heliox: Where Evidence Meets Empathy π¨π¦β¬
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Heliox: Where Evidence Meets Empathy π¨π¦β¬
π¬We Taught a Machine to Stop Guessing and Start Wondering
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
The Five-Paper Journey to Building the First Autonomous AI Scientist
There is a particular kind of vertigo that comes from realizing something you assumed was thinking was, in fact, only remembering very well. For a few years now, this has been the quiet, uncomfortable secret sitting underneath the astonishing performance of large language models. They write. They code. They pass exams that would humble most graduate students. And yet, ask one to explain what happens when a familiar rule of the universe simply stops applying, and something in the machinery seizes up. It doesn't discover. It confabulates. It reaches for the nearest plausible-sounding thing in its training data and asserts it with the same confidence it uses for everything else.
This episode of Heliox traces the five-paper path that got us from that uncomfortable diagnosis to something genuinely new: an AI system that can behave, in a narrow but real sense, like a scientist. Not a stenographer of scientific consensus. An actual inquirer.
References:
and five other papers
This is Heliox: Where Evidence Meets Empathy
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Disclosure: This podcast uses AI-generated synthetic voices for a material portion of the audio content, in line with Apple Podcasts guidelines.
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Spoken word, short and sweet, with rhythm and a catchy beat.
http://tinyurl.com/stonefolksongs
You know, it is wild when you stop and think about where we are with artificial intelligence right now. Oh, absolutely. Like, just take a second and look at the device sitting in your pocket or the computer on your desk. You have access to a system that can do things that would have seemed like pure, unadulterated science fiction barely a decade ago. Yeah, the progress is just... These systems can even pass the bar exam, diagnose medical conditions from x-rays, write functional software code. They're essentially superhuman at recognizing patterns. Right. But this is a massive fundamental... But here, if you take that exact same brilliant, supposedly genius AI and you drop it into a completely novel situation. A scenario it has literally never encountered in its massive training data. Yeah, exactly. And you ask it a really simple "what if" question, it often just... to stop acting like a giant passive pattern matcher and start acting like an actual scientist. It's a huge paradigm shift. It really is. We are going to explore the creation of an AI that can generate its own hypotheses, design its own original experiments, and discover the hidden laws of nature all by itself. It is just a phenomenal journey. And to really grasp how monumental this shift is, we are going to follow the intellectual breadcrumbs that led up to a truly groundbreaking piece research the August 2026 paper yes published by Kevin Murphy titled model discovery agent but you know we can't just parachute right into Murphy's final system right because the model discovery agent or MDA as we'll probably call it today it didn't just appear out of nowhere not at all it is the brilliant synthesis of several massive sequential roadblocks that the scientific community had to smash through over the last Well, the first roadblock was fundamentally philosophical, but it had really deep mathematical consequences. Oh, interesting. Yeah. It was all about defining what it actually means to understand something. So back around 2024, researchers at DeepMind, specifically Jonathan Richens and Tom Everett, Well, if you want an AI to survive and adapt in a changing universe anyway. Wait, why is it a dead end? If the line passes through all the docks, doesn't that mean the AI kind of figured it out? Not at all. Because if the underlying rules of the universe shift even slightly what scientists call a shifting distribution, that beautiful line the AI drew becomes completely useless. Oh, I see. Richens and Everett actually proved mathematically in theorems 1, 2, and 3 of their paper that that if an agent is going to survive unknown environmental shifts, it cannot just be a curve fitter. It means something more foundational. Right. It must learn the actual causal mechanisms generating the data. And to explain this, they leaned on a framework created by Judea Pearl. Oh, the causal inference guy. Exactly. A giant in the field. Pearl created something called the causal hierarchy, which separates intelligence into three distinct levels. I do love a good hierarchy. What's level one? Level one is association or simply seeing. This is where current AI lives and breathes. It observes. Okay. Let's use a classic example. An AI observes that every single morning a rooster crows and shortly after the sun comes up. Right. So it associates the two events perfectly. Its curve fits the data. Exactly. So the AI predicts that rooster crows cause sunrises. Well, or at least it predicts they will always happen together. Yes. But then we move to level two of Pearl's hierarchy, which is intervention or doing. Doing. Okay. This is where you don't just watch the world. You act upon it. You ask, what happens if I step in and change something? Okay. So sticking with the rooster. right what if we put a soundproof box over the rooster so it can't crow well i mean the sun is still going to rise exactly but the level one ai would be completely lost because its passive pattern has been broken right it would predict eternal darkness because it has literally never seen a morning without a rooster crowing precisely it doesn't understand the underlying mechanics of planetary rotation that makes a lot of sense of sense. And just to complete the hierarchy, level three is counterfactuals or imagining. Imagining, like what if scenarios. Yeah, asking what would have happened yesterday if the rooster had been silenced. Most AI models cannot even begin to touch level three. and they struggle immensely with level two. Because they can see patterns, but they don't know why the patterns exist. Exactly. You know, is it sort of like the difference between sitting in the passenger seat of a car for 10 years versus actually getting behind the steering wheel? Oh, I like where this is going. Like, imagine, I've never driven, but I've sat in the passenger seat of your car every day for 10 years. Just by passive observation, I can build a totally flawless predictive model in my head. Right. You've seen me do it a thousand times. Exactly. I know that every time your hands turn the wheel to the left, the car's chassis moves left. Every time your foot presses the right pedal, the car speeds up, my curve fitting is perfect. Yes. But I don't actually know why. I mean, I have no concept of a steering column or brake pads or a combustible engine. I just know the associations. That is an absolutely brilliant way to conceptualize it. And so if one day we're driving and the steering wheel suddenly locks up, or I don't know, we hit a patch of black ice and the wheels lose traction, my passive model just completely shatters. It's useless. Right. I have no idea how to correct this kid because I've never actually intervened to see where the physical limits of the car are. Yes. As a passenger, you have a mountain of data, but that data leaves the underlying physical mechanisms totally unidentified. You're just watching a movie, basically. Exactly. If you want to figure out how the car actually works or how the universe actually works, you cannot just sit in the passenger seat and watch. You have to take the wheel. You have to intervene. You have to turn it too hard on a slick road and see what happens. In the language of science, you have to poke the world. Poke the world. I like that. You have to design an experiment, apply a specific perturbation, and observe how the system responds to your direct action. And Richens and Everett's math proved this. Beyond a shadow of a doubt, they prove that for an AI to truly understand reality, it requires an interventional framing. It has to distinguish between what will normally happen and what would happen if I specifically chose to act. Okay, so that is the first massive conceptual roadblock cleared. We know the AI needs to stop just watching internet videos and start actually poking the world to find causal laws. Right. But the moment you say that, a huge practical problem jumps out at me. The cost. Exactly. Poking the world in real science is not like poking a physics engine in a video game where you can just reset the level. No, it is not. Running a clinical trial for a new cancer drug or launching a satellite to probe Jupiter's atmosphere. Right. or building a subterranean particle accelerator, these experiments cost millions, sometimes billions of dollars. And they take years, sometimes decades. Right. You can't just unleash an AI to randomly poke the real world a billion times to see what happens. It would literally bankrupt the entire scientific enterprise in an afternoon. And that very realization brings us to the second major roadblock in our journey. Okay. If we accept that AI must conduct experiments to learn, how on earth do we make sure it conducts the right experiments? How do we make it efficient? Right. This problem was tackled head on by researchers at Oxford, led by Adam Foster, who developed a framework called Deep Adaptive Design, or DAD framework. for short. Deep adaptive design. Okay, so how do they solve the efficiency problem? By leaning into a concept called Bayesian optimal experimental design. Oh boy. Okay, you're going to have to walk me through the Bayesian part of that. Yeah. Don't worry, let's strip away the intimidating math. The core philosophy here is about maximizing something called the value of information, or VOI. Value of information. It's an incredibly intuitive idea. It basically states that out of all the possible experiments you could run in the universe right now. you should always choose the one that you expect will reduce your current ignorance the most. So don't run an experiment that just confirms what you already know. Exactly. You want to run the experiment that is most likely to break a tie between your competing theories. Okay, give me an example. Let's say you are trying to figure out if a new material is superconductive. You have two theories. Theory A says it superconducts at room temperature. theory B says it only super conducts at absolute zero got it if you test the material at a hundred degrees Celsius both theories predict it will just act like normal metal testing it there gives you zero new information the value of information is near zero right because neither theory cares about a hundred degrees exactly but if you test it at room temperature theory a predicts a massive change and theory B predicts nothing That experiment has a massive value of information because whatever the result is, it instantly kills one of the theories. Okay, it reminds me of playing the game "Guess Who." Oh, that's perfect! Like, if you have 20 faces on the board, you don't ask, "Does your person have a nose?" Right, everyone has a nose. Exactly, the value of information is zero. You ask, "Is your person wearing a hat?" because it instantly eliminates half the board. That is a perfect parallel. You want to eliminate the most uncertainty with a single move. But, applying that logic to high-level science introduces a crushing computational bottleneck. Oh, because of the math involved. Right. In a real scientific scenario, you aren't just looking at 20 faces. You might be tracking hundreds of variables and thousands of competing hypotheses. I put a lot of hats and noses to keep track of. To calculate the mathematical value of information for your next move, a traditional computer system has to simulate the outcome of every possible experiment you could run. Cross-reference it against every possible hypothesis you hold, calculate the probabilities of all those outcomes, and then rank them. Which means the computer has to pause between every single experiment to do just a mountain of math. Yes. I would imagine that if an AI is trying to do this in real time, like dynamically adjusting a laser during a live physics experiment, It would just freeze up trying to calculate the absolute optimal next move. It totally would. Traditional Bayesian methods are agonizingly slow. You run a step, collect the data, and then you wait hours for the computer to update its beliefs and tell you what to do next. Which makes real-time adaptive experimentation basically impossible. Exactly. And this is the exact bottleneck the Oxford team solved with their D&D framework. They solved it through a concept called amortization. Amortization. My brain immediately goes to paying off a mortgage or a car loan over time. That's not wrong. How does that apply to AI doing math? The underlying logic is actually really similar. It's about restructuring when you pay a massive cost. Think about it this way. Instead of forcing the AI to do all that heavy computational lifting during the live experiment while the clock is ticking, what if we force it to pay that computational cost up front, offline, before the experiment even begins? Oh, like studying endlessly for months before a final exam, so that when you sit down at the desk, the answers just flow instantly. You've already done the hard work. Yes, exactly. Or think of a chess grandmaster. A grandmaster spends 10 years in their bedroom studying millions of board configurations. They pay a massive cognitive cost over a decade. So when they finally sit down at a real tournament and their opponent moves a night, the Grandmaster doesn't have to calculate every possible ramification of that move from scratch. They recognize the pattern of the board and they instantly know the optimal counter move. Exactly. Their calculation cost has been amortized. Ah, okay. So how does the AI actually do this? How does it study for the experiment? The Oxford team used what they call an amortized design network. They take a neural network and they train it offline on millions of simulated experimental scenarios. So it's practicing. Heavy practicing. They use a mathematical trick called contrastive information bounds to teach this network how to be a strategic thinker. Contrastive information bounds. Basically, they throw countless vague experiments at it, letting it learn which types of questions yield the highest value of information in different scenarios. Gotcha. By the time it's done training, it isn't just a calculator, it's a trained experimental policy network. So when you actually deploy it into the real world? When you deploy it live, the magic happens. You conduct step one of your experiment, you feed the raw data from that step into the trained neural network, and with a single lightning fast-forward pass of its circuits, taking mere milliseconds, it spits out the absolute optimal design for step two. That's incredible. No pausing, no hours of calculation. Exactly. The heavy math was already done in the simulation phase. That is incredibly clever. Yeah. And, you know, to ground this, let's look at the specific examples of this in action from the Oxford paper. Let's do it. One of the tests they ran was a 2D source localization task. Essentially, a highly complex mathematical game of finding a hidden object on a two-dimensional grid by taking signal measurements. Right. It's a very common problem in fields like sonar detection or environmental monitoring. You have a vast area, and you are trying to find the exact source of a faint signal. Right. And when they tested the DID system on this, the way it behaved was... totally alien to how a human or a rigid algorithm would do it. Yeah, humans are quite predictable, Eric. We are. A standard human design search pattern might just sweep the grid methodically, row by row, like mowing a lawn. Sure. But DAD had learned an adaptive dynamic strategy. It would take a few sparse measurements across the board to get a general feel. In the absolute millisecond, it detected even a hint of a signal in one quadrant. It would instantly cluster its next sequence of measurements tightly around that specific area to pinpoint the exact source. It was dynamically adapting. It was hunting. Yeah. And it was doing it in a fraction of a second, whereas the old Bayesian math methods would have taken hours just to decide where to look next. It was optimizing its information game dynamically. And, you know, they also tested it in a completely different domain psychology. Yes. The hyperbolic temporal discounting experiment. That's the one. That sounds like a mouthful. But it's basically mapping out a human being's threshold for delayed gratification, right? Exactly. It's the classic question. Would you rather have $10 today or $15 next week? Right. Psychologists use this to understand impulsivity and decision making. But every human is different. Some people will wait a week for an extra 50 cents. Some people won't wait a day unless you double the money. Yeah, we're all wired differently. So the goal of the experiment is to find the exact indifference point for a specific patient... the exact dollar amount and time delay where their brain genuinely cannot decide which option is better. And finding that exact point is really tricky. If you just ask random questions, the patient gets bored, they start paying attention, and your data gets incredibly noisy. They just start clicking, whatever. Exactly. You need to ask the perfect sequence of pinpoint questions... to narrow down their threshold as fast as possible. And that is exactly what the D-DAD system did. It was hooked up to human subjects, and it generated these highly adaptive, perfectly calibrated questions in real time in under a tenth of a second. It's just brilliant. Based on the human's answer to question one, it instantly calculated the question with the highest value of information for question two. It's that amortization network. Right. It was so efficient that it completely outperformed the specialized handcrafted questioning strategies that human psychologists had spent years designing for this specific test. So if we look at our journey so far, we have cleared two major hurdles. From DeepMind, we learned the AI must actively intervene and conduct experiments to escape the trap of curve fitting. Poke the world. Poke the world, exactly. From Oxford, we got the DID framework, which gives the AI the ability to design those experiments efficiently in milliseconds without getting bogged down in endless math. So we have the foundation for an active, strategic AI. We do. But, and there's always a, but here we are still operating in a somewhat idealized mathematical space. True. Finding a hidden point on a 2D grid or asking humans multiple choice questions, those are relatively clean environments. They are. What happens when the AI steps out of the theoretical math lab, and tries to apply this to the natural world, because the natural world is chaotic. It is messy. And a lot of times, the things we want to study are completely opaque black boxes. You've hit on the third massive roadblock, messiness. The messiness of reality. And this is where we turn to a 2025 breakthrough by a researcher named Jacob Deisler and a large collaborative team. Okay, what did they do? They published a comprehensive guide on something called simulation-based inference, or SBI. This framework was explicitly designed to tackle the chaos of real-world science. Let's get into the mess. What is the specific chaos they were trying to fix? Well, to understand the fix, you have to understand how standard Bayesian math works. Okay. In a traditional setup, if you want to figure out how likely your scientific theory is, you need a neat, clean mathematical equation. This equation is called an explicit likelihood function. Explicit likelihood function. Got it. It essentially says, "If my theory of gravity is correct, then the mathematical probability of this apple falling at exactly the speed is exactly 98.4 percent." It's like the physics equations we learned in high school. Force equals mass times acceleration. Clean, tidy, and mathematically absolute. Right. But in modern cutting-edge science, we almost never have those clean equations. We don't. Think about astrophysics when you're modeling the collision of two galaxies. Or climate science modeling the global weather over a decade. Or neuroscience trying to understand the firing patterns of a billion interconnected neurons. Yeah, that's way too complex for a chalkboard. Exactly. These systems are far too complex to be written down as a single neat equation. We don't have explicit likelihood functions for them. So what do scientists use instead? We build massive, incredibly complex computer simulators. We take all the tiny physical rules we think are operating how a single drop of water evaporates. Or how a single neuron fires, we program millions of them into a supercomputer and we just press play. And see what happens. We let the simulator run and see what kind of global data it spits out. But here is the critical problem. These simulators act as giant black boxes. Okay. You can put parameters in and you can get data out, but you cannot peek inside the code and extract a simple likelihood equation to do your Bayesian math. If you don't have the equation, how do you judge if your theory is right? Like, how do you update your beliefs when you get new data from a telescope or a brain scan? Historically, scientists relied on brute force. They used algorithms like Markov chain Monte Carlo or MCMC. And why is MCMC a problem? Because when you tie MCMC to a massive, complex simulator... It is agonizingly, painfully slow. Let me give you an analogy. Please do. Imagine you are dropped into the middle of the Himalayas, totally blindfolded, and your goal is to find the absolute highest peak. Okay, that sounds terrifying. It is. The MCMC algorithm basically says, take one step in a random direction. Did your foot go up or down? If it went up, take another step that way. If it went down, go back and try a different direction. Taking one step at a time, blindfolded, across the entire mountain range. Exactly. And now remember that in our scenario, every single step requires you to run the massive supercomputer simulator to see the result. Oh, yeah. If your climate simulator takes an hour to run a single scenario and your MCMC algorithm needs to take a million tiny steps to map the topography of the probability mountain. You are going to be wandering blindfolded for years just to test one hypothesis. Decades sometimes. It's totally impractical. That is completely unsustainable. You'd spend entire careers waiting for a loading bar. So how does simulation-based inference, or SBI, fix this? SBI represents a total paradigm shift. It takes the neural network technology we've been developing and uses it to replace the blindfolded wandering entirely. Okay, how? SBI says, "Instead of taking tiny steps during the experiment, let's run the simulator a bunch of times up front across a wide, scattered range of parameters. Let's generate a massive, messy data set of thousands of different possible outcomes." We mapped the whole mountain ahead of time. Sort of. Then we train a deep neural network on that data set to learn the hidden relationships between the input parameters and the output data. Wait, so the neural network studies the simulator's output, and then the neural network essentially becomes the inference engine? Yes. Once it is trained on all that simulated data, you can take this neural network out into the real world. You can hand it a real empirical observation, say, a messy, complicated brainwave recording from a human patient. It doesn't need to run the slow simulator again. It doesn't need to wander the mountain. It already knows the terrain. Exactly. It just does a single lightning fast calculation and instantly outputs the probability distribution. It looks at the brainwave and says,"Based on my training, There's an 80% chance this was caused by low dopamine and a 20% chance it was caused by high cortisol. That is mind-blowing. It just instantly maps the real-world chaos back to the hidden parameters. And it does it for systems where traditional math completely breaks down. the Deisler paper highlighted a specific example in neuroscience that really makes this concrete. Oh, the 31 parameter thing? Yes. They were looking at a biophysical model of neural circuits. Now this wasn't a simple on/off model. This simulator had 31 different biological parameters. So 31 dials you could Exactly. Things like the microscopic diameter of the dendrites, the exact density of sodium ion channels, the resistance of the cell membrane. It is a massive, high-dimensional parameter space. That sounds impossible to track. Trying to find the exact right combination of those Cumi1 dials to perfectly match a real biological brain recording seems completely impossible. And with the old blindfolded MCMC method, it would take centuries. But by using SBI, they didn't just find a single best fit set of dials. The neural network was so adept at understanding the deep structure of the simulator that it revealed something profoundly biological. It mapped out something called co-tuning. Co-tuning? What is that? It's the biological reality that nature is flexible. There isn't just one perfect way to build a functional neural circuit. The SBI analysis showed that these 31 biological parameters can actually compensate for each other. Oh, I see. So if one dial gets turned up too high, the system doesn't just crash. Right. If, say, the density of a certain ion channel is abnormally high, the circuit can still function perfectly normally if another specific parameter, like the membrane resistance, scales down proportionally to balance it out. Like a natural balancing act. Yes. The neural network, because it had analyzed the entire landscape of the simulator data, was able to map out these complex, hidden 2b conditional relationships between the parameters. It could see the balancing acts that nature performed. So what does this all mean? It means it allowed the researchers to take an impossibly complex, noisy reality and instantly understand the hidden mechanics driving it. Okay, so let's take a breath and look at the Frankenstein's monster we have built so far. It's quite the monster. It really is. We know the AI has to actively intervene to learn causal models. That's the deep mind realization. Check. We know it can design those interventions rapidly and cheaply using amortized Bayesian networks. That's the deep adaptive design from Oxford. Cool. And we know it can evaluate the results of those interventions against a messy black box reality. by using neural networks as inference engines. That's the simulation-based inference. We have successfully built the computational machinery for an autonomous AI scientist. We have the engine, but you know, an engine on a test block is different from a car on the track. Very true. Can this AI actually use all this machinery to discover something genuinely new? To find out, researchers had to design an ultimate test. And this brings us to the fourth major milestone, a 2026 paper by Wiemann and colleagues introducing the Discover Physics benchmark. The ultimate test. And honestly, this might be my absolute favorite part of this entire deep dive, because the premise of this test is straight out of a sci-fi novel. It is a brilliantly devious way to test an AI because, you know, the fundamental problem with testing a state of the art AI on Earth physics is that the AI has already read every physics textbook ever written. Right. It knows it already. It has ingested all of Wikipedia, all of Artea, all of human scientific history. If you drop an AI into a simulation and ask it to figure out how gravity works and it spits out Isaac Newton's equations, you have no idea if it actually discovered gravity or if it's just reciting a textbook from its memory banks. Exactly. To prove that the AI is actually doing science and not just engaging in elaborate data retrieval, Wyman's team realized they had to take the AI out of our universe entirely. So they created 22 curated alien worlds. These aren't just random environments with slightly different gravity. These are carefully, meticulously designed universes where the fundamental laws of physics have been systematically broken, altered, or replaced with entirely new concepts. I want to spend some time here because these alien worlds are just fascinating. Paint a picture for me. What kind of universes are we talking about? Let's start with the Yukawa world. Yukawa world, okay. In our universe, gravity is a force that extends infinitely, getting weaker over distance, but never truly disappearing. In the Yukawa world, gravity is what physicists call a screened force. What does that mean practically? Like, if I'm standing in the Yukawa world and I throw a baseball, what happens? Well, at first, it behaves exactly like you'd expect. The ball leaves your hand, and for the first hundred feet, it arcs toward the ground under a familiar gravitational pull. Normal baseball stuff. Right. But the Yukawa force has an exponential decay built into its math. It literally has a cutoff point. So mid-arc, as the ball reaches a certain distance from the planet, the gravitational pull just... shuts off completely. So the ball just stops falling. It stops accelerating downward. Whatever trajectory and momentum it had at that exact cutoff boundary, it just maintains forever. It would float away into space in a perfectly straight line, entirely ignoring the massive planet below it. Wow. It is entirely counterintuitive to everything the AI learned from Earth physics. That is wild. What else did they build? They built the ether world. This world breaks a foundational rule of our physics called coordinate invariance. Coordinate invariance. Yeah. In our universe, if you drop a bowling ball in New York and you drop an identical bowling ball in Tokyo, they fall at the same rate. The laws of physics don't care where you are standing. But in the ether world? In the ether world, the laws of physics are tied to absolute locations in space. The rules literally change depending on your coordinates. A pendulum might swing differently on the left side of the room than on the right side, simply because of its position. That would drive an Earth physicist completely insane. And that's the point. They also created the dark matter world. In this universe, there are visible particles that the AI can track, but there are also hidden, invisible particles swarming around them. Uh, spooky. These invisible particles exert strong physical forces on the visible ones knocking them off course Accelerating them inexplicably it would look like magic. It would look like objects are just moving themselves to solve the dark matter world The AI has to mathematically infer the existence of objects It cannot even see purely based on the unexplained anomalies in the data Yeah, it's a massive cognitive leap. And they also built worlds with nonlinear velocity dependent friction, where the faster you go, the rules of drag change in bizarre, chaotic ways. So the setup is brutal. The AI gets dropped into one of these 22 alien universes with zero instructions. It just has a blank slate. A total blank slate. Its job is to act like an explorer. It has to design experiments like saying, I'm going to place particle A at these coordinates, give it a velocity of 50 meters per second, and observe where it goes. Right. It watches the raw numerical trajectory data come back. All right. And then based purely on that data, it has to write a functional Python program that mathematically defines the new alien law of physics it just uncovered. And as a final hurdle, it has to explain that law conceptually in plain English. It is the ultimate test of autonomous scientific reasoning. Weeman's team unleashed 11 of the most advanced AI models in the world onto this benchmark. And the results were incredibly revealing about where AI currently stands. They really were. How did they do? Did they crack the alien of physics? Well, the massive proprietary frontier models, systems like Claude Opus 4.7 and GPT 5.5 perform the best. But even these multi-billion dollar models could only successfully solve about half of the 22 alien worlds. Only half. And what about the open source models like Lama 3.3 or Quinn 3.5? They lagged massively behind. They struggled to design informative experiments, often just repeating the same basic tests over and over. And they largely failed to adapt to the non-canonical physics. You know, what fascinates me is how these AI models failed. Because women's team didn't just give them a pass-fail grade on the math. They used a dual evaluation framework. Yes. Metric 1 was the Trajectory Means Squared Error, or MSE. The MSE. This is the purely mathematical grade it asks. how well does the Python equation the AI wrote actually predict the future movement of the particles right if the AI writes an equation and that equation can accurately guess where the ball is gonna be in 10 seconds he gets a good MSC score but metric two was the explanation score an independent panel of judges evaluated the plain English description the AI wrote to a company company it's math. They wanted to know. Does the AI conceptually understand why the particles are moving this way? And the divergence between these two scores is the punchline here. A high predictive accuracy, a low MSE, did not guarantee a high explanation score. In fact, they were often completely disconnected. Exactly. The AI models would often slip right back into their old bad habits of curve fitting. They were faking it. They were. An AI might observe a particle moving weirdly in a non-linear friction world. Instead of figuring out the true physical rule, the AI would just take a standard Earth equation, aggressively tweak a bunch of random mathematical parameters until the line sort of match the weird trajectory and submit it. So the math would technically predict the movement well enough to get a good MSE score. Yes. But when asked to explain it in English, the AI would write total nonsense. It didn't uncover a concept, it just brute forced a numerical approximation. It achieved low error without a single shred of true scientific insight. There was a specific failure mode in these tests that I think perfectly highlights why current AI isn't ready to be an independent scientist. The researchers called it being prior-bound. This is a devastating flaw for an explorer. Being prior-bound means you are too attached to what you already believe. It's hilarious, but also kind of tragic. You drop these AI models into the ether world or the Yukawa world, and they were so profoundly stubborn. So stubborn. They were so anchored to their training data on Earth physics that they would just try to force Newtonian gravity into a universe where it absolutely did not belong. When the AI ran an experiment and the data clearly contradicted standard gravity, instead of having a "eureka" moment and realizing it was in an alien world, the AI would often just dismiss the data as experimental noise. The data must be wrong. Exactly. It would rather believe its eyes were lying than revise its core hypothesis. The universe must be broken because my textbook is right. Precisely. It couldn't let go of its priors. And the models failed completely when dealing with latent structures like the dark matter world. Where things are invisible. Right. Inferring the existence of a hidden, invisible force acting on the visible world requires a leap of conceptual imagination combined with rigorous mathematical elimination. Current large language models on their own cannot do it. They just aren't built for that yet. They aren't rigorous enough to systematically hold three or four competing hypotheses in their head, design the exact experiments needed to eliminate them one by one, and generate a truly novel concept that exists outside their original training distribution. So we have reached the dramatic climax of our journey. We have seen that the current frontier models, as miraculous as they seem when writing poetry or generating images, hit a hard limit when faced with the unknown. They do. They are not autonomous discoverers. They are too biased. They are too easily confused by hidden variables. And they are far too eager to just curve fit a mathematical line rather than uncover the true mechanical gears of reality. They have all the individual pieces, but the pieces aren't working together. Which perfectly sets the stage for the breakthrough we have been driving toward, the culmination of everything we have discussed. In August 2026, Kevin Murphy at the University of British Columbia published the paper introducing the model discovery agent. or MDA. The MDA. This is where the Frankenstein's monster finally wakes up and walks. He is. Kevin Murphy basically looked at all the roadblocks, all the tools, the causal logic, the Bayesian experimental design, the simulation-based inference, and he engineered them into a single, cohesive, self-correcting system. He did. He built an architecture that forces the different parts of the AI to constantly check and balance each other. but his true stroke of genius the thing that makes mda fundamentally different from everything that came before it was how he utilized the large language model within this framework okay to understand why it's so brilliant we have to talk about the m open problem the m open problem this sounds like the theoretical crux of the whole system let's break it down. In traditional Bayesian statistics, you operate in what mathematicians call the M-closed setting. Think of it like walking into a restaurant and being handed a menu. A menu, okay. The menu has three items, hypothesis A, hypothesis B, and hypothesis C. The Bayesian math engine is an incredibly rigorous, flawless judge. You feed it experimental data, and it evaluates the data against the menu. Got it. It will tell you with absolute mathematical certainty, based on this data, hypothesis B is 90% likely, hypothesis A is 10% likely, and hypothesis C is impossible. It's the perfect referee. It just shuffles the probabilities around based on the evidence. Exactly. But what happens if the true law of nature, the actual reality of the universe you are studying, isn't on your menu? Okay. Uh-oh. What if the truth is hypothesis Z, a completely new concept you've never even thought of? The Bayesian referee is useless. It's going to confidently tell me that hypothesis B is the best answer, even though it's completely wrong, simply because it's the best of the bad options on the menu. You've hit the nail on the head. Standard Bayesian math cannot invent a new hypothesis. It can only shuffle probabilities among the ideas you already gave it. So that's the M-open problem. This is the M-open setting when the actual truth lies completely outside your current hypothesis class. And this is exactly what happens when you drop an AI into an alien universe or when a scientist encounters a truly novel biological phenomenon. So if the math can't invent new ideas, how does Kevin Murphy's MDA system solve the M-open problem? By radically redefining the role of the large language model. In the MDA system, the rigorous Bayesian math engine is constantly running in the background. It is constantly doing out-of-sample predictive checks. Predictive checks. That means it looks at the current data and asks, are any of our current theories actually doing a good job of predicting what happens next? It's constantly checking its own homework. Yes. And when the math engine sees the predictive error spike, when the particles start moving in ways the current theories can't explain the system flagged its own inadequacy, it realizes we are in an M-open situation. The truth is not on our menu. And that is when it tags the LLM in to help. Exactly. But it doesn't ask the LLM to do the math. This is crucial. Wait, really? Yeah. It hands the failing data in the environmental context to the LLM and essentially says, "I Our current theories are broken. Look at this unexplained residual error. Be creative. Dream up entirely new mechanistic structures, entirely new physical concepts that might explain this weird behavior. I want to make sure I fully grasp this, because it seems like the key to avoiding the failures we saw in the discover physics tests. The LLM, the AI that writes the poetry and passes the bar exam, is not running the Bayesian adates. It's not calculating the probabilities or tweaking the error margins. Not at all. And that division of labor is why MDA completely avoids the prior-bound stubbornness and the curve-fitting laziness that ruined the other models. That's fascinating. In MDA, the LLM is treated purely as the creative spark. It acts as the right brain of the scientist. It looks at the anomalies, draws on its vast knowledge base, and writes raw Python code proposing completely novel functional forms for the physical laws. It rewrites a new menu. Yes, it writes a new menu of hypotheses. And then it immediately hands that new menu back to the rigorous mathematical judge, the Bayesian machinery. Teamwork. The math engine then evaluates these new proposals, calculates their exact likelihood using simulation-based inference, and then crucially uses the value of information framework to design the absolute best, most highly targeted next experiment to test those specific new ideas. The AI proposes, and the math disposes. It is a perfect autonomous loop of discovery. And to prove how beautifully this works, Murphy's team ran MDA on something called Force Bench, which is a harder adaptation of that Discover Physics alien universe test. It's a brutal test. They took MDA and they dropped it into the Yukawa world, the universe, and the universe. with the screened gravity that abruptly shuts off. And remember, when they tested a standard, state-of-the-art LLM agent on its own in the Yukawa world, it completely failed. Right. Let's walk through exactly how the solo LLM failed and then how MDA succeeded, because the contrast is incredible. It's night and day. When the solo LLM got dropped into the Yukawa world, it started running a few basic, short-range experiments. Now, if you only look at objects falling from a few feet away, Yukawa's screen gravity looks almost identical to standard Earth gravity. It just looks like a simple mathematical power law. The short-range data fits perfectly into the AI's prior expectations. Right. So the solo LLM saw the short-range data, confidently guessed the standard power law, tweaked the parameters a bit to curve fit the numbers, got a decent MSE score on those short-range tests, and proudly submitted the wrong answer. It stopped looking. It was completely blind to the fact that 100 feet further out, gravity just ceases to exist. It never even bothered to look. But the MDA system, driven by its Bayesian design engine, didn't fall into that trap. Because it was optimizing for the value of information. The MDA system had its LLM propose several different ideas. The LLM said, well, maybe it's a standard power law like Earth, or maybe it's a screened force that cuts off. Right, it generated the mention. The Bayesian math engine looked at those two competing theories in the menu and realized something profound. It said, if we keep running short-range experiments, these two theories look mathematically identical. The data won't help us choose. Exactly. The only way to break this tie is... The only way to get a high value of information is to design a long-range probe. It actively hunted for the exact edge case that would break its own assumption. Yes. MDA deliberately fired a particle out into the distant, unmapped reaches of the simulated universe. It forced the universe to reveal its secrets. And the moment it did that, the moment it saw the numerical data showing the particle just floating away in a straight line instead of arcing back down, the true law snapped into focus. The data forced the shift. The mathematical evidence for the simple Earth gravity collapsed to zero. And the evidence for the screened Yukawa force guyrocketed. The researchers noted that MDA grokked the physics. It didn't just fit a curve. It understood the hidden mechanism. Here's where it gets really interesting. The numbers on this are staggering. When tasked with uncovering these alien laws, MDA hit a 74% exact mathematical recovery rate. Wow. 74% of the time, it flawlessly deduced the alien physics. The baseline, solo LLMs only hit 31%. A huge leap. But here's the absolute kicker. MBA achieved that massive 74% success rate while running five times fewer experiments. That is the astonishing power of combining a creative proposer with a rigorous, value-driven experiment designer. It doesn't waste time on redundant tests. Every single move it makes is calculated to slice away uncertainty. The LLM proposes a weird mechanism. The Bayesian math identifies exactly how to stress test it. The new data comes in to sharpen the forecasts. And if any unexplained weirdness remains, it triggers the LLM to dig deeper and discover the next subtle layer of truth. It had a tireless cycle of scientific rigor. But Kevin Murphy and his team didn't just stop at physics. Kept going. Oh, yeah. If MDA is truly a generalized autonomous discovery engine, it shouldn't just work on simplified particle physics in a simulation. Yeah. it should be able to tackle the complexities of other scientific disciplines. And the paper proves exactly that by marching the agent straight into the fields of chemistry and biology. Let's start with chemistry. They tested MDA on a benchmark called ChemBench, which was derived from a real-world platform called AutosciLab. Yes. The challenge here was for the AI to discover enzyme kinetic rate laws. Enzyme kinetics is a notoriously difficult field. It's all about understanding how biological catalysts' enzymes speed up or slow down chemical reactions. And it is incredibly complex. Right. The rate of a reaction depends on a massive web of variables. Okay. The concentration of the original substrates, the presence of various inhibitors, the ambient temperature, the pH level. It's a high dimensional soup. It is. Identifying the exact mathematical equation, the rate law that governs how fast a specific chemical reaction will happen in that soup is a grueling process for human chemists. So they said MDA loose on the chemistry data. How did it do? It was almost shockingly data efficient. MDA hit its absolute peak accuracy in uncovering the correct, complex, symbolic rate laws after conducting only about eight experiments. Eight experiments to map out a complex chemical reaction. Yes. To put that in perspective, the previous state-of-the-art AI system on this exact same benchmark required 60 experiments just to reach a significantly lower level of accuracy. So it's not just better, it's vastly faster. Exactly. MDA was vastly faster. But more importantly, and this echoes what we saw in the physics benchmark, MDA found the right kind of answer. It found the underlying truth. What do you mean by the right kind of answers? An equation is an equation, isn't it? Not in science. The researchers compared MDA against another highly advanced system called PYSR. PYSR is a symbolic regression tool. Okay. It is basically the ultimate curve fitter. It is incredibly good at crunching numbers and finding mathematical equations that perfectly fit a data set. But PYSR doesn't understand chemistry. It doesn't know what an enzyme is. Okay. When tasked with finding the rate law, PYSR would just throw out bizarre, incredibly convoluted mathematical expressions, equations with variables dividing by zero or raising concentrations to impossible negative powers. Yikes. The equations technically fit the numerical data perfectly, achieving low error, but the laws they described were physically impossible in the real world. It was drawing lines that matched the dots, but the lines didn't represent reality. Exactly. MDA, on the other hand, operates completely differently. Because it uses the large language model as proposer, it is guided by vast amounts of domain knowledge. The LLN knows the foundational rules of chemistry. Right. It's read all the textbooks. So when MDA investigated the data, it found actual interpretable biological mechanisms. Could you give me a specific example of what it found? Sure. It correctly identified a phenomenon called substrate inhibition. Substrate inhibition. In a normal reaction, the more reactant or substrate you add, the faster the reaction goes. But in substrate inhibition, adding too much reactant actually clogs up the enzyme and slows the whole process down. It's a very specific counterintuitive mechanism. And MDA didn't just stumble onto a math formula that simulated a slowdown. It actively proposed substrate inhibition as the physical reality and then designed the experiments to confirm it. Precisely. It didn't just find a math equation. It found the true chemical mechanism. It was doing genuine chemistry. OK, so MDA conquers alien physics and it maps out complex chemistry. Is there a harder test? Does it get more difficult than chemical soup? It gets much, much more difficult. Which brings us to the final and perhaps most impressive demonstration in Murphy's paper, biology. Biology. Biology is the ultimate boss fight for an AI scientist. To prove MDA could handle it, Murphy's team had to create an entirely brand new benchmark, which they called NeuronBench. Biology is the boss fight. I love that. What makes biology so much harder than physics or chemistry?
Two massive problems:partial observability and stochasticity. NeuronBench was designed to maximize both. Let's break those down. First, the setup of NeuronBench. What is the AI actually trying to do? They created six mystery neurons. These were highly complex, simulated biological cells based on generalized Hodgkin-Huxley models. which is the gold standard for modeling how neurons fire electrical signals. Okay. Each of these six mystery neurons was secretly hiding a novel, unknown ion channel. An ion channel is basically a tiny microscopic gate in the cell wall that opens and closes to let electricity ions flow in and out of the cell, right? That's how neurons talk to each other. Exactly. And the devious catch built into this benchmark is that under standard textbook testing, like just giving the cell a basic electrical shock to see what it does, You might have to apply a very specific sequence of electrical current clamps, or maybe simulate injecting chemical blockers to shut down all the normal known ion channels, just so you can isolate the activity of the mystery one. Wow, it's a biological puzzle box. It really is. And that covers the partial observability. You can only measure the overall electrical output of the whole cell. You can't peer in and see the individual microscopic gates opening and closing. You have to infer they are there. But what about the stochasticity? The noise. That is the real nightmare. Physics simulations are clean. If you run a physics simulation twice, you get the exact same numbers. Biology is wet, squishy, and infinitely noisy. Oh, right. In a real biological cell, ion channels do not open and close with perfect mathematical precision. They flutter. They open randomly due to thermal noise. They misfire. So the data you get back the electrical voltage trace of the cell isn't a smooth, clean line. It's a jagged, erratic mess of random biological noise. It is a chaotic mess. So the AI has to solve a partially observed puzzle box, and it has to do it while looking through a blizzard of random static. How on earth does a machine process that? If there is no clear equation and the data is just static. How does it do the Bayesian math to figure out if its theories are right? Traditionally, to do inference on a hidden, noisy system like this, statisticians use something called the particle filter. I feel like we're about to hit another computational bottleneck. We absolutely are. A particle filter is a powerful tool. It essentially tries to average out the noise by simulating thousands of parallel realities. Parallel realities. It says, okay, let's simulate this cell 10,000 times. each with slightly different random noise, track the hidden states across all 10,000 timelines and average them out to find the truth. That sounds incredibly thorough, but also overwhelmingly slow. It is paralyzingly slow. If the MDA system has to run a massive particle filter for every single crazy hypothesis the LLM proposes, across every possible experimental design it wants to test, the entire autonomous system just grinds to a halt. It would take weeks to run a single biological test. test. So Kevin Murphy needed a final trick, a way to bypass the particle filter. Exactly. And the paper details a genuinely brilliant workaround. To handle the noisy biology, MDA uses something called learned summary statistics. This is a beautiful piece of modern AI engineering. Instead of trying to painstakingly track every single microscopic flutter and random spike in that chaotic voltage trace, MDA learns how to compress the data. Compress it, okay. It uses a 1D convolutional neural network, a deep learning model, to look at the massive, messy voltage trace, ignore the meaningless static, and extract just the most crucial, informative features. It summarizes the data into a handful of key numbers. Okay, let's think of an analogy for this. It's like, if I want to understand how the global stock market performed today, I don't need a spreadsheet tracking every single penny fluctuation of all 10,000 companies second by second. No, that's way too much noise. Right. I just look at the Dow Jones and the S&P 500. Those are summary statistics. They compress millions of chaotic data points into a few numbers that tell me the overall truth of the system. That is a very apt comparison. By using this convolutional neural network to generate a clean summary, MDA can evaluate its biological hypotheses roughly 10,000 times faster than using a particle filter. It just bypasses the bottleneck entirely. It can exile. 10,000 times faster. But wait. There is a well-known trap with AI trying to summarize data like this. The paper mentions MDA had to overcome a massive hurdle called representational collapse. What is representational collapse? And why didn't it ruin MDA's summary? Representational collapse is a very common, very frustrating failure mode when you ask a neural network to compress data on its own. Let's use a non-math analogy. Imagine you have a deeply lazy student, and you assign them to read a 500-page dense novel and give you a one-sentence summary. Okay, you know a few of those. The lazy student realizes that the easiest way to finish the assignment without thinking is to just write. This is a book about people doing things. Technically true, but completely useless. Exactly. The student has collapsed the representation of the novel into a constant, meaningless value. Neural networks do the exact same thing if you aren't careful. They get lazy too. They do. If you just tell a network, make this giant file of biological data smaller, the network might realize that the most mathematically efficient way to minimize its processing effort is to just output the number five for every single test. All voltage traces equal five. Right. The data representation collapses into uselessness. To fix this, AI engineers usually have to resort to incredibly complex, hacky tricks things called stop gradients or running exponential moving averages across different networks just to force the AI to pay attention. So how did MDA avoid raising a lazy student? How did it force the network to actually learn the biology? By anchoring the student with a highly specific supervised objective, MDA doesn't just hand the data to the network and say, summarize this. It gives the network a strict test. What's the text? It says, "You must compress this data, but you must compress it in such a way that you can still use the summary to accurately predict the exact underlying physical parameters, like the ion channel density that generated the data in the first place." Ah, so you tell the lazy student, "Give me a one-sentence summary, but based only on that sentence, you must be able to tell me the main character's name, their motivation, and how the book ends." Yes. because the neural network is forced by that supervised objective to predict the underlying physical truth it cannot be lazy it cannot collapse it has to learn and preserve the specific features in the static that actually matter that's brilliant and the most elegant part of the system is that because it's a dynamic learning network it automatically adapts on the fly What do you mean? Imagine the LLM proposer suddenly comes up with a wild new hypothesis about a weird type of calcium ion channel that no one has ever seen before. The moment it proposes that, the convolutional neural network instantly retrains itself on simulated data of that new channel and it learns exactly which new features to look for in the static to prove if that specific thing works. theory is true. So did Dax in real time. It adjusts its own summary metrics instantly without a human scientist ever having to write new mathematical formulas. It is a completely 100% autonomous discovery engine. And the impact on the neuron bench test was phenomenal. MDA was able to drive its forecast error down to the absolute single-trial noise floor of the biological cell itself. Meaning it extracted literally every single drop of useful information that physically existed in that noisy environment. It completely dominated the tests. It consistently beat the baselines, correctly identifying the hidden silent ion channels time after time. by aggressively designing the exact sequence of electrical shocks and chemical blockades needed to force those hidden channels out into the open. It played the noisy puzzle box perfectly. It is a profound, almost breathtaking demonstration of what becomes possible when you synthesize the creative generative power of modern language models with the rigorous, skeptical, self-correcting engine of Bayesian experimental design. It really is. It feels like we are looking at the actual blueprint for the next century of scientific discovery. Let's take a step back and recap this epic journey we have been on, because the amount of ground we have covered is just staggering. It's been a long road. We started with the fundamental realization that AI can no longer just be a passive curve fitter. If it wants to understand the shifting nature of reality, it has to learn causal world models by actively intervening, by getting behind the wheel and poking the world. That was step one. Then we saw how the deep adaptive design framework solved the cost of those interventions, allowing the AI to calculate the value of information. and design pinpoint experiments in milliseconds by amortizing the heavy math up front. Step 2. We explored how simulation-based inference allowed the AI to test those theories not against clean equations, but against the messy black box reality of supercomputer simulations and real-world noise. And Step 3. We watched early highly advanced AI models fail spectacularly in the alien universes of the Discover Physics benchmark. blindly forcing standard gravity into worlds where it didn't belong, proving that language models alone lacked the rigor for true science. They needed a partner. Right. And finally, we arrived at Kevin Murphy's masterpiece, the Model Discovery Agent, a system that conquers the M-Open problem by using an LLM to creatively dream up entirely new hypotheses when old ones fail, uses Bayesian math to rigorously judge those wild ideas, and designs highly efficient, active experiments to actively prove itself wrong. It's incredible. We watched it uncover the hidden mechanics of screened gravity, define the complex rules of chemical enzyme kinetics, and successfully hunt down silent ion channels in the chaotic, noisy static of wet biology. And I think the essential takeaway for anyone listening to this isn't fear. It isn't that MDA or systems like it are going to render human scientists obsolete or replace researchers. Far from it. What this technology represents is the creation of a tireless, rigorously logical partner. A partner that isn't intimidated by massive, noisy, 31-dimensional biological data sets. A partner that doesn't suffer from human bias, that can boldly navigate the unknown, propose completely left-field hypotheses that a human might be too conservative to suggest, and systematically design the exact experiments needed to map out the complex causal realities of our universe. It is a tool designed to exponentially accelerate human understanding. Exactly. A tireless partner in the dark, shining a flashlight into the unknown corners of physics, chemistry, and biology. Which leaves me with a final thought for you to mull over as we wrap up this deep dive. Okay, let's hear it. We have seen today that an AI agent can be dropped into a simulated alien universe with zero instructions, realize its own Earth-based assumptions are wrong, mathematically deduce the existence of a screened Yukawa gravity field, and actively design long-range probes to prove it. We did. We've seen it uncover hidden, silent mechanisms in the messy static of biology. It has proven it can uncover the hidden causal rules of complex systems. So what happens when we eventually point these causal discovery agents away from cells and particles and focus them on the incredibly messy, stochastic, deeply complex data of our own human systems? Oh, wow. If an agent like MDA can find the hidden laws of alien physics... Could it eventually analyze massive data sets to discover fundamental underlying social laws or economic mechanisms that we are currently too close or too prior bound to see? Could it find the hidden math of human behavior? That is a question that opens up entirely new frontiers of scientific curiosity. If the tools of discovery are universal, the next decade is going to be incredibly illuminating. It certainly is. We will leave you with that thought to chew on. Thank you for joining us on this exploration of the model discovery agent and the incredible journey to build an artificial scientist. We've moved away from the muddy, confusing waters of passive AI observation, and we are finally starting to see the jagged, visible underlying gears of true causal understanding. Until next time. Keep diving deep.
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Hidden Brain
Hidden Brain, Shankar VedantamAll In The Mind
ABC Australia
What Now? with Trevor Noah
Trevor Noah
No Stupid Questions
Freakonomics Radio + Stitcher
Entrepreneurial Thought Leaders (ETL)
Stanford eCorner
This Is That
CBCFuture Tense
ABC Australia
The Naked Scientists Podcast
The Naked Scientists
Naked Neuroscience, from the Naked Scientists
James Tytko
The TED AI Show
TED
Ologies with Alie Ward
Alie Ward
The Daily
The New York Times
Savage Lovecast
Dan Savage
Huberman Lab
Scicomm Media
Freakonomics Radio
Freakonomics Radio + Stitcher
Ideas
CBCLadies, We Need To Talk
ABC Australia