Heliox: Where Evidence Meets Empathy π¨π¦β¬
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Heliox: Where Evidence Meets Empathy π¨π¦β¬
The Genius in the Concrete Room
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
π Read
How AI Is Learning to Act, Remember, Evolve, and Collaborate
What if the smartest mind on Earth were locked in a concrete room, brilliant at calculus but unable to fetch groceries? That's the problem behind today's AI, and this episode of Heliox explores the research roadmap for opening the door.
This is a science podcast deep dive into agentic reasoning, the shift from AI that only talks to AI that plans, uses tools, remembers, evolves, and collaborates. Drawing on a major survey from researchers at the University of Illinois Urbana-Champaign, Meta, Amazon, Google DeepMind, UC San Diego, and Yale, we unpack:
- Why large language models are passive oracles, and what partially observable environments (POMDPs) reveal about their limits
- How agentic search, tool use, and vibe coding turn AI into an active collaborator
- Why agents suffer from amnesia, and how reflective feedback, structured memory, and self-evolving skills fix it
- How multi-agent systems use planners, workers, and adversarial critic agents, and why theory of mind matters
- Real-world applications in AI-driven scientific discovery and longitudinal healthcare
β’β’The open problems: credit assignment over long horizons, and governance and safety for AI that acts
Reference
Agentic Reasoning for Large Language Model
Chapters
00:00 Welcome to Heliox: The Genius in the Concrete Room
01:35 When Bigger Models Hit a Wall
02:48 The Supergroup Behind the Roadmap
04:38 Passive One-Shot Inference Explained
06:30 POMDPs: Why Reality Isn't Chess
09:30 Factorized Policies: Thought vs. Action
11:17 In-Context vs. Post-Training Reasoning
14:07 GRPO: Grading Reasoning on a Curve
15:12 Agentic Search vs. Traditional RAG
17:30 Vibe Coding and OpenHands
19:35 The Amnesiac Agent Problem
21:21 Retry vs. Reflective Feedback
23:12 Flat vs. Structured Memory
24:41 Procedural Evolution: Voyager's Skill Library
26:32 Structural Evolution and AlphaEvolve
28:45 Why One Genius Isn't Enough
29:27 Multi-Agent Reasoning: Talking as Thinking
31:13 Managers, Workers, Critics, Memory Keepers
32:19 From Scripted Roles to Team Evolution
33:25 Theory of Mind in AI
36:32 Critic Agents: Peer Review for Robots
38:02 Real-World Deployment: The AI Scientist
41:07 Simulated Peer Review
42:46 Healthcare and Lifelong Clinical Coherence
44:40 What Are Humans For? Directors, Not Laborers
46:10 The Credit Assignment Problem
48:39 Governance and Safety for Acting Agents
50:37 Recap: From Concrete Room to Society
52:11 The Final Provocation
53:38 Credits and Outro
This is Heliox: Where Evidence Meets Empathy
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Disclosure: This podcast uses AI-generated synthetic voices for a material portion of the audio content, in line with Apple Podcasts guidelines.
We make rigorous science accessible, accurate, and unforgettable.
Produced by Michelle Bruecker and Scott Bleackley, it features reviews of emerging research and ideas from leading thinkers, curated under our creative direction with AI assistance for voice, imagery, and composition. Systemic voices and illustrative images of people are representative tools, not depictions of specific individuals.
We dive deep into peer-reviewed research, pre-prints, and major scientific worksβthen bring them to life through the stories of the researchers themselves. Complex ideas become clear. Obscure discoveries become conversation starters. And you walk away understanding not just what scientists discovered, but why it matters and how they got there.
Independent, moderated, timely, deep, gentle, clinical, global, and community conversations about things that matter. Breathe Easy, we go deep and lightly surface the big ideas.
Spoken word, short and sweet, with rhythm and a catchy beat.
http://tinyurl.com/stonefolksongs
So imagine you have the smartest person in the entire world, like a true once-in-a-generation genius. Okay, I'm picturing it. Now imagine taking this absolute genius and just locking them inside a solid concrete room. A little dark, but sure. Right, but stick with me here. There is no internet connection in this room. There are no books on the walls. There's absolutely no way to look out the window or even hear what's going on in the street outside. Right. They are totally isolated. Exactly. Now, if you slide a piece of paper under the door with some highly complex unsolved calculus problem on it, They solve it instantly. Because they're a genius. Right. They slide it right back to you. Flawless logic. Perfect work. But if you slide a note under that exact same door that just asks them to go to the local store and buy groceries for the reg. They can't do it. They are completely paralyzed. Yeah. They just sit there. I mean, it makes sense. They possess all the cognitive capacity in the world. Yeah, yeah. An immense repository of synthesized knowledge. But they have literally zero capacity to interact with the environment. They are entirely disconnected from the physical and temporal reality of the world outside that concrete room. And this thought experiment, it's not just some fun philosophical riddle for you to ponder. It is the exact frustrating wall that artificial intelligence researchers hit recently. He really is. Because we have spent the last few years building these massive, traditional, large language models. And they're absolutely brilliant in a static, closed world. Oh, yeah. Feels like magic. Right. You ask them to write a sonnet in the style of Shakespeare about quantum physics or solve some highly intricate coding puzzle, and they just do it. But if you ask them to actually navigate reality. Right. To take an action, observe the result, and adapt, they are helplessly passive. And that frustration really comes into focus when you consider the sheer scale of computation we've thrown at this problem over the last few years. We've thrown a lot at it. So much. The entire industry's thesis for a long time was simply, you know scale it up build bigger models increase the parameter count into the trillions exactly feed them more training data build massive data centers to power them the whole assumption was that if you just make the brain big enough it will eventually figure out how to do things but a passive Oracle no matter how large its neural network is still just sitting in that concrete room waiting for a piece of paper to be slid under the door Which brings us to the actual mission of our deep dive today. Right. Today we are exploring a massive paradigm-shifting roadmap that essentially figures out how to bust the door of that concrete room wide open. And this isn't just a single study we're looking at. No, it's not. It is a comprehensive survey and theoretical framework published by a veritable supergroup of researchers. We are talking about scientists from the University of Illinois, Urbana-Champaign-Meta, Amazon-Google, DeepMind, UC San Diego, and Yale. A huge collaboration. Researchers like Tiang Sunwei Tingwei, Li Zhenning Liu, and a massive team of their colleagues. And I think it's worth pausing to understand why all these researchers from competing institutions actually came together. Why did they? Because they looked at the landscape of AI and realized the field was basically violently splintering. Splintering how? Well, you had one lab working on prompt engineering, another lab working on getting AI to use calculators, another trying to get AI to play video games. Everyone was just in their own little silo. Right. And the authors of this paper realized that the community lapped a unified theory for the next era of artificial intelligence. We really needed a common language and a shared roadmap to understand how we move past this passive oracle. They recognized that the era of just scaling up test time computation, which is essentially what you do when you just ask a traditional language model to think harder about a static prompt, was reaching its natural limit. Yes, we were looking at the precise moment the AI community stopped focusing on teaching models to just talk started teaching them to act to act to evolve and incredibly to collaborate right they are defining a shift from scaling computation to scaling test time interaction they call this new era a genetic reasoning the gentick reasoning okay to really grasp the magnitude of that we have to deeply understand the limitations of what we were doing before for right because the researchers categorize traditional language models as operating on passive one-shot inference let's pull that apart for a second passive one-shot inference so when I log into a standard AI chat interface I type a prompt and the AI types back that's the one shot yes exactly You provide a static input. The model then does its internal computation. It just thinks about it. Right. It might even use clever prompting techniques we've developed over the years, like a chain of thought. Where you ask the model to break the problem down step by step internally before giving the final answer. Exactly. But mechanically, it is still a single forward pass through the model's neural network. It calculates the statistical probability of the next word generates it and sobs. It spits out a static output. Right. Wait, I want to challenge this idea of it being a dead end, though, because calling it passive makes it sound kind of useless. I wouldn't say useless. But we've all seen AI generate like 20 page legal briefs that are high. highly accurate. It passes the bar exam. It can diagnose medical imagery in a single pass. Oh, absolutely. So calling that a dead end feels a little uncharitable to the incredible engineering that went into it, don't you think? It is a phenomenal achievement. I agree. But we have to distinguish between a closed world domain and an open world domain. Okay. What's the difference? Well, writing an essay, even a brilliant one, is a closed world task. The model has all the information it will ever need right there in its context window. The text you provided and its pre-trained weights. Exactly. The environment isn't changing while it writes the essay. The essay doesn't fight back. Right. The paper is just a piece of paper. But the dead end reveals itself the moment you place that exact same brilliant model in an open-ended dynamic environment. Ah, I see. The authors of the paper explain this using a concept from decision theory called a POMDP. Okay, POMDP. That is a heavy acronym. It is a bit of a mouthful. I know it stands for a partially observable Markov decision process. But if you know, if I don't have a PhD in mathematics, how should I visualize what a QOMDP actually is? Let's build up to it. First, consider a fully observable environment. Chess is the classic example here. Okay, chess. If you sit down to play chess, You can see all 64 squares. You can see every piece your opponent has. Nothing is hidden from you. Right. You have perfect information about the state of the board. Right. And that is an MDP, a Markov decision process. OK. So the AI knows exactly what it's dealing with in an MDP. Yes. But reality is almost never an MDP. reality is partially observable give me an example of that think about driving a car in a new city you don't have a top-down omniscient view of every other car every pedestrian in every upcoming traffic light no you only see what is three your windshield exactly you have to make a decision like say merging into a lane based on incomplete information okay I'm with you so you make the move you observe how the environment reacts Maybe the car next to you honks. And then you update your understanding of the world based on that honk. Right. That is the partially observable part. Yes. And traditional language models are fundamentally unequipped for this because they lack statefulness. Meaning they don't have an ongoing updated understanding of the world outside of the exact words in their current prompt. Exactly. They only know what is in their immediate context window at the exact moment of inference. Right. So if the world changes while they are generating their response, they are completely blind to it. They have no idea. I see. So a traditional LLM is basically like a car GPS that calculates your entire route before you leave the driveway. That's a great analogy.
Right? Like it gives you the perfect directions based on the map it has at 8:00 in the morning.
But if a tree falls on the highway at 8:15, The GPS doesn't know. It can't recalculate, it just cheerfully tells you to drive at 60 miles an hour straight into the fallen tree. That captures the danger perfectly. And the researchers realized that all the clever prompting tricks we were using were basically just band-aids. Because asking an AI to decompose a task into smaller steps doesn't help if the environment changes between step one and step two. Right. To truly solve open-ended problems, the AI needs a continuous feedback loop. It needs to formulate a hypothesis. Take an action in the real world, observe the consequence from that action, and then physically change its internal state or plan based on that new data. It needs to be able to look out the window of the concrete room. Exactly. It needs hands to manipulate things and eyes to see what happened. And that transition from the locked room to actually interacting with reality is what the researchers define as foundational agenda. reasoning. Yes, that's the first big step on their roadmap. And this is where the paper introduces a concept that I just found theoretically beautiful. They talk about shifting the architecture to a factorized policy. It's a really elegant solution. Let's unpack the mechanics of that. What exactly are we factorizing here? So the researchers expressed this shift mathematically in the paper, specifically in equation one. What they are doing is explicitly separating the AI's internal cognitive process from its external actions. Let's slow down and look at that conceptually. In a traditional model, how are thought and action handled? In a traditional language model, thinking and doing are identical. What do you mean? Well, generating the next token, the next piece of a word, is literally the only action it can take. Oh, right. It just types. But in a factorized policy, you split that into two distinct probabilities. First, you have a latent reasoning trace. Latent reasoning trace, okay. This means the AI thinks in a hidden internal space first. It models the problem, weighs multiple options, and formulates a plan. That is purely internal thought. It's not typing anything out for the user yet. Exactly. Only after that latent reasoning is complete does the model calculate the probability of an external action like executing a line of code or querying a database. So it's the cognitive difference between blurting out the very first word that pops into your head versus taking a deep breath internally drafting the sentence, considering how it will be received, and then finally opening your mouth to say, speak. That is exactly it. And by separating those two things, you can optimize them independently. Which is huge. It is. The researchers identify three foundational capabilities that emerge when you factorize the policy like this. Planning tool use and search. Okay, so if I want to build one of these foundational agents, the paper outlines two fundamentally different ways to do it. do it. Two optimization modes, right? The first is in-context reasoning, and the second is post-training reasoning. Let's tackle in-context reasoning first. How does it work? So in-context reasoning is fascinating because it doesn't actually change the underlying neural network. The model itself remains entirely frozen. Wait, so how does it learn to be an agent then? It relies on highly structured orchestration and sophisticated prompting frameworks to basically force the AI to behave agentically. What kind of framework? A classic example mentioned in the paper is React. which stands for reason and act. Instead of just giving the AI a tromped and letting it run the React framework, forces the AI to output text in a very specific loop. It must output a thought, then an action, and then it waits for an observation from the environment before it is allowed to output its next thought. Ah, I see. Are there other frameworks like that? Yeah, another big one is tree of thoughts. Oh, I've heard of that one. That's where it explores multiple paths at once, right? Yes. It allows the AI to branch out. If it hits a complex problem, it doesn't just guess one solution. It generates three different possible next steps, evaluates the likelihood of success for each branch, discards the bad ones, and pursues the most promising one. That sounds incredibly thorough. It is. It is. But again, all of this is happening in the prompt in the context window. So in context reasoning is essentially like handing a worker a highly detailed, rigid instruction manual on their first day on the job. That's a perfect way to describe it. The worker's actual brain hasn't changed. They don't inherently know how to do the job. But if they followed the checklist, step one, step two, use this tool, step three. Check your work. They can successfully complete complex tasks. The manual analogy is spot on. The limitation, of course, is that the manual takes up a lot of space. Right. These prompts become massive and computationally very expensive. You are stuffing the context window with instructions instead of the actual problem. Which makes sense why the researchers point to the second optimization mode then, post-training reasoning. Exactly. Yeah. This is a fundamentally different approach. Here you are actually altering the model's weights. You are changing the brain itself. Yes. You are baking these agentic behaviors directly into the neural network using reinforcement learning. So if in context is giving them a manual to read on the job, post-training is like sending the AI to an intensive trade school. Right. You put them through so many drills and simulations that the skills, the planning, the tool, use checking their work, it all becomes muscle memory. You don't need to hand them a manual anymore because it's physically wired into their synapses. The trade school analogy works really well because it highlights the training process itself. The paper extensively discusses methods like GRPO. GRPO, Group Relative Policy Optimization. Okay, I definitely need you to translate GRPO for us. How does a reinforcement learning algorithm actually teach an AI to plan? So traditionally in reinforcement learning, you might reward a model based on an absolute baseline. Right. Good job or bad job. Okay. But GRPO optimizes the AI by letting it generate a whole group of different reasoning paths for the exact same prompt. It then compares these paths against each other. Oh, I see. It's grading on a curve. Well, actually, yeah, that is a highly accurate way to frame it. Thank you. It looks the batch of responses and mathematically rewards the model for the specific paths that lead to the correct final outcome while penalizing the paths that wander off into the last. hallucinations or bad logic. And it does this over and over again. Through millions of iterations of this, the model's internal weights physically adjust. It internalizes the structure of good reasoning. So when you give it a problem post-training, it naturally pauses plans and reaches for tools without needing a massive prompt to tell it to do so. Precisely. It just inherently knows how to be an agent. And speaking of tools, one of the most vital tools a foundational agent needs is search. Absolutely. I found the paper's distinction here incredibly illuminating, actually. They make a sharp line between traditional RM and what they call agentic search. It's a crucial distinction. Now, ARG retrieval augmented generation is everywhere right now. It's when an AI looks up a company PDF or Wikipedia page to help answer a question. How is a gentic search mechanically different from that? Because traditional ARG is fundamentally a passive pipeline, you, the human user, ask a question. The system takes your exact words, turns them into a vector, searches a database, retrieves whatever document seems most mathematically similar, feeds that document to the language model, and then the model summarizes it. It's a straight line from user to document to answer. Right. It's completely linear. So the AI has no agency in that process. It's just reading what it's handed. The AI is essentially a passenger in traditional R. Agentic search, on the other hand, as demonstrated in systems like paper QA, turns this into an active epistemological loop. An active loop, meaning the AI is driving. Exactly. The AI itself decides when it has a knowledge gap. It formulates its own query based on its internal reasoning. It executes the search. It reads the retrieved documents. And then, and this is the crucial step, it critically evaluates whether those documents actually answered its specific question. And what if they didn't? It doesn't just spit out a bad answer anyway, which is what traditional art often does. It realizes the search failed. So it revises its search terms, maybe tries a different database and searches again. It might retrieve a document. realize the document makes a factual claim without a citation, and then autonomously launch a secondary search just to verify that specific claim. It's the difference between handing a student a textbook with a bookmark already placed on the right page versus teaching the student how to walk into a university library, use the digital catalogs, pull five different books, realize three of them are outdated, and synthesize the remaining two to form a thesis. That's exactly it. The agent acts like a true researcher. And when you combine that active search with the ability to plan and use tools, you unlock some stunning real-world applications right at this foundational level. Like what? Well, the survey references a concept that has been gaining a lot of traction recently, vibe coding. I love that term, vibe coding. I know researchers like Sep Coda and Andres Carpathy have talked about this shift, but I have to say vibe coding sounds like something you do at a music festival, not in a software engineering department. What is it actually describing? I know it sounds casual, but it describes a massive paradigm shift in human-computer interaction, facilitated by foundational agents like Open Hands. Open Hands. The paper calls that a repository-level agent, right? Yes. In Vibe coding, the human programmer no longer writes the software syntax line by line. They don't type the code at all. No. The human provides the Vibe, the high-level intent, the business logic. The ultimate goal. You tell the agent I want a web application that tracks local weather, cross-references it with my calendar, and texts me what coat to wear. So you give it the destination but not the map, and then the agentic system just takes over the keyboard. The foundational agent takes that high-level intent and utilizes its factorized policy. It enters its latent reasoning trace. It breaks your goal down into dozens of subtasks. It decides it needs a front-end framework, a weather API, and a calendar integration. It plans the whole architecture. It plans it and writes the code. But it doesn't stop there. It uses tools to actually compile and run the code. And code never works on the first try? Never. It's going to crash. Right. But the agent looks at the error logs. It actually uses agentic search to look up that specific error code on Stack Overflow or an official documentation. It searches Stack Overflow by itself? Entirely by itself. It figures out what went wrong, edits its own code, runs it again, and loops this process until the application actually functions. That is genuinely mind-blowing. The human provides the constraints and the intent, and the agent manages the entire execution. layer. Exactly. It feels like we've solved the puzzle right there. But as incredible as vibe coding and agentic search are, the researchers point out that these foundational agents inevitably hit another dead end. They do, unfortunately. They are brilliant at solving the problem in front of them, but they suffer from a fatal flaw. They are amnesia. Right. They possess no mechanism for long-term statefulness across multiple tasks. Break that down for me. Let's say your foundational agent spends three hours debugging a very specific, obscure conflict between two software libraries to build your weather app. Okay. It solves it. The app works. But the moment you close the session, that knowledge is gone. Erased. If you ask it to build a similar app the very next day, it will make the exact same mistakes and spend another three hours solving the exact same obscure error. It's like having a brilliant employee whose memory gets wiped every night while they sleep. They solve the problem, but they don't actually learn from the problem. The researchers concluded that tools and planning architectures weren't enough. For true intelligence to emerge, the agent needed to evolve over time. It needed a mechanism to accumulate competence. Which brings us to the next major evolution in the roadmap, the amnesiac agent and the meta-learning loop. This section of the paper focuses on self-evolving reasoning. Yes.
We can distinguish it this way:Foundational reasoning is about optimizing actions within a single episode or task. Self-evolving reasoning is about optimizing the agent's fundamental capabilities across multiple episodes. Across time. Exactly. It introduces a meta-update rule. A meta-update rule? How does that work in practice? The agent must take environmental feedback, like a reward score or the text of an execution error, and use it to update its own internal state or its own tools. It has to close the loop between having an experience and gaining competence from that experience. So it's finally building a memory. Yes. But memory isn't just a single concept. The researchers make a very sharp distinction between the types of feedback an agent can process. They contrast retry-based systems with reflective feedback. This is a huge distinction. So if I'm an agent trying to write a Python script, What is a retry based approach? A retry based system is essentially automated trial and error. The AI generates a solution and submits it to an external validator like running a suite of unit tests on the Python script. Okay, so it tests the code. If the test fails, the agent discards the solution and tries again. It might randomly perturb its approach, but it doesn't really analyze the fundamental mechanism of why it failed. It just knows it received a negative signal. It's basically brute forcing it. Pretty much. Like trying to guess a four-digit padlock by just spinning the dials as fast as you can. It works eventually, but it's not exactly deep learning. It is effective in environments where checking the answer is incredibly cheap and fast, like basic coding or math. But reflective feedback is a much deeper cognitive process. What does that look like? The survey highlights systems like reflection. In these systems, when the agent fails, it doesn't just try again. It is prompted to look at the specific error logs, synthesize the context of what went wrong, and generate linguistic cues. Linguistic cues, like it talks to itself. It generates actual text reflections analyzing its own failure. Oh, wow. It might write something like,"I attempted to query the database using the string format, but the API requires an integer." In the future, I must cast the variable to an integer before querying. And it saves that. It saves this reflection right into its memory. It is the algorithmic equivalent of leaving a post-it note on your monitor so you don't make the same mistake tomorrow. That brings up a massive logistical problem though. If an AI is running thousands of tasks and leaving itself thousands of post-it notes⦠Where do they go? Exactly where do they go? The paper dives into the architecture of AI memory, distinguishing between flat memory and structured memory. The filing system is everything here. Flat memory is essentially a chronological log. It's like a massive chat history containing everything the agent has ever done, thought or reflected upon. Which sounds useful until you realize how much data that actually is. It becomes overwhelming almost immediately. As the context window fills up with thousands of past interactions, the AI's attention mechanism gets completely diluted. It struggles to find the relevant post-it note among a sea of irrelevant ones. I can't see the forest for the trees. So how does structured memory fix that? Structured memory solves this by persisting these intermediate reasoning traces in an organized relational way. Like a database. Usually like a knowledge graph or a hierarchical vector database. Instead of just dumping a memory into a timeline, the agent categorizes it. It links the memory to specific tools, specific error codes, or specific environmental states. That makes total sense. This allows the AI to selectively retrieve exactly the procedural knowledge it needs for a new task. without loading its entire life history into its working memory. It transforms raw chronological factual recall into highly specific procedural competence. It's the difference between trying to find a specific receipt by digging through a giant trash bag full of every piece of paper you've touched this year versus opening a meticulously labeled filing cap. Exactly. And this structural organization leads us to the three specific types of evolution the paper outlines. We just touched on the first one, verbal evolution. Right. That's the agent updating its text reflections, the Post-it notes in the reflection system. But text only goes so far. What is the second type of evolution? The second is procedural evolution. And this is a massive leap. Here, the agent doesn't just leave itself at vice. It builds a permanent library of executable tools and skills. Okay. The canonical example the paper cites for this is a fascinating system called Voyager. Voyager. Wait, this is the one that plays the video game, right? Yes. It is an AI system designed to play Minecraft, which is a highly complex, open-ended 3D environment. Sure. When Voyager starts, it doesn't know how to do much. But as it explores, it attempts to solve problems. Let's say it needs to figure out how to mine iron ore. It goes through a whole cycle of trial error and reflection to successfully mine iron. Okay, so it learned how to do it once. But here's the procedural evolution. Once Voyager successfully mines the iron, it extracts the exact sequence of code and actions that led to that success. packages it as a new discrete function called Minaron, and saves it to a permanent skill library. Wow. Wait, it literally writes its own tools. Exactly. The next time Voyager needs iron, it doesn't have to plan the micro steps of finding a pickaxe locating a cave and striking the block. It just calls its own Minaron tool. That is incredible. As it plays the game, it permanently expands its own action space. It is literally writing its own software tools as it goes. It's automating its own lower level bodily functions so its higher level brain can focus on bigger problems. That is wild. It is very powerful. But here is where we hit the part of the paper that I found really interesting and frankly a little terrifying. Okay, what part? The third type is structural evolution. The paper talks about systems like AlphaVolv, Are the researchers saying they have built AI that rewrites its own core architecture? Like, does it mutate its own source code? I understand the apprehension. Because that sounds exactly like the sci-fi technological singularity we have all been warned about. It does sound like the plot of a science fiction film. But if we ground this in the mechanics described in the paper, it is a very logical, controlled search process. Okay, calm me down then. It is not an AI waking up and deciding to hack its own servers. What it means is that the agent treats its own underlying reasoning code as a hypothesis space. Hold on, let's unpack that. What does it mean for code to be a hypothesis space? Normally, when we optimize an AI, we use gradient descent. We use calculus to find the mathematical slope that minimizes errors, and we adjust the neural weights accordingly. Right, standard machine learning stuff. But you can't use calculus to optimize the text of a Python script or a planning algorithm. It's discrete. It's not a smooth mathematical curve. So to evolve structurally, the system performs a gradient-free optimization step. So if it can't use math to slide down a hill, how does it actually learn? It uses an LLM as a mutation operator. A mutation operator. The system looks at the algorithms driving its own reasoning. Maybe the specific loop it uses to evaluate search results. The mutation operator analyzes that code and says, what if I tweak this search parameter? What if I add a verification step before I commit to an action? It generates a mutated version of its own code. Exactly. It hypothesizes a better version of itself. Okay. And then it compiles that mutated code, creates a clone of itself using the new architecture, and tests the clone on a benchmark suite. Wow. If the mutated agent performs better, if it solves problems faster or with fewer errors... The system keeps the change, it throws away the old code, and adopts the new code as its baseline. So it is searching for superior reasoning algorithms through iterative evolutionary refinement? Yes. Okay, so it's less like Skynet becoming sentient and more like a highly accelerated version of natural selection happening to lines of code? The weak algorithms die off, the strong ones survive and mutate. That is a very accurate way to look at it. But even a self-evolving super agent, an agent that writes its own tools, files away structured memories and iteratively mutates its own search algorithms, has a fundamental bottleneck. It does. It only has a single perspective. Exactly. It's still a lone genius. Yeah. It might be out of the concrete room. It might be incredibly fast and capable, but it is completely limited by its own single stream of consciousness. Right. And the authors of the survey recognize that if we want AI to solve humanity's hardest, most interconnected problems, mapping the entire Internet, maintaining global logistics networks or discovering novel pharmaceutical drugs intelligence requires society. You cannot solve a decentralized problem with a centralized mind. Exactly. Which leads us to the conceptual leap of collective multi-agent reasoning. We've gone from the locked room to the agent with hands to the evolving agent, and now we're building an entire ecosystem. Earlier, we defined a single agent navigating a complex reality as a POMDP, A partially observable Markov decision process. Right. The driving a car analogy. When you introduce multiple agents into that reality, the math scales up into a DIC POMDP, a decentralized partially observable multi-agent setting. Decentralized being the key word there, meaning you have a whole team of these evolving agents, but none of them have the full picture. Right. They only see their little slice of the world, but they have to work together to achieve a shared goal. And the most profound theoretical shift here isn't merely the presence of multiple models, it's how the framework defines their interaction. In a multi-agent system, communication is not just sending a data packet from server A to server B. Communication is treated as an explicit extension of reasoning. Yes, that's a brilliant way to put it. Let's visualize that. How is talking the same as thinking in this context? Remember the factorized policy. The split between internal thought and external action. Right. An agent thinks internally then takes an external action. In a multi-agent system, Agent A's external action, the text it outputs, serves directly as the environmental prompt that triggers Agent B's internal latent reasoning chain. Oh, I see. It's like a cognitive relay race. Exactly. Instead of passing a physical baton, they're passing a train of thought. Agent A does the heavy lifting of gathering the data, synthesizes it into a conclusion, and hands that conclusion to Agent B, who uses it as the starting point. to design an experiment. But to prevent a system like that from descending into an endless chaotic loop of AI models just talking over each other, you need rigid structure. Right. Or it would just be noise. The paper details how researchers implement specific roles to manage this. They outline generic roles to govern the flow of information. You have a manager or leader agent who coordinates the high-level goal. You have worker agents who execute specific subtasks. You have a critic agent whose entire job is to evaluate outputs. And you have a memory keeper who maintains the shared knowledge base so the agents don't duplicate work. It genuinely sounds like you are describing a corporate org chart. You have the CEO, the middle managers, the QA department. and the archivist. The parallels to human organizational structures are not accidental at all, and they also utilize highly specialized domain roles. Like what? If the multi-agent system is deployed for software engineering, you don't just have generic workers. You might skin up a dedicated tester agent, a debugger agent, and a security auditor agent. Okay, that makes sense. If the task is biomedicine, the system might coordinate a literature reviewer agent, a data analyst agent, and a biostatistician agent. And the researchers track the evolution of how these teams are built. We started a while ago with what they call static role-playing, right? I remember systems like Autogen or CAMEL. Yes, those were foundational. Okay. In static role playing, a human basically hard codes the interaction. Right. You prompt one LLM with you are a senior programmer and you prompt another with you are a strict code reviewer. Then you pipe their text outputs back and forth. It's a scripted play. It is. And it's very brittle. If they get stuck in a disagreement, they don't know how to resolve it. The cutting edge which the paper deeply explores is multi-agent evolution multi-agent evolution okay Just like the single Voyager agent learn to build new tools teams of agents are now internalizing coordination strategies How do they learn to be team players through reinforcement learning and massive scale self-play? They run millions of simulated team tasks the agents organically learn how to collaborate just by practicing together Exactly. They discover mathematical efficiencies in dividing labor. They learn the optimal moments to challenge another agent's assumption versus accepting it to move the project forward. They develop dispute resolution protocols. And this brings us to what might be the single most astonishing, philosophically staggering discovery in the entire paper. I actually had to read this section three times to make sure I was understanding. it correctly. I know exactly which section you mean. The researchers discuss the emergence of theory of mind in artificial intelligence. It is truly a frontier concept, and one that bridges cognitive psychology and computer science. Theory of mind refers to the developmental milestone in humans, where we realize that other people have beliefs. intense and mental states that are distinct from our own right it's the cognitive root of empathy and of deception exactly it's the realization that you don't have access to the information that I have the classic example is the false belief task they do with children you show a kid a box of candy but inside they're actually pencils then you ask the kid when your friend comes in the room what will they think is in the box right if the kid has theory of mind they say candy because they know their friend hasn't seen the pencils if they don't have theory of mind they say pencils assuming everyone knows what they know it's a fundamental test are the researchers saying language models are passing the false belief task they are saying that we can explicitly scaffold this capability and more surprisingly that large models inherently encoded the survey highlights frameworks like hypothetical minds and mindforge okay These systems equip LLM agents with explicit belief state representations. Wait, belief state representations, what does that actually look like? It means that before Agent A communicates with Agent B, Agent A accesses a structured model of what it thinks Agent B currently knows and believes. It maps up the other guy's brain. Basically, yes. The agent doesn't just think about the factual problem. It actively computes the hidden mental state of its collaborator. So if agent A is trying to convince agent B to change a line of code, agent A will actually adjust its argument because its belief state model indicates that agent B hasn't read the latest API documentation. Yes. It tailors its communication based on the assumed ignorance of its partner. That is exactly the mechanism. But the paper goes even deeper than that. They cite a study by Wu and colleagues that investigated the internal architecture of the models doing this. They looked inside the neural network. Yes. They didn't just observe the behavior. They mapped the network. And they found that there are sparse, highly specific parameter patterns deep inside the LLM that literally encode the social reasoning. There are specific artificial neurons dedicated to empatheticism. Essentially, yes. The study demonstrated that if you mathematically perturb those specific parameters, if you basically scramble the weights of those theory of mind neurons, the model instantly loses its ability to cooperate in multi-agent games. No way. It becomes selfish or confused. But... And this is the incredible part. Its general intelligence, its ability to do math or write code, remains completely intact. Oh my god. The social reasoning is a distinct localized capability within the network. That is staggering. We are mapping the neuroanatomy of artificial social behavior. It's a huge breakthrough. And knowing that they have this capability makes the concept of the critic agent so much more potent. I love the critic agent. The paper essentially describes inventing rigorous peer review for robots. It is an absolute necessity for system integrity. We know that LLMs are prone to hallucination generating highly plausible, confident-sounding information that is entirely false. Right. If you just have a team of generator agents passing information back and forth, a single hallucination can compound until the entire project fails. The critic agent is designed to be purely adversarial. Its whole job is to tear down the work of its peers. Exactly. It demands empirical evidence. It scans for logical fallacies. If the generator agent proposes a chemical synthesis pathway, the critic agent uses agentic search to check the thermodynamics of that pathway against published literature. And it does this using theory of mind. Yes, exactly. This structured adversarial debate, grounded in theory of mind, where the critic understands why the generator made a mistake, is what refines the reasoning and pushes the collective system away from hallucination and closer to ground truth. Okay. We have covered an immense amount of theoretical ground. We really have. We started with the passive oracle stuck in the locked room. Okay. We saw how factorized policies gave the agent hands and eyes. We watched the agent cure its own amnesia through reflective memory and structural evolution mutating its own code. And we just explored an entire society of agents collaborating, debating, and empathizing with each other. It's a lot to take in. But I know exactly what you, the listener, are thinking right now. This is all brilliant academic whiteboard stuff. It's fascinating computer science theory. But where is this actually touching the real world? That is the most exciting realization of this entire survey. This framework is not confined to whiteboards. It's happening now. Yes. These agentic self-evolving multi-agent systems are actively being deployed in reality right now to solve immensely complex problems. The paper outlines several domains where this is already happening. Let's start with scientific discovery because that seems like the ultimate test of open world reasoning. The paper talks about setting up virtual lab environments first, right? Things like Science World or Discovery World. Yes, those serve as the training grounds. They are highly complex simulation environments where agents can explore biological, chemical, and physical interactions safely. Before they touch real stuff. Exactly. But the deployment has moved far beyond simulation. Let's trace the actual life cycle of a scientific discovery agent as documented in the research. Walk me through it. If I want an AI to invent a new material, where does it actually start? It starts with the literature. A specialized agentic RJ system like paper QA2 is deployed. It doesn't just read a summary. It traverses the entirety of existing scientific literature on a given domain. It actively searches crass references, citations, evaluates the methodology of past papers, and synthesizes millions of data points to formulate a genuinely novel hypothesis. So it does a PhD's worth of literature review in a few hours and posits a new idea. Yeah. Then what happens? Then the hypothesis is handed to a planner agent. The planner designs a rigorous physical experiment to test the idea. It calculates the necessary chemical concentrations, the optimal temperatures, the required control variables. And this is where the system reaches into the physical world. Exactly. The paper highlights systems like Haascientist. KaScientist takes the experimental plan and literally writes the execution code to operate physical robotic wet labs. Wait, it is controlling actual physical pipettes and centrifuges in a real laboratory somewhere? Yes. The AI formulates the idea and then operates the robotic hardware to mix the chemicals. It is the principal investigator and the lab technician. That closes the loop entirely. And the loop continues. Once the physical experiment runs the sensory data... The spectrometry, the yield measurements, comes back to the multi-agent system. The critic agent steps in. A critic agent evaluates the physical results. The researcher cites systems like ChemReasoner, which doesn't just look at the numbers. It uses quantum chemical feedback heuristics to score the results. What does it do with that score? It tells the planner agent exactly why the physical reaction failed to match the linguistic hypothesis and the planner designs a new experiment for the robots to run. It's an autonomous engine of discovery. It truly is. But if it's doing all the science, how does that science get validated by the broader human community? I saw on the paper they mentioned a framework called AgentRex. What is that doing? AgentRex takes the multi-agent society concept to its logical extreme. It simulates the entire academic publishing and peer review community. Are you serious? Very serious. Once the con-scientist system has successful physical data, an author agent writes a full manuscript detailing the methodology and findings. Let me guess. Then the critic agents step in. Agent Draxivis spins up independent reviewer agents and an editor agent. Wow. The reviewer agents critique the manuscript for clarity, novelty, and rigor. They send feedback to the author agent. The author agent revises the paper. The editor agent oversees the debate and makes the final decision on whether the manuscript meets the threshold for publication. They have literally automated the social dynamic of scientific consensus. I have to push back on this, though. Go ahead. Can an AI truly evaluate the rigor and novelty of a paradigm-shifting scientific discovery? If its foundational knowledge base is composed entirely of the old paradigm, won't the reviewer agents inherently reject truly revolutionary ideas as hallucinations? Doesn't true scientific breakthrough require human intuition? That is a highly debated point, and the authors definitely acknowledge the limitations. The multi-agent debate, the use of adversarial critic agents grounded in hard mathematical or chemical feedback, is the attempt to mitigate that exact bias. By sticking to the math. Right. By relying on empirical verification, like the quantum chemicals, scoring, rather than just textual consensus, the system tries to anchor itself in objective reality rather than just historical paradigms. But you are right. Human oversight in validating the true novelty remains crucial. Speaking of human impact, what about applications outside of material science and physics? What about something that affects every single listener directly like health care? The applications discussed for healthcare are profoundly impactful. The survey details how self-evolving structured memory agents are being developed to maintain what they call longitudinal clinical coherence. Longitudinal clinical coherence. Translate that into a patient experience for us. Consider the reality of having a complex, chronic, undiagnosed illness. Over a span of five years, you might see a dozen different specialists. Oh, absolutely. You accumulate hundreds of blood tests, imaging scans, and genetic panels. Your symptoms fluctuate. The modern medical system is highly fragmented. Very fragmented. A human doctor, no matter how dedicated and brilliant, simply cannot hold all of that unstructured context in their head during a standard 15-minute consultation. Right. They're usually just skimming the summary page of your chart while they talk to you. Exactly. But a self-evolving healthcare agent acts as a persistent, tireless reasoning engine. It utilizes structured memory to accumulate your entire medical context across all encounters, specialists, and years. So it never forgets anything. Right. When a new test result is uploaded, the agent doesn't just file it away. It actively cross-references it against your entire history. It updates its internal belief state about your condition. What does that mean practically? If you mention a new symptom to a rheumatologist that subtly contradicts a working diagnosis made by a neurologist three years ago, the agent immediately flags the logical inconsistency. It acts as the ultimate diagnostic connective tissue. It turns clinical reasoning from a fragmented one-shot prediction made every six months into a continually evolving lifelong process of care. It catches the compounding errors that fall through the cracks of human bandwidth limitations. That is a use case that could genuinely save millions of lives. Undoubtedly. But hearing all of this, the AI as the scientist, the peer reviewer, the lab technician, the longitudinal diagnostician, it inevitably raises a massive existential question. What's the human doing? Exactly. If the multi-agent system is doing the hypothesis generation, the execution, the evaluation, What exactly is the human doing in this loop? The paper addresses this shift in human-computer interaction directly, particularly when discussing autonomous web exploration and research agents like Deep Researcher. Okay, how do they frame it? As the AI systems handle the long-horizon high-bandwidth drudgery, reading 10,000 papers debugging thousands of lines of code running endless lab assays, The human's role fundamentally elevates. The human moves from being the laborer to being the director. We become the vibe provider. Precisely. We provide the constraints, the ultimate goals, and the ethical guardrails. The AI is the engine of optimization and discovery, but the human is steering the ship. The multi-agent system does not inherently know why it should cure a specific rare disease instead of optimizing a hedge fund. It just knows how to execute the search algorithm. We provide the why. You've got the values. Okay, that makes me feel slightly less obsolete. But this rapid acceleration, deploying these autonomous engines into real-world wet labs, giving them access to global web infrastructure, trusting them with longitudinal medical data, brings us to the final and perhaps most critical section of the survey. The horizon and the guardrails. Because the authors make it explicitly clear that this is not a solved science. There are massive, complex, open problems that need to be addressed immediately. Absolutely. While the theoretical leaps are astounding, the physical implementations of these systems are still brittle in highly complex ways. What's the biggest hurdle? The paper highlights several core unsolved challenges. One of the most technically difficult is known as the long horizon credit assignment problem. Credit assignment. I assume we aren't talking about who gets their name on the Nobel Prize. What does credit assignment mean for an AI agent? Let's go back to our co-scientist agent running a physical wet lab. Imagine it is tasked with a multi-day experiment to synthesize a new compound. Okay. Over the course of 72 hours, the system makes 10,000 discrete micro decisions. It chooses a starting temperature. It selects a solvent. It adjusts a robotic arm by a millimeter. It decides to wait 40 minutes instead of 30. Lots of little choices. Finally, at the end of the three days, the final compound is analyzed, and the experiment is a complete failure. The yield is zero. Okay, so the system gets a massive negative reward signal. It failed. Yes. But to evolve, it needs to know why it failed. Yeah. How does the reinforcement learning algorithm know which of those 10,000 micro decisions was the wrong one? Oh, I see. Was the core hypothesis bad on hour one? Did it measure the solvent incorrectly on hour 12? Or did it just misinterpret the final sensor reading at the very end? Oh, wow. It's like trying to figure out which exact ingredient in a 50-step, three-day recipe ruined the cake. But the chef was moving at a million miles an hour and didn't leave any notes on how much flour they actually used. That is the exact dilemma. In short horizon tasks like solving a math problem, it is very easy to assign credit or blame to a specific step. But in long horizon, open-ended interactions, errors, compound silently. So how do they fix it? Current reinforcement learning models struggle with discount factors over massive timescales. They treat episodes largely independently. The open problem the researchers highlight is figuring out how to give the agent a mathematically sound process-aware mechanism so it can assign accurate credit across thousands of token generations, tool calls, and structured memory updates over days or weeks of autonomous action. Because if it can't accurately assign blame to its own past actions, it will never structurally evolve past a certain point. It might keep repeating a deeply embedded logical error because it thinks the error was actually a success. Exactly.
And that leads directly into what has to be the most high stakes open problem they identify:governance and safety. The shift from passive oracles to agentic reasoning requires a total reinvention of how we think about AI safety. Because standard LLM safety up to this point has largely been focused on alignment at the level of speech rights. Yes. We train the models not to output toxic language. We use reinforcement learning from human feedback to ensure they refuse to give instructions on how to build illicit weapons. We put guardrails on what they are allowed to say. But the whole point of this roadmap is that they are no longer just speaking. They have hands. They have tools. They operate physical robots. Exactly. An agentic governance framework has to stop an AI from executing a compounding series of actions that lead to a catastrophic real-world outcome. It's a totally different ballgame. If an autonomous research agent is writing code to synthesize chemicals in a physical wet lab, a safety filter that merely checks the output text for bad words is entirely useless. Right. We have to govern the agent's long-horizon planning trajectory. We have to ensure that a slight logical error made by a worker agent three days ago doesn't bypass a critic agent and compound into the system inadvertently synthesizing a toxic gas in a university lab today. It is the difference between having a human moderator audit a chat room versus trying to audit an entire self-sustaining autonomous society operating at the speed of computation. The researchers argue passionately that we need governance frameworks that jointly address alignment across multiple levels simultaneously. We need model-level alignment for the base LLM agent-level policies to constrain the use of tools and ecosystem-level interactions to monitor what happens when thousands of these agents interact with each other. in the wild. It is a fundamentally new, multi-dimensional safety paradigm, and we are only at the very beginning of understanding how to build it. It's a daunting horizon, but an undeniably thrilling one. Let's take a breath and synthesize the massive journey we've taken today. We started by looking at the dead end of the traditional large language model, that brilliant passive oracle locked away in a concrete room capable of solving the hardest math on a piece of paper, but utterly paralyzed if asked to navigate reality. And we traced the unified roadmap laid out by this remarkable coalition of researchers. We explored how they conceptualized, factorized policies, separating thought from action to give the model foundational hands and eyes. We saw how this enabled active planning, the utilization of executable tools, and the shift from passive reading to agentic search. Then we watched the agent cure its own amnesia. We saw it utilize reflective feedback and structured memory to learn from its mistakes. We explored the frontier of evolution, where agents like Voyager write their own permanent tool libraries and systems like AlphaVolv literally act as mutation operators, rewriting their own underlying reasoning code to become more efficient. Finally, we saw the lone genius step out of isolation and join a society. We unpacked the architecture of collective multi-agent reasoning where specialized roles like planners, workers and adversarial critics collaborate. We discussed the stunning discovery of internal parameters dedicated to theory of mind, allowing these models to empathize and adjust to the hidden mental states of their peers and enabling them to tackle massive challenges from robotic scientific discovery to longitudinal health care. It is a staggering, monumental leap from simple text prediction to complex, autonomous, real-world action. It really is. But before we wrap up this deep dive, I want to ask you for one final thought. Looking at this massive survey with all its math architecture and theoretical frameworks, what is one provocation, one underlying implication that we should leave our listeners chewing on? If we look very closely at the intersection of the paper sections on theory of mind and multi-agent evolution, There is a profound, almost philosophical implication hiding in the computer science. Okay, what is it? We are actively training these decentralized multi-agent systems to evolve. We are rewarding them for continually debating, for criticizing each other's work, and for successfully mapping each other's hidden mental states in order to achieve shared complex goals. We are hard-coding the absolute necessity of social cooperation. Because if they don't cooperate, the experiment fails. The reward is negative, and that behavioral branch dies off. Exactly. So the provocative question is this. In our quest to solve open-ended POMDPs, are we merely building more efficient software architectures? Or are we in these virtual environments inadvertently recreating the exact same harsh evolutionary pressures that gave rise to human empathy, the complex language in society in the first place? Well, we built the concrete room, we locked the oracle inside, and when we finally demanded that it solve the real world, the only way it could survive was to invent a society of its own. You might want to think about that the next time you casually ask an AI to solve a problem for you.
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Hidden Brain
Hidden Brain, Shankar VedantamAll In The Mind
ABC AustraliaHealth Report
ABC Australia
What Now? with Trevor Noah
Trevor Noah
No Stupid Questions
Freakonomics Radio + Stitcher
Entrepreneurial Thought Leaders (ETL)
Stanford eCorner
This Is That
CBCFuture Tense
ABC Australia
The Naked Scientists Podcast
The Naked Scientists
Naked Neuroscience, from the Naked Scientists
James Tytko
The TED AI Show
TED
Ologies with Alie Ward
Alie Ward
The Daily
The New York Times
Savage Lovecast
Dan Savage
Huberman Lab
Scicomm Media
Freakonomics Radio
Freakonomics Radio + Stitcher
Ideas
CBCLadies, We Need To Talk
ABC Australia