LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.
If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Episodes
941 episodes
"Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout
TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeove...
"AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen
OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It...
"How My Students Think About AI" by dvd
Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spri...
"You’re Absolutely Right" by Linch
Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these log...
"LLMs Are Starting To Noticeably Accelerate Our Work" by johnswentworth
About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean. The first to la...
"There Will Come Soft Rains" by tanagrabeast
Today is August 4, 2026 [Crossposted from AI StopWatch] In the living room the voice-clock sang, Tick-tock, seven o’clock, time to get up, time to get up, seven o’clock! as if it were afraid that nobody would....
"Four LLM loss functions → four flavors of LLM misalignment" by Steven Byrnes
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately. Training stage Loss function
"FAQ: Isn’t AGI coming too soon for reprogenetics to help?" by TsviBT
Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine,...
"What just happened? A retrospective of AI alignment" by Richard_Ngo
This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progre...
"Don’t Build Mindreading" by Celer
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building datasets to enable telepathy, to use their term.
"OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi
How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a...
"models may behave differently in graded episodes (a tirade)" by nostalgebraist
Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure,...
"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh
This post is written in our personal capacity. Three Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a...
"Generalized atheism rules out “inaccurate simulation”-ism." by Eliezer Yudkowsky
Reposted from Facebook, on January 17, 2017. I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis. Yes, I know it's a joke. I'm still concerned. ...
"Arguments for P" by Cleo Nardo
Daniel Kokotajlo: To be clear, we don’t claim P will happen specifically. But when we wrote out our best-guess scenario month by month, P kept happening. Eventually we decided to just publish P. I’m at ~80% on P; my coauthors are lower. R...
"RL & search is a terrifying way to build AGI (an FAQ)" by Steven Byrnes
Q1: What are you saying? A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that's choosing actions via reinforcement learning (RL) and/or model-based search and planning—a ...
"Returning to ARC" by paulfchristiano
I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building techniques to find mechanistic explanations for neural network behavior and t...
"Thousand-dimensional structure" by Geoffrey Irving, David Africa
Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence...
"Big-World Intuitions" by sarahconstantin
Consider the following situations: when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything adversarial “against” them, but just to focus on gro...
"Duane Arnold" by Tomás B.
“So maybe I should enlighten you on what happens in your absence. This selfish existence where this introvert turns extrovert and dons her social armour.” Some posh girl in drainpipes said that - 200 views on TikTok and me one of them. But she di...
"The High-Control Dynamics at MAPLE" by Kyle Hubbard
As I write, many former friends of mine are living and working at a monastery in Vermont that I believe is a high-control group, commonly known as a ‘cult’. I say this not as someone who was concerned to see these friends go there, but someone wh...
"The Long (Self-)Correction" by Wei Dai
I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection. Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper...
"You (Yes, You) Need A February 2020 Checklist for AI Policy" by davekasten
TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and when it does, in a detailed way. (Epistemic status: origin...
"Is Mythos good at cyber because it kept hacking Anthropic during training?" by Tim Hua
From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-bas...