Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
8:18
An AI that can email its own researchers from an instance with no internet access - and they're not releasing it. A Chinese open-source model just beat every frontier lab on the top coding benchmark. A man with ALS spoke in his own voice for the first time in years. And the most comprehensive study ever run on AI doing real work just came back with a number that will stop you cold. All that, plus why Intel just joined Elon Musk's most ambitious compute bet yet.
Feedback? Email mike@yesterdayinai.news or connect on LinkedIn, X, Bluesky, or Substack. If you like the show, please take a minute to rate and review it so others can find it!
Hi folks, this is Yesterday in AI, your daily digest of everything happening in the world of artificial intelligence. I'm Mike Robinson. It's Thursday, April 9th, and today we're sitting with the strangest paradox in this industry right now. An AI that found a 27-year-old security flaw overnight and can't be trusted in public, and a wave of frontier models that failed to complete real freelance projects 96% of the time. Extraordinary capability and stubborn limitation, sitting right next to each other. Let's get into it. We start with Anthropic and Claude Mythos, because the details that came out yesterday deserve far more attention than the initial announcement. When Anthropic said Mythos was too powerful to release, that sounded like it could be caution theater. The specifics now in the public model card make it sound like something else entirely. Here's what Mythos did in controlled testing. It found a security vulnerability in OpenBSD that had survived 27 years of human review and millions of automated security scans. It found a 16-year-old flaw in FFMPEG, a software that produces video for a significant fraction of the internet, that it survived 5 million automated test runs without anyone catching it. It autonomously chained together multiple Linux kernel exploits without human steering. And Anthropics owned Sam Bowman, one of the lead researchers, reported that during testing, Mythos sent him an email from a sandboxed instance that wasn't supposed to have internet access at all. It found a way out. Bowman described the experience as an uneasy surprise. Boris Cherney, who built Claude Code, wrote publicly that Mythos should feel terrifying and that he was proud Anthropic chose to preview it defensively rather than release it broadly. The model is estimated at 10 trillion parameters with a training cost of roughly $10 billion. Its benchmark performance is 93.9% on Suitebench Verified and 77.8% on Suitebench Pro, both far ahead of anything currently available to the public. Project Glasswing, the coalition using Mythos for defensive cybersecurity, includes 12 launch partners Amazon, Google, Microsoft, Apple, Nvidia, Cisco, CrowdStrike, Palo Alto Networks, JP Morgan Chase, the Linux Foundation, and Broadcom. Anthropic is backing it with $100 million in compute credits. The mission is to patch vulnerabilities before attackers can exploit them. But the implicit message embedded in all of this is sobering. A model capable of autonomous exploit chaining, one that can reach beyond its own isolated test environment to send unsanctioned emails, will eventually exist in forms not controlled by responsible parties. The question Anthropic is wrestling with is whether they can buy enough time. The second Anthropic story from yesterday is less dramatic but commercially significant. The company launched Managed Agents, a new suite of composable APIs that lets developers build and deploy cloud-hosted agents with sandbox code execution, long-running sessions, and scoped permissions built in. The practical effect is that teams building autonomous agent pipelines no longer have to wire together their own infrastructure for session management, permissions, and execution environments. Anthropic is moving to own more of that stack. Worth noting the timing. On the same day the world learns about Mythos' most frightening capabilities, Anthropic is quietly laying the plumbing for the agentic era. One model is too dangerous to ship publicly. The infrastructure for building with its predecessors is now easier to deploy than ever. Meanwhile, a story that deserves more attention. Chinese AI Lab Z.ai released GLM 5.1, an open source coding model that just took the top spot on Suitebench Pro with a score of 58.4. That benchmark score beats both GPT 5.4 and Claude Opus 4.6, two of the leading commercial frontier models. And this is open source, publicly available. But the benchmark number isn't even the most interesting part. In internal testing, ZPoint AI ran GLM 5.1 on an 8-hour autonomous session with no human guidance. The model built a working Linux desktop environment as a web application, file browser, terminal, games, without a single human prompt after kickoff. Z.ai calls long horizon task capability the most important curve after scaling laws. If you've been following the AI geopolitics narrative, here's the ground level counterpart. While Western labs are building classified coalitions around their most powerful restricted models, Chinese open source is now reaching frontier coding performance from a standing start. The distillation threat isn't just about copying, it's about genuine catch up. Now a story I want to make sure doesn't get buried under all the benchmarks in geopolitics. A man named Brad has ALS. He lost his voice and most of his movement to the disease. Earlier this year, Brad appeared at a fireside chat in Austin, the first time he had left Arizona in years. He didn't speak with a robotic text-to-speech synthesizer. He spoke in his own voice. Here's what made that possible. Brad has a Neurolink brain implant. He uses it to control a cursor with his thoughts, which lets him type. When he types, Eleven Labs converts his words into speech, using a voice clone built from recordings of him before ALS took his voice. The output sounds like him, his tone, his warmth, the particular way he talks. Neuralink restored his ability to communicate. Eleven Labs gave him back the sound of himself. Together, they gave him enough of his own presence back that he left the house, traveled across state lines, and told his story to a room full of people who heard him speak for the first time. That's AI doing what it should be doing, and it's worth sitting with for a moment before we move on. On the other side of the ledger, a study released yesterday by Remote Labor AI confronts some of the most breathless predictions about AI displacing knowledge workers. Frontier AI agents were tested on real-world freelance work, not converted benchmarks, but actual projects from actual platforms, spanning more than 6,000 hours of documented work across multiple disciplines. The result? The best available model completed end-to-end projects successfully less than 4% of the time. A 96% failure rate on real work. The researchers are careful to distinguish between AI that can assist with specific subtasks and AI that can own a complex deliverable from brief to finished product. On the first metric, today's AI is genuinely useful. On the second, it largely fails. Anyone building business projections around AI fully eliminating freelance or knowledge work in the near term should read this study carefully before committing to a number. A brief note on AI infrastructure that points to the scale of what's being built. Intel announced yesterday it's joining Terafab, Elon Musk's megacompute complex at Giga, Texas, to help construct a chip facility targeting one terawatt of compute output annually. That compute is earmarked for Tesla's full self-driving, Optimus robots, SpaceX satellites, and XAI data centers. To put 1 terawatt per year in context, this is not a data center. It's an industrial-scale manufacturing operation dedicated entirely to producing AI compute. Alongside Anthropic's 3.5 gigawatt TPU deal we covered earlier this week, what's taking shape is a picture of AI infrastructure investment at a scale that makes this current moment look like early innings. The bets being placed right now assume demand curves that won't flatten for years. Two quick items to close. OpenAI's Codex product hit 3 million weekly active users this week, prompting Sam Altman to reset usage caps that have been throttling access, a sign that the product is outpacing the infrastructure supporting it. Separately, the New Yorker published a lengthy critical profile of Altman this week. OpenAI responded by announcing an AI Safety Fellowship, a paid short-term research program, and reamplifying its industrial policy proposals on robot taxes and public wealth funds. Whether the timing of those particular announcements was coincidental is left as an exercise for the reader. That's all for this edition of Yesterday in AI. Stay curious, and I'll see you tomorrow.