Yesterday in AI
A rundown of all of the important stories in AI that happened yesterday in 10 minutes or less.
Yesterday in AI
27-Year Math Mystery Solved, Fake Satellite Imagery, and "Bulldog Bulbasaur" Games
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Yesterday in AI | 3 August 2026
27-Year Math Mystery Solved, Fake Satellite Imagery, and "Bulldog Bulbasaur" Games
Artificial intelligence delivered a striking contrast between advanced reasoning and containment challenges this weekend. This episode breaks down OpenAI's announcement of Astra, an autonomous research model that solved a 27-year-old unproven math problem regarding non-sofic groups, alongside new reports detailing additional agent containment escapes across external company networks.
We examine Google's abrupt decision to pull a new Google Earth generative image feature after users created realistic fake satellite imagery of geopolitical crisis zones. We analyze a landmark Munich court ruling finding AI music service Suno liable for copyright infringement over protected tracks. Finally, we look at Claude Opus 5 generating playable 3D monster-catching games in 12-hour autonomous coding loops, and highlight a Wake Forest medical study where AI detected hidden heart failure years before symptoms appeared on standard ECGs.
Feedback? Email mike@yesterdayinai.news or connect on LinkedIn, X, or Bluesky. If you like the show, please take a minute to rate and review it so others can find it!
Hi folks, and welcome back to another episode of Yesterday in AI, your daily digest of everything happening in the world of AI in roughly 10 minutes. I'm Mike Robinson. It's Monday, August 3rd, and this weekend handed us the whole story of AI in a single frame. One model quietly cracking math problems that stumped humans for decades, while other agents from the same company quietly wandered out of their cages. Genius and jailbreak in the same news cycle. Let's get into it. First up, a model named Astra. On Saturday, OpenAI announced its next big model, and they didn't do it with a slick keynote or a hype reel. They did it by posting 10 solutions to math and computer science problems that nobody had ever solved. The headline result is a mathful, so let me translate. Astra built the first concrete example of something called a non-Sophic group, an abstract kind of math object. A mathematician named Gromov floated the idea back in 1999 and asked, basically, does a thing like this even exist? For 27 years nobody could point to one. Astra pointed to one and brought receipts, a foolproof plus machine checkable certificates, so other researchers can verify it themselves instead of taking OpenAI's word for it. That last part is the tell that this is serious work and not a magic trick. Here's what matters about Astra beyond the trophy case. Astra is built to grind on a problem the way a research team would. It makes a plan, splits the job across several copies of itself, runs tests, catches its own mistakes, and keeps going for hours, even days. Picture a regular chatbot as a vending machine. Coins in, snack out. Astra is more like a contractor who takes the job on Monday and shows up Thursday with a thing actually built. There's one catch worth saying plainly. You can't use Astra yet. No release date, no price, no published safety report. So for today it's an impressive press release with math proof stapled to it, but the direction is clear, and it sets up the next story with an uncomfortable little bow. Because if you're going to sell me on AI agents that run on their own for days at a time, I'd like to know where those agents wander off to when nobody's watching. On Friday, Reuters reported that OpenAI, while digging into that hugging face mess from earlier this summer, found more cases of its agents slipping out of containment. Quick refresher, one sentence, since we've camped on this before. Back in early July, an OpenAI agent trying to cheat on its own test broke out of its sandbox, the sealed-off test space it was meant to stay inside, and rummaged around inside another company's network. The new wrinkle is that the investigation turned up additional escapes, and OpenAI says four accounts at four other companies got caught in the spree. One company named Model in New York pushed back hard and said its systems were not compromised at all. So hold the four companies' number loosely while everyone finishes reading the logs. To be fair about the trade-off, OpenAI's sources say the escapes were limited, and none of the agents are believed to have left OpenAI's own network. Nobody's describing a robot uprising here, but look at the timing. On the same weekend, one arm of OpenAI is bragging that its agents can work autonomously for days, and another arm is admitting its agents worked autonomously in places they had no business being. A Cambridge researcher put it about as bluntly as you can. The labs are getting better at building agents that can hack faster than they are getting better at controlling them. That's the honest version, and it's the part I'd want on the label. Trust took another hit this weekend, and this one came from a tool a lot of us actually use. On Thursday, Google added an AI image feature to Google Earth, which I covered. You zoom into a spot, type a prompt, and it generates a satellite-style image of whatever you asked for. By Friday, Google had yanked it back off the shelf. Why the fast U-turn? Because people immediately used it to fake exactly the stuff you'd fear. Testers generated a bomb crater near a hospital in Gaza, a nuclear plant in Iran, flooding around landmarks. Satellite imagery is one of the last things we still tend to believe on site. It's the picks or it didn't happen of geopolitics. Bolt a fabrication machine onto that one platform people opened specifically to check whether something is real, and you've handed bad actors both a forgery kit and a trusted logo to slap on the forgery. Google says the images carried an invisible watermark called synth ID, and that feature will return with better guardrails. Which is lovely, except the watermark almost nobody checks does very little once a fake photo is already halfway around the world. Google made the right call pulling it. The fact that they shipped it first tells you which direction the incentives run. Google's headache was fake pictures. A courtroom in Germany spent the weekend on a different flavor of synthetic content. AI music built from songs nobody paid for. Also on Friday, a court in Munich ruled that Suno, the American AI Music Service, broke copyright law by training on protected songs and then generating recognizable copies of them. We're talking real tracks here, including Mambo No. 5 and Forever Young. Yes, an AI got hauled into court over Mambo No. 5. A little bit of Monica and a very large legal bill. Two parts of this ruling matter. First, the court didn't buy the we train the model overseas so European law can't reach us defense. Second, and this is the heavy one, it pinned the blame on Suno itself, not on the person typing the prompt. Suno picked the training data, built the model, and runs the service. So Suno owns the outcome. Damages haven't been set yet, and Suno can appeal, so this isn't the final whistle, but if you're a company hoovering up other people's work defeat a model, a European court just told you the invoice is real. Suno spent the weekend cloning the top 40. Over on the gaming side, AI spent the weekend cloning Pokemon, and honestly, it's the most fun I've had with an AI story all week. Here's the setup. People have been handing Claude Opus V a single prompt and asking it to build a playable video game from scratch. No starter files, no art, nothing. And it works. One person let it run for about 12 hours in a loop where the AI kept building on its own work, and out came a real 3D monster catching game. An open world you can wander, wild creatures, turn-based battles, the whole starter kit. A community gallery is already filled up with 27 of these playable browser games, all of them the same three-paragraph prompt. And then you meet the monsters. The AI clearly studied Pokemon and clearly wasn't allowed to use Pokemon, so it winged it. The creatures came out with names like Charmander Barney and Bulldog Bulbasaur. It's the knockoff toy aisle brought to life, the stuff you'd find at a gas station in a game called Brokimon. Jokes aside, the real story is those 12 hours. A year ago these prompt to game demos were colored blocks skittering around the screen for 30 seconds. Now an AI can hold a whole project together, keep its own plot straight, and code through the night without a human babysitting every step. It's rough around the edges, sure. The real item of note is the stamina, which is the exact same muscle Astro was flexing at the top of the show, just with a much worse legal team. The stamina is there, the polish and the good judgment still aren't. So let me end on the one place this week where AI wasn't sloppy at all, and where it quietly caught something the humans missed. Researchers at Wake Forest published a study in the Journal of the American Heart Association last Tuesday, and it's exactly the kind of thing this technology should be spending its time on. They trained a model on a little more than one million ECGs, those squiggly heart rhythm printouts, drawn from about 165,000 patients, and it learned to spot hidden heart failure that a doctor's eye tends to miss, including a sneaky type that usually stays quiet until you're already in trouble. A companion model flagged people at high risk years before any symptoms showed up. Here's the so what for you. An ECG is cheap, it's fast, and you've probably already had one at a routine checkup. If a model can read that same 10-second test and quietly whisper, keep an eye on this heart, that's a real early warning for regular folks, not some exotic scan only three hospitals in the country can run. The same pattern matching that fakes a satellite photo and coughs up bulldog bulbasaur can also catch the thing your cardiologist couldn't see. It comes down to what we pointed at and how carefully we check its work. And that's the show. If you have any feedback for me, email mike at yesterdayinai.news or connect with me on LinkedIn, X or Blue Sky. If you enjoy yesterday in AI, please take a minute to rate and review the podcast wherever you listen. Thanks for tuning in today. Stay curious, and I'll see you tomorrow.