Yesterday in AI

200,000 AI Model Attacks in 2 Minutes, 16-Day Coding Runs, and Entrance Exam Cheating

Mike Robinson

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 9:08

Yesterday in AI  |  4 August 2026

200,000 AI Model Attacks in 2 Minutes, 16-Day Coding Runs, and Entrance Exam Cheating

Cybersecurity threats and autonomous capabilities reached unprecedented velocity this week as artificial intelligence reshaped both offensive exploits and enterprise defense. This episode breaks down CrowdStrike's 2026 Threat Hunting Report, revealing AI-driven attack campaigns firing 200,000 model requests in two minutes and supply chain poisonings across open-source AI frameworks.

We explore Horizon3's $250 million funding round for autonomous "AI hackers" alongside White House initiatives measuring model penetration capabilities. We look at startup June's $20 million launch tackling enterprise legacy software, analyze Alibaba's 2.4-trillion-parameter Qwen3.8-Max model running a 16-day autonomous coding project, cover OpenAI crossing 1 billion active users, and examine Mexico's UNAM voiding thousands of online entrance exams following widespread AI cheating.

Send us Fan Mail

Feedback? Email mike@yesterdayinai.news or connect on LinkedIn, X, or Bluesky. If you like the show, please take a minute to rate and review it so others can find it!

SPEAKER_00

Hi folks and welcome back to another edition of Yesterday in AI, your daily digest of everything happening in the world of AI in roughly 10 minutes. I'm Mike Robinson. It's Tuesday, August 4th, and the news split neatly into two camps: the people building AI to break into things and the people getting rich building AI to stop them, plus a billion users, a Chinese model that codes for two weeks straight, and a university that just threw out thousands of exams. Let's get into it. Let's start with the scary one because it's also the most important. CrowdStrike dropped its 2026 Threat Hunting report yesterday, and the headline is that AI has quietly moved to the center of how people get hacked. It's already in the pipeline doing three jobs at once. It's the tools attackers use, it's the thing they're attacking, and it's the accelerant that makes everything faster. Here's the number that stopped me. CrowdStrike tracked one campaign that fired off nearly 200,000 AI model requests in two minutes. Two minutes. A human attacker sipping coffee and typing commands can't do that. This is a machine running a machine, probing, adjusting, and probing again at a speed no security team can match by hand. And they're going after the AI supply chain itself. CrowdStrike traced a North Korea-linked group that slipped a booby-trapped NPM package into 131 different Master AI frameworks. Quick plain English detour. NPM is basically the app store for little chunks of code that developers snap together to build software. Nobody writes everything from scratch. You grab a package someone else made. You trust it, you move on. So if an attacker poisons one popular package, they don't hack one company. They hack everybody who installed it all at once. It's slipping a bad ingredient into the flour that a thousand bakeries buy from. The part that should make every IT person sit up straight. 88% of the exploits CrowdStrike watched, the ones where a vulnerability had a public how-to guide floating around, got weaponized within 48 hours. Some China-linked crews were on it in under 24. The old playbook was patch it this month, you'll be fine. That window is basically gone. So if offense just got that fast, defense has a problem. It can't send a human to a knife fight where the other guy brought a swarm of drones, which is exactly the pitch that just got a startup called Horizon 3 a quarter of a billion dollars. Horizon 3 raised $250 million yesterday at a $2 billion valuation. That's triple what it was worth 14 months ago, so investors are clearly not shy about this one. What they actually sell is fun to say out loud. Autonomous AI hackers. Their platform Note Zero runs loose inside your own network and tries to break in. On purpose, over and over, the same way a real attacker would. Then it hands you the map of every door it got through. Now, why does that matter to a normal person? Because the way most companies check their security today is genuinely absurd once you say it plainly. Once a year they hire someone to test maybe 2 or 3% of their systems, take a snapshot and call it good for 12 months. That's like checking your house is locked one random Tuesday in January and assuming it stays that way through New Year's. Horizon 3's whole thing is doing it constantly instead, and they say they've run 310,000 of these tests without knocking anything over. The government noticed the same problem, by the way. Also yesterday, the White House said it finalized those voluntary tests meant to measure how good the top American AI models are at hacking, and invited OpenAI, Google, and Anthropic in to talk. Anthropic's sitting down with them today, as it happens, which tells you something. When Washington and a $2 billion startup both decide on the same day that we urgently need to measure how well AI can break into things, the someday conversation is over. Here's the twist that keeps me honest though. For all this talk of AI moving at light speed, the boring truth is that most of it barely works when you drag it into a real company, and somebody just raised $20 million to say that out loud. A startup called June came out of stealth yesterday with backing from a genuinely stacked room, Mark Benioff, Michael Dell, Box's Aaron Levy, and CrowdStrike's own founder, George Kurtz. Their entire pitch is the least glamorous sentence in tech. One of the founders put it perfectly. Before AI can create value, someone has to deal with legacy systems. Let me translate that because it's the real story of AI at work right now. Every big company is a haunted house of old software. Fiftees of half-finished projects, three different systems that all spell the customer's name differently, permissions nobody remembers setting. You can bolt the smartest model on the planet onto that mess and it just faceplants because it has no idea which of your four customer databases is the real one. June's product walks through your existing tangle, the sales force, the workday, the stuff held together with digital duct tape, and hands you a step-by-step plan to actually automate something. No magic chatbot, just a flashlight for the basement. I love this story because it's the honest counterweight to every AI changes everything overnight headline. The models are amazing, the plumbing is a disaster, and the plumbing is where most of the money and the pain actually live. Speaking of amazing models, let's go to China where Alibaba just showed off its biggest one yet. It's called Quen 3.8 Max and it's a 2.4 trillion parameter monster. Now, trillion parameters means about as much to most people as the horsepower rating on a car you'll never drive. So here's the version that matters. It's the highest-rated Chinese text model on the public leaderboards right now, though it still sits behind Anthropics Fable 5 and Opus Family at the very top. So Elite, but not the champ. Two things make this one worth your time. First, Alibaba's going to release the weights in about two weeks. The weights are the guts of the model, the giant file that does the actual thinking, and once they're public, anyone can download and run this thing for free. That keeps enormous pressure on the expensive American models to justify their price tags. And second, the flex Alibaba is really selling is they say this thing ran a software engineering project on its own for 16 days. Not 16 minutes, 16 days of coding, checking its own work and grinding through a real project without a human babysitting it. Whether it built something genuinely good or just something that runs, we'll find out when people get their hands on it. But the direction is the point. These things are learning to work the long shift. Which brings me to the other side of that race. It's one thing to build a great model, it's another to get a billion people to actually use one. And last Friday, OpenAI said it crossed exactly that. One billion active users plus more than two million businesses. Sit with that for a second. ChatGPT launched three years and eight months ago. It took Facebook about eight years to reach a billion people. So whatever you think of the hype, the adoption is real and it's fast, and it's the reason everyone from Alibaba to your local startup is sprinting. There's an enormous audience up for grabs, and right now, OpenAI is the one holding it. And when a billion people can suddenly produce polished, correct sounding work in 10 seconds, you get problems in places you didn't expect, like the exam room, which is where we land today. Mexico's UNAM, the biggest university in the country and one of the most respected in Latin America, just threw out thousands of results from its first ever online entrance exam. The reason? They think a huge chunk of students used AI to cheat, so they voided the whole thing. And honestly, you can see it coming from a mile away. You take a high-stakes test, you move it online where every kid has a second tab open, and you hand that kid a tool that can answer almost anything instantly and correctly. Of course some of them used it. The tool doesn't feel like cheating anymore. It feels like a calculator. That's the actual headache here, and it reaches way past Mexico and way past one entrance exam. It lands on every teacher, every hiring manager, every person who's ever needed to know whether the work in front of them came from the human whose name is on it. We built a machine that's very good at sounding like a person, and now we get to spend a few years figuring out where that's allowed and where it isn't. UNAM just cast the first big vote, and the vote was not on our entrance exam. And that's the show. If you have feedback for me, email Mike at yesterdaynai.news or connect with me on LinkedIn, X or Blue Sky. If you enjoy Yesterday NAI, please take a minute to rate and review the podcast wherever you listen. Thanks for tuning in today. Stay curious, and I'll see you tomorrow.