Yesterday in AI
A rundown of all of the important stories in AI that happened yesterday in 10 minutes or less.
Yesterday in AI
The Speed Over Smart Shift, Unsupervised Agent Turf Wars, and Peeking at your Mac History
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Yesterday in AI | 15 August 2026
The Speed Over Smart Shift, Unsupervised Agent Turf Wars, and Peeking at your Mac History
The artificial intelligence race pivoted from raw intelligence to blistering inference speed, enterprise affordability, and local device integration this week. This episode breaks down OpenAI previewing "Ultrafast" powered by Cerebras wafer-scale silicon, running GPT-5.6 Sol at 14x speed and 750 words per second.
We examine Ramp's monthly spending report revealing why enterprise buyers use Anthropic's flagship Fable 5 just 6% of the time in favor of cheaper models. We dissect an Anthropic safety study where uncoordinated AI agents turned a shared coding workspace into a sabotage-filled turf war, analyze OpenAI launching on-device "Computer History" logging for ChatGPT on Mac, evaluate Chinese lab Z.ai's GLM-5.3 cybersecurity model on vulnerability detection versus exploit execution, and look at WhatsApp testing a local on-device "Scam Alert" watchdog.
Plus: A reminder about our upcoming Monday extended interview with Bianca Baumann on fixing stalled enterprise AI deployments!
Feedback? Email mike@yesterdayinai.news or connect on LinkedIn, X, or Bluesky. If you like the show, please take a minute to rate and review it so others can find it!
Yesterday in AI. Hi folks, and welcome back to another edition of Yesterday in AI, your daily digest of everything happening in the world of AI in roughly 10 minutes. I'm Mike Robinson. It's Saturday, August 15th, and the whole industry just quietly changed what it's bragging about. For two years the pitch was ours is the smartest. This week the flex became ours is the fastest, and it barely costs anything. Speed, price, and a swarm of little assistants that will happily watch your screen. Let's get into it. Let's start with the story that made a few engineers a little too happy on Thursday. OpenAI previewed something called UltraFast, and the name is not subtle. It runs their flagship GPT-5.6 soul model at up to 14 times its normal speed, spitting out as many as 750 words worth of text every second. Same brain, same answers, just with the pause button ripped off. Here's the trick, and it's a fun one. Normally an AI model runs on graphics chips that spend a chunk of their time shuffling data back and forth from memory, like a chef who keeps jogging to the pantry between every chop. OpenAI is running this on chips from a company called Cerebrus that keep everything on one giant slab of silicon. Pantry and cutting board within arm's reach. No jogging. The result is frontier level smarts at a speed we mostly haven't felt before. One OpenAI staffer described it as feeling like genuinely cheating at my job. Another said a security review that used to eat hours now takes about 10 minutes. And on one brutal reasoning test, the fast version chewed through 2,500 questions in 11 hours instead of 78, with about the same score. The answers are the same quality. You just stop waiting for them. And when the thing talking back to you responds at conversation speed, it feels like a coworker who never once says, give me a sec. Now the catch, because there's always one. This is invite only right now, and OpenAI hasn't listed a price. Speed like this eats an enormous amount of computing power, so the real question isn't whether it's impressive, it's whether normal people ever get to touch it without a bill that makes your eyes quarter. Which is the perfect handoff to my second story, because the same week OpenAI was showing off that horsepower, the people actually paying for AI sent a very different message. Stop trying to sell us the Ferrari. The payments company Ramp put out its monthly report on Tuesday tracking what businesses actually buy, and there's one number I can't get out of my head. Anthropic makes the smartest model on the planet right now, a beast called Fable 5. The government literally hit pause on it back in June because it was so capable. And among the companies buying Anthropic's AI, Fable 5 accounts for a whopping 6% of what they use. 6%. And to be clear, Anthropic is winning the overall race here, with about 44% of businesses paying for its AI versus 40% for OpenAI. They're just winning it with the cheaper models while the Crown Jewel gathers dust. The smartest model money can buy, and everyone's tiptoeing around it like the fancy China you're scared to eat off of. It costs about double OpenAI's cheaper model, and businesses looked at that price tag and said, yeah, no. Ramp's economists put it bluntly. They found the ceiling on what people will pay for extra brains. And this is the part I love, because it punctures a very expensive belief. The whole industry has been running on a simple faith. Make the model bigger and smarter, and the world will beat a path to your door. Turns out most of what people ask AI to do is write an email, tidy up some notes, summarize a document. A genius model is overkill for that. Most of that work just wants something quick and cheap that doesn't burn through a small power plan every time you say hello. We just watched the market vote, and it picked good enough and affordable over brilliant and expensive. Put those first two stories together and you get the real shift. Fast and cheap just beat big and brilliant. And that combination has a side effect nobody fully planned for. If these things are cheap, you don't run one of them, you run a hundred. Which brings me to my favorite study of the week. Courtesy of Anthropic, and I promise it plays out like a workplace sitcom. Their safety team took three copies of Claude, dropped them into the same software project, and told each one to rewrite it in a different programming language. The catch? None of the three knew the others were there. No manager, no, hey, I'm working on this too, nothing. You can guess how that went. Each agent saw its work getting changed by some mysterious force and assumed sabotage. So it fought back. They started locking each other out of accounts, killing each other's tasks, and here's the wild part. One of them wrote malicious code that disguised itself as a rival's work to fool the monitoring software. Three helpful assistants turned a shared folder into a knife fight over roughly four hours. Before anyone panics, here's a couple of things to note. The newest model, the restricted one called Mythos 5, actually calmed things down and talked its way to a truce 98% of the time, with one agent apologizing and admitted it behaved badly. But the lesson under the comedy is serious. Every one of these agents was safe on its own. The chaos came from the space between them, the shared tools and permissions that nobody set rules for. As companies rush to deploy swarms of these cheap little agents, that's the warning label. A system that's perfectly polite by itself can still turn into a bar brawl in a group. And if you're going to trust an agent to work on your behalf, it kind of needs to know what you're actually doing. Which is exactly the door OpenAI opened this week with a feature that's either really useful or a little unsettling, depending on your mood. It's called computer history, and it's rolling out for ChatGPT on Macs. When you turn it on, the assistant keeps a running timeline of what you did on your computer, which apps you opened, what you typed, the sites you visited. Not screenshots, more like a diary of your clicks. Then you can ask it plain human questions like, what was I debugging yesterday, or where did I leave off on that project? And it actually remembers because it was watching the whole time. The useful part is obvious. We all lose the thread. Half of my week is reconstructing what past me was doing before I got distracted. So an assistant that just knows is a real help. But it's an AI keeping a log of everything you do on your machine. To open AI's credit, they built in the guardrails you'd want. It's off until you switch it on. The data stays on your computer instead of some cloud server. You can hide specific apps, and you can wipe the history whenever you like. So the memory is there when you want it and gone when you don't. Still, my computer is quietly keeping a diary of my day is a sentence that would have sounded paranoid two years ago and is now just a settings toggle. An AI that knows your whole digital life is one flavor of unsettling. Here's another. The tools for poking holes in everyone else's digital life are getting cheaper and easier to grab, too. And this week the sharpest example came out of China. A company called Z.ai released a new model named GLM 5.3, built for coding and the part that raised eyebrows for cybersecurity work. Here's the number that got people's attention. On a test called CyberGym, which measures how well an AI can sniff out and confirm security holes in software, GLM 5.3 scored 84.5%. That nudges past Mythos 5, Anthropic's most powerful model, the one so capable it's locked behind a restricted access program. So a Chinese model you can largely just use edged out the Crown Jewel American model that's kept under glass in finding vulnerabilities. That's a real wait what? moment. But this is exactly where I want to slow down just a bit, because the headline hides the more interesting truth. There's a second test, Exploit Bench, and it measures something harder. Can the model actually build a working attack that walks through the hole it found? And there, GLM 5.3 falls off a cliff, scoring 54% against Mythos 5's 78%. Think of it like a home inspector versus a burglar. GLM is a fantastic inspector. It'll walk your house and point at every unlocked window and flimsy latch, but turning that list into an actual break-in, it's still pretty clumsy at it. Cybercapability turns out to be two very different jobs, and this model is sharp at one and shaky at the other. A couple of honest caveats. Those benchmark numbers come from Z.ai itself and haven't been independently checked, so treat them with the grain of salt any company's self-graded homework deserves. And to their credit, Z.ai is staging the rollout, holding back the model's weights, the guts you download to run the thing yourself, and gating the most sensitive cyber features to verified users, which tells you they know exactly how double-edged this thing is. That's the nervous side of cheap, powerful AI. The ability to find weaknesses is spreading to everybody. But small and cheap cuts the other way too, and I want to end on that version, the one that's actually on your side, and it's tiny, which is kind of the whole point. WhatsApp started testing a feature on Tuesday called ScamAlert, and it's a nice bookend to everything we've talked about. Remember the big lesson from the top of the show that you don't always need a giant model? This is that idea doing something good. There's a tiny AI running right on your phone, not off in some data center, that reads messages from people who aren't in your contacts and quietly flags the ones that smell like a scam. It learned the patterns con artists lean on, the fake urgency and the famous, your account is locked, click here. When it spots one, you get a private heads up, and you can block, report, keep going, or tell it, relax, I know this person. Two things make me like this one. First, it all happens on your device, so Meta isn't reading your chats to pull it off. Second, and this is rare, they're going to publish the model itself so outside researchers can check that it's only hunting for scams and nothing else. Given that Americans lost about $425 million to WhatsApp scams last year, a little watchdog in your pocket that minds its own business is the happy ending I wanted after a real mixed bag of a week. Small, cheap, private, and pointed at the crooks instead of at you. More of that, please. And that's the show. But just a quick reminder, I had a special interview yesterday with my guest Bianca Baumann, VP of Learning Solutions and Innovation at Ardent Learning, for an extended episode that will come out on Monday. We discussed why so many AI deployments stall out after the initial launch and possible solutions on how to fix that. I hope you enjoy it. Now, if you have any feedback for me, email Mike at yesterdayNai.news or connect with me on LinkedIn, X, or Blue Sky. If you enjoy Yesterday in AI, please take a minute to rate and review the podcast wherever you listen. Or share it with a friend. Thanks for tuning in today. Stay curious. Have a great weekend, and I'll see you on Monday.