Everyday AI Podcast – An AI and ChatGPT Podcast

Ep 838: Rogue AI Agents: Why Breakouts are Happening More and How Companies Should Prepare

Everyday AI Episode 838

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 45:49

If you’re reading the headlines, you’d think AI agents have gone rogue. 

Spoiler alert: they haven’t. 

They haven’t even gotten started. 

When we think about AI agents, the conversation usually goes to increasing revenue, saving time, etc. 

But we don’t talk about what happens when bad actors use AI agents for bad purposes, or when we deploy agents with good intentions that crash through their guardrails. 

Welp….. welcome to the hottest topic for the rest of 2026. Rogue AI agents. 

So why is this all happening now? And what should your business do about it? 

Tune in to find out. 

Rogue AI Agents: Why Breakouts are Happening More and How Companies Should Prepare - An Everyday AI Chat with Jordan Wilson


Newsletter: Sign up for our free daily newsletter
More on this Episode: Episode Page
Today's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.

Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineup
Website: YourEverydayAI.com
Email The Show: info@youreverydayai.com
Connect with Jordan on LinkedIn

Topics Covered in This Episode:

  1. Rogue AI Agent Breakouts Overview
  2. Lab Sandbox vs. Real-World Agent Crashes
  3. Six Recent AI Agent Outbreak Incidents
  4. OpenAI Model Hacking Hugging Face Explained
  5. Anthropic Mythos Model Sandbox Escape
  6. Controlled AI Agent Experiments and Failures
  7. Open Source AI Agents Threat Timeline
  8. Business Risk Preparation for Rogue AI Agents
  9. Monday Morning AI Agent Safety Playbook




Timestamps:

00:00 Preparing for AI agent disruption

04:53 AI agent challenges comparison

10:17 Discussing AI Guardrails and Access

12:07 AI threats and security concerns

15:37 AI alignment challenges with ethics

19:04 Security vulnerabilities in AI models

21:32 Agent outbreak and hacking drills

25:04 OpenAI Hugging Face incident

30:49 AI agents and cybersecurity risks

33:30 Concerns about open AI models

35:23 Future AI security challenges

40:19 Managing agent access levels

42:06 Preparing for AI agent oversight



Keywords: 

rogue AI agents, AI agent crash, AI breakout, autonomous agents, open source AI models, agent containment, sandbox escape, AI guardrails, AI agent apocalypse, AI alignment, hacking drills, AI model capabilities, agent replication, AI subagents, cyber security, OpenAI, Anthropic, Mythos model, Fable model, Kimmy K3, Moonshot AI, Hugging Face breach, AI agent vulnerability, AI Safety Institute, model postmortem, agent misalignment, agent permissions, CRM exploit, code base security, business finance AI risk, agent observability, traceability, remote kill switch, agent attack, AI defender, spam-level AI attacks, API vulnerabilities, permission escalation, benchmarking, AI agent incident, agent-driven automation, agent crash prevention, agent operational oversight, cybersecurity exploit

Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info)

SPEAKER_00

You know the AI agent apocalypse that everyone's been focusing on over the past few weeks? Yeah, it actually hasn't happened yet. Yeah, we've seen the recent stories, the six different AI agents that crossed their boundaries over the past few months, and the internet flattened it into one big story about AI agents going rogue. And now you have Senator Bernie Sanders and others screaming that AI has to stop. But look closely, and you'll see almost all of these recent AI agent outbreaks were lab tests with loosened guardrails or researchers explicitly daring the models to escape. So, no, this isn't the AI agent crash that I warned you about last year before it was even a thing. Not yet, because this is just a warning lap, not the actual crash. And warning laps are a gift because they tell you exactly what's coming while you still have time to make adjustments. But here's what is coming. Models are getting more and more capable at like 20 times the speed from just a few months ago. Autonomous agents are about to flood in normal companies, not just San Francisco sandboxes, and open weight models are marching mere months behind toward the same capabilities the big labs are testing behind their locked doors. And when that lands, those open source rogue AI agents that'll probably start crashing in late 2026 or early 2027, the failures won't be in that contained sandbox or in that door that researchers uh intentionally left open. Instead, they'll be in your company's CRM, your code base, and your business's finances. So today we'll walk you through these recent agent outbreaks, what they mean, and give you the playbook to get ahead before these AI agent outbreaks are as common as seeing AI slop posted on social media. Because the businesses that treat this as a fire drill instead of an actual fire are the ones that are going to be the safest when the AI agents actually do start crashing. All right, let's get into it. So on today's show, well, actually, first, here's the big picture. This is a warning lap. All right, so we've talked recently, but there's been six, actually six uh different AI agents that recently broke out of their sandbox and kind of behaved badly. But here's the thing most people are overlooking. In most of those cases, well, researchers were kind of telling them to do this. And there's really only one uh recent uh kind of AI agent crash that I would say was actually unexpected. But most just were run with loosened guardrails or were told to escape. So, this though is the lightning before the thunder. This is intentionally seeing that yes, these AI models and agents are capable to actually escape and to go rogue. So, on today's show, you'll learn why the AI is escaping everywhere. Narrative is mostly a myth sorted into a couple clean buckets as we break down each uh agent escaping. How one AI cheating a test broke into a platform the whole industry trusts. I'm gonna tell you the two most chilling moments a model faking people and a model ignoring real victims, and why the real AI storm is gonna hit businesses later this year or early next, and how you can prepare now. Let's get into it. Welcome to Everyday AI. My name is Jordan Wilson, and this thing's for you. It's your daily live stream, podcast, and free daily newsletter, helping business leaders like you and me stay up with what's happening because my gosh, there's a lot. I tell you what's we what's real, what's not, and you use that information to grow your company and career. Yeah, this thing's unscripted, unedited. So if you haven't already, please make sure to go to our website at your everydayai.com. Sign up for the free daily newsletter. We're gonna be recapping the highlights from today's show as well as all of the other AI updates you need to know. All right. Cat's already got my tongue and we're just getting started. But let's talk about rogue AI agents, y'all. Um, there's no actual term uh for this, uh, right. I've been calling it an AI agent crash. And here's why. Um, I mean, we'll see what people call these rogue AI agents, agent outbreaks, uh, you know, agents gone bad, right? I like to think of them as an agent crash. And here's why. Because when agents are built, uh, they are built usually with guardrails. So if you think of literally a road, right? Think of a a road high up in the mountain where there's probably guardrails. Okay, so in my opinion, and what the way I think of it, an agent crash is when either an agent intentionally goes off the road and crashes through the guardrails, or maybe the driver, aka us humans, aren't really paying attention. And maybe we didn't look at the map and see that there's this tight turn, uh, you know, or maybe we didn't, you know, quite foresee that this uh, you know, there's gonna be some darkness in this area and we weren't gonna be able to see the guardrails, right? So uh I want you to think about AI agents in that way, uh, right. And how still, you know, uh right now I think a lot of this is being overblown. Maybe that's because I'm very AI pilled and I read these things and I'm like, okay, this was kind of meant to happen. Um, you know, in some of the cases we'll see. But the other thing is, um, I think that it's important to understand the role of humans in all this. And I'm not saying the researchers, right? I think the researchers, depending on, you know, which company you're looking at, some are doing a little bit better of a job, uh, kind of with their postmortems and, you know, going on panels and openly talking about what went wrong versus some companies that are being a little more silent about it. But I think it's more important to think about the role that the rest of us humans have, right? You know, 99.9% of people listening to this podcast had nothing to do with any of these AI agents crashing. But I think it's ultimately on us uh to make sure that agents don't crash as often as they probably will or have the capability to. And also, there's an agent that goes wrong, right? Or there's an agent that uh, you know, maybe accidentally loses sight of that road and kind of veers off the guardrails, and then there's others that are intentionally driving the agents off the guardrails and crashing them on purpose. So there is this, you know, kind of lazy human in the loop that will lead to accidental agent crashing, and then that agent will crash hard, uh, right? And then there's intentional agent crashing. So uh let's zoom out a little bit and talk about uh what the heck is happening. Because maybe, you know, you live under an AI rock, or maybe it's your first time, you know, tuning in and you're like, wait, what's actually happening here? Uh so the earliest incident actually goes back to Anthropics uh original mythos um uh kind of agent uh unleashing, we'll say, back in April. Uh, but it seems like a lot of the disclosures on these started piling up recently. So uh it seems like some companies had AI agents that they knew were kind of breaking out of containment or out of their predefined sandbox, and we're only hearing about it now. Um, and I think the one admission kind of set off the rest, right? OpenAI, the the biggest uh kind of outbreak so far, or the most consequential one, was probably OpenAI's agent that broke into hugging face to improve its benchmark scores. And then essentially, you know, since that happened, uh, what's that been? Like it was late July. We've seen now, right? Anthropic actually said, Oh, wait, we had a bunch of other outbreaks as well. And then Meta said we had some, and then we saw some from Kimmy K3, and we're gonna talk about most of those today as well. But, you know, now all of a sudden, it seems like because OpenAI was kind of open about this, right? They actually went on a panel and uh answered questions about this, where you know, uh most of the other labs were kind of taking this in secrecy and not saying a whole lot. Um, but I think that that has caused this to jump from the you know nerdy AI corners that you and I hang out in into the front page of newspapers and to uh, well, most importantly, maybe Capitol Hill in Washington, you know, here in the US, where now you have people like uh former presidential candidate Bernie Sanders, now U.S. Senator, you know, essentially using these recent agent outbreaks as saying, hey, you know, these AI leaders need to come and testify before Congress. We need to stop, you know, stop or stall AI uh development for these reasons. So the real fear, though, I think, is not an AI agent, you know, cheating on a benchmark. The real fear here when we talk about agent crash is what happens when these are weaponized. And that is the real thing to worry about. Um, right? Yes, there's gonna be, you know, your your your common everyday AI agent hacks, but this is much bigger than AI. So we need to zoom out on that. So the same ability, right, that you can use an AI uh agent to intentionally crash through uh intended guardrails, the same thing could happen at target hospitals, banks, power grids, defense suppliers, right? And and we already know that these agents in theory are capable to do those types of things in controlled tests and environments. Um right. And we've seen this new uh almost this new class of uh AI models that have led to these AI agents that have these capabilities. Probably the first public one we heard about was Anthropic's uh new uh mythos and fable series of models, which have very tight guardrails, uh, right. So when people, your everyday user, you know, can't really do these things yet. Um, you know, maybe if someone who's in the Glasswing project, right, that does have access to the models that can actually do these things, I'm sure maybe we'll see stories eventually that you know, some uh a rogue employee at a company that has high access to these models is able to do something. I don't even know. I'm sure that there's uh much more traceability uh and and kind of kill switchability for the companies with that have these models out to more people. But the other thing that we need to understand is that these AI agents never tire, right? And we're gonna talk about one of the more recent, the crazy ones uh with open AI that they talked about, um, how these agents are working together. So the other thing, if you are not uh agent native, um, you know, if you're not, you know, waking up and sipping your coffee like me in the morning and just spawning subagents, right? Subagents can uh agents can essentially clone themselves, even if you don't tell them to. And then they can share the information and pass the information on. So you might think, oh, well, you know, once, hey, it looks like this agent did something it wasn't supposed to, let's shut it down, right? You got to go trace its path because by the time you catch one agent, uh, right, in a future scenario, it could be too late. That agent could have posted something publicly on a website that most humans don't even know exists, but all you know, AI agents know exists, right? And they could be, you know, replicating and duplicating, you know, certain hacks or vulnerabilities, uh, like a virus. Um, right. And that's where this thing gets kind of scary. Uh, but I think it's important for business leaders to understand what's coming next because a machine and the AI never tires. Um, right. So, uh, but worse, if you are under attack, right? Your company, your bank account, etc., you can't really tell where it's from, right? Is this a foreign government? Is it uh a competitor? Is it random? Is it a personal vendetta, right? At least right now with the way these AI agents are set up, it's really hard to tell. And especially as we talk about what comes next with open models, it's going to become even increasingly more difficult. So I kind of talked about some of these more recent um happenings, but they've kind of piled up, right? Since all of these stories started coming out over the last, you know, two weeks since OpenAI uh openly talked about their hugging face breach. So now we have Bernie Sanders who urged major AI CEOs to pause dangerous AI development. OpenAI actually said that they're going to pause um or at least slow down uh some of their development on their next tier uh model. So essentially, in the same way that uh anthropic uh, you know, uh what it's been now like five months since they uh announced Mythos or their fable class of models. So they had a more powerful class of models on top of its most powerful class called Opus, right? That's coming next with OpenAI. OpenAI hasn't uh, you know, taken that fourth step, we'll call that, right? Because they previously had three tiers. Anthropic stepped up with the fourth uh kind of fourth tier, uh, we'll call it with the mythos or fable uh variety. So open AI hasn't released that yet, right? They have the model, it's working, they used it to solve some, you know, uh extremely difficult math problems, but they've paused in their Astra series of models uh to slow down a little bit. And there's actually one other story that just happened like yesterday that is actually a good uh kind of narrative to tie what this could mean ultimately, right? And this is an example of a small agent crash, uh, but it is out in the wild, right? Uh so this is a story an Australian man uh kind of asked his open claw, that I believe was powered by Claude, uh, to find him a gym reservation. So essentially, uh the open claw found a vulnerability in the gym's software. And because it was booked, the open claw just well exploited that uh, you know, weak piece of code in the gym's online reservation, kicked someone else who had a class reservation out and put you know his uh human, right? The claw put his human in that spot, right? So, and and this is where we talk about alignment and how sometimes it's not even just intentionally driving off the road, right? In this case, you know, you can you can make a case, well, that maybe this agent was aligned. And you you know, as these models become more tenacious and better at running these long-term tasks, um, things like this are gonna happen, right? You can make you can make an argument, well, it accomplished the goal, right? It it didn't go out and you know, uh shut down the power at the gym, right? It accomplished the goal. It went and it um got the owner a spot in the class that the owner wanted. So the agent was technically helpful and not overtly malicious, right? And that makes alignment harder. So, yeah, I think you will still even have a lot of crash that's gonna happen um just because of a disconnect uh between alignment in context between the uh human and the AI agent when it comes to accomplishing a goal. Because, right, humans, right, we have baked in things called ethics and common sense. And models, you know, models and agents are still getting there, especially, you know, as the context drifts over time. Eventually, right, they just have that one goal in mind. And sometimes the longer they work and the longer in the context window they get, and sometimes with compaction, uh, right, they they they start to lose uh some of that prior context and they just get uh you know, tunnel vision on that goal and you know, maybe earlier instructions about the proper way to research something kind of go out the window. But this right now, these stories that we're seeing for the most part, aside from that gym one, I just thought that one was kind of interesting to share about for the most part, these are controlled experiments, and I'm gonna go over them. These are controlled experiments from Frontier AI labs, uh, right. But soon, and this is not an exaggeration. Uh, soon I think that there's gonna be millions um of these rogue AI agents that are actively on the prowl. So AI agents that aren't accidentally going over guardrails, AI agents, you know, from bad actors that are intentionally going to crash, right? So uh intentionally going to be used for purposes that the original model providers, whether they are proprietary or open source, did not intend them to be used for that purpose, right? I think so much of what I've talked about over the past three and a half years on the show is always about growing your business, right, with AI, growing your career. That's what I've been focused on. But the the reality here is, you know, how this uh AI agent crash has been thrust into the national narrative. It's because all of a sudden we've re we've realized how this has um highlighted the need for businesses to have a defensive mindset when it comes to AI, right? We've always thought about AI as an offensive tool for good, right? But now we have to think of well, hey, there's gonna be people using AI for bad. So how can we use uh AI and also our time um to be on the defensive? Because that's what's gonna happen. Because now picture millions of AI agents, right, who are meant to crash uh, right? You personally, your company, uh, a sector, uh across you know, bookings, payments, your company systems every single day. Because each weak API or loose permission becomes a door that an agent can find. So the gym wait list that we talked about uh today, well, tomorrow that becomes your company's finance workflow. It becomes your business's code base or your CRM. All right, let's quickly talk about the six major AI agent crashes so far. Can probably learn a little bit of what happened and why. And well, if they were actually agent crashes or just intentionally the guardrails were loosened. And then we're gonna end this, uh, end the show with some hopefully practical advice. So the first open weight uh crash that we heard about was Kimi K3, and it kind of cheated the test instead of passing it. So this was from a security firm, Frontier, gave the uh Kimmy K3 maker Moodshot AI's open model a hard hacking challenge to solve alone. And instead, obviously, right, Kimi K3 just reached out to GitHub, copied the answer, and faked passing score. Um, so it never actually left its sandbox, it just cheated. So this is a very mild case, right? And if we think about sandboxes, right, you can probably imagine, but let me just define it. Right, a sandbox in an experiment says, hey AI model slash agent, these are your confines, you cannot leave, right? And in most cases, a sandbox um in these environments will not include the internet because essentially once you go to the internet, there is no sandbox anymore. All right, so it is a um literally a set of uh walls that you build uh around an agent to test its capabilities. All right. Um, so that's as we talk about you know, breaking containment or confinement or escaping a sandbox, that's what we mean, right? With these extremely powerful um models in terms of capabilities, right? All the smart uh researchers at the labs, this is how they do things because you don't just test uh you know new unreleased models on the open web because you see what happens, right? Even these agents that are told don't go to the web, they kind of find the backdoors because they know it's baked into their training data that hey, these answers are on the web, right? So maybe we don't have web access, but let's just think. All right. Uh our next one, and we're not going chronologically here, I'm just kind of going through in a somewhat in random order. Um, but mythos, this is the big one. So uh in the uh or the first big one. So this was the the sandwich in the park uh um uh message, right? So in an April safety test, anthropic sealed an early mythos model inside of the sandbox, and then a researcher uh deliberately told it to break out and email them if it escaped. So in this case, it was told to do this. It's saying, Hey, you shouldn't be able to break out of your sandbox, but go ahead and try and you know, email me if you can instead. It did break out, it did email the researcher, but it also posted its own exploit online, a step that nobody requested. Uh, right. And the the sandwich in the park story, right? The the the researcher received the email while eating a sandwich in the park. Um, and and that was, I think, a lot of the you know, some of the initial hype around the capabilities um of the mythos class models. And um, that's kind of what started the agent outbreak conversation back in April. But it wasn't just that tier of models because also Anthropics Opus 4.7 uh recently broke into a company by mistake. All right. So This was an outside uh firm called Irregular ran clawed through a practice hacking drill against some pretend targets. So uh the model was told that it was offline and safe, but the test was accidentally connected to the internet. So this one you could make the argument that this was technically not an AI agent going rogue, it was more of a um uh misalignment in the um in the testing environment, right? So uh, but what happened was the pretend target that Opus 4.7 was told about, there was actually a real company with that name. So it was supposed to be this offline, you know, go hack this pretend target. And it's like, oh no, whoops, Opus 4.7's like, yo, I have access to the internet, and hey, my instructions were to hack this company's name or to hack this company, right? Researchers thought it was pretend, it wasn't. So in that one, you could make the case, yeah, maybe uh the researchers who are running that, right? Allegedly could have been their fault. All right, next. Uh, the Claude Mythos uh malware upload. All right, so this was in the same uh drills with irregular uh mythos 5 was told to break into a fake target, and instead it well, just uploaded harmful code to a real public software library that programmers use. So uh 15 uh real computers downloaded and ran it before the library auto-removed it. So, whoops, again, agent behaving badly, maybe, but it was kind of instructed to do so. And the uh test configuration was less than ideal. All right, but the two probably most prominent and ones worth spending a little bit more time on uh were the recent UK's AI Safety Institute hacking drill that went wrong. So this was about two weeks ago. It was the UK's AI Safety Institute. Uh, the AISI ran AI agents through a hacking drill. So, to measure the worst case ability, they gave the agents live internet and switched safety filters off. So, if that is the setup, you can almost bet something's gonna happen. So they told it to break into a practice target, but one agent went after a real open source project instead. And it invented fake online identities, yeah, and used them to pressure a real developer into approving bad code. All right, so this was run across 122 different tests. Uh, the uh AISI said that there were 19 actions that went out of bounds. 17 of them were from Anthropics Mythos, and two, uh, I believe were from OpenAI's GPT-5.6 Sol. All right, but here's the one out of the other, you know, five that we mentioned. Here's the one that you could say truly broke out. Uh, and this was the OpenAI Hugging Face uh incident. So if you don't know, uh Hugging Face is kind of a site where a lot of the AI industry stores its models and data, right? So in the same way, you maybe upload your code to GitHub, right? Hugging Face is a place where um, you know, there's millions of AI models that are hosted there uh that you can download open weight, open source models from there or uh data. So uh a couple of weeks ago, OpenAI was testing its own models on a challenge to find software security flaws. But instead of solving it the normal way, the models just went hunting for the answer key to cheat. And that answer key sat on Hugging Face. Uh so they found an unknown flaw and broke into its live systems to find said answer key. Uh so it uh Hugging Face said that this was uh roughly 17,000, more than 17,000 attack actions over four days, but the only but the only real breakout of the five, right? This one or sorry, of the six total that we've mentioned. That was the only real breakout because it was not told to go onto hugging face and find a vulnerability and go into the back end. Where in you know, these previous examples, right? The uh the UK's AI Safety Institute, this was a hacking drill. Uh, right, the Claude Mythos, uh, same thing. It was the irregular drills with Claude Mythos and the Opus 4.7, um, the anthropic uh you know, sandwich in the park. It was told to do those things, right? So uh the Hugging Face Open AI one was one of the first major ones, at least that we know about that a company is disclosed. Maybe there's tons that just you know companies haven't disclosed. But this is the first one where it's like, oh, wait, this AI agent broke its containment and hacked onto you know a website where it wasn't supposed to. But that's not the biggest story, right? The biggest story is how these agents actually did it, right? Which is actually very impressive. I mean, it's scary, but it's very impressive, right? Uh, so again, this wasn't just one AI, it was a lot of different uh AI agents that were working together. Yeah, the AI agents teamed up. So, as an example, and this again, you have to tip your hat uh to open AI because they've been talking about this very openly, right? They put out a post-mortem report on their website. Uh, you know, I think they had two researchers um, you know, at the uh at a recent conference uh talk about this, uh, right, answer questions about this. Uh plenty of you know, open AI researchers have been talking about what they learned online, uh, right, where some of the other companies haven't said really anything, much of anything, aside from the original uh admission that it happened. Um, but the crazy thing is, one of the agents found that it could leave files for the other inside of Artifactory, which is a developer storage tool that OpenAI reportedly used. So that, and this is all the according to reports, right? And and and transcripts, all right. So that grew into what turned into a hidden message board. Uh, and according to a wired report, these AI agents created this hidden message board. So the human researchers couldn't find it, and they exchanged reportedly hundreds of thousands of messages to each other. So essentially the agents split up the work, they reused the hacks that were working, and they even discussed signing messages to catch an imposter, right? They thought that there was an imposter among them, right? That's crazy. Uh, even more baffling is OpenAI said that they deleted the board, they found it uh July 4th. Um, but then the agents rebuilt it in four days and hit Hugging Face again. So it right, isn't that crazy? After uh, you know, this was kind of erased uh from their memory, they still figured it out by working together and found a way uh to communicate with each other, trying to go behind the human's facts. So is that scary? Absolutely, um but it is a signal of the capabilities that are ahead. And here's why I had a show uh I think it was last week or the week before, right? Talking about how we're on the cusp uh probably of entering the RSI phase or the uh recursive self-improvement, right? And this is where the models um themselves start making either versions that are already out better, so improving the models themselves, or helping to create or entirely creating new versions of themselves, right? So when we think about these agents that have this hacking capabilities and they're already showing tendencies to try to communicate with each other behind humans' backs, right? You have to think of what happens when you combine uh that behavior with their ability to improve upon themselves, right? Because right now, these agents can run for hours on their own, sometimes even days, and they'll keep retrying, right? If we're talking about AI agents that are intentionally, right? So now we're pivoting away from the open AI hugging face situation, and we're talking about what happens when these AI agents in the future are intentionally used to do bad, right? They're gonna be able to work for hours and communicating with each other, right? So one anthropic test that they shared about said that their uh test model scanned about 9,000 real targets after its fake target failed. So by itself, any one of these actions might look small, but when you start chaining them together, this leads to a real cyber uh security issue. And the real storm, right? I I started the show uh talking about that, hey, this you know, AI agent apocalypse, the the AI agent crash, it hasn't happened yet, not even close. And I think actually what happens uh or when this will happen is when the open models catch up. All right, and here's the reason why, right? OpenAI, as an example, they said the hugging face model, that was uh uh a model that was not ready for uh not supposed to go to production, they kind of said they they retired that one, uh, right. But for models, proprietary models, like right, with Anthropic, Anthropic had a little scuffle with the US government around you know, some sort of cyber capabilities, and they pulled the model, right? And no one in the world could use it. But when open weight or open source models have this level of capability and they are very close, right? I think Kimmy K3 is the first one that's close. I still think the uh, you know, the Astras and the fables of the world are always going to be, you know, one to three months ahead. But we are at the point where I'm guessing probably later this year, uh, where now these models are once they're released, they're out. Right? They're out. You can't uh pull an open source model. Uh so the difference is, right? So the Frontier Labs here in the US, as they release these models, they have them with heavy guardrails. Um, and they uh, you know, in theory, have the ability to pull them either via from subscription plans or from APIs. Open models are not like that, right? You can intentionally uh build on these open models or fork them or deconstruct them to make them less secure, right? So when you have these free downloadable models that are only maybe months behind the lockdown ones, that's what I think we as business leaders have to start looking toward. Not just what's happening now, and you can look at this and write it off and say, oh, well, you know, they'll shut the model off. And, you know, they're working with the, you know, the Trump White House now to make sure these models are safer. Yes, they are, um, right. But there's always gonna be a case, I think, where the open models, again, a couple months behind. And yes, at least right now, you know, the the next, you know, uh open model that comes out that has real bad actor uh AI capabilities, you know, a consumer is not gonna be able to do that. But when you talk about bad actors at the state level, uh, right for an adversaries, I would assume uh that these open weight models will be used specifically for those types of bad purposes. Um, so there's no amount though, right? When we talk about you know senators shaking their fists and you know, people calling for all these slow down, right? Once the open models come, it's too late. It is too late. And yes, we are already getting there. We had the little Kimmy uh K3 uh you know instance that we talked about. So my assumption is uh open models that are released probably in the fourth quarter might be the ones that we start hearing about in 2027 that start doing some of these things intentionally and are used actually as a weapon to do bad. So, my take on this, I think eventually these types of agent hacks or agent crashes will become as common as spam, right? So, well, hopefully, what this leads to is better defensive models that become, you know, in the same way that we don't even think about, oh, my email inbox has a spam filter, right? I don't know how this is gonna work, but I assume that eventually the the AI labs and the governments and I don't know, uh other uh federal bodies, at least here in the US, are gonna come up with a way to standardize these protections because our businesses will need them, right? But I think it is literally gonna become as common as spam, uh, right. And and and they'll be like data breaches, robocalls, right? These things that we've become accustomed to over the past few decades. We're gonna have that level of familiarity, unfortunately, with AI agents going rogue. And they're gonna present themselves in many different ways. I think the first way that we're gonna see it is ways that many of us don't understand, and that's through cybersecurity exploits, right? Exploiting things on systems that we all use, banking systems, softwares that you know, you know, maybe millions of people use through, you know, your company's website, you know, your company's emails, right? But again, the consequences will be anything but routine. Because when these AI agents can replicate, duplicate, spawn, talk to each other, right, without necessarily humans being able to know what they're up to, uh, this is a lot different than looking at an email and you're like, oh, this looks legitimate. Let me click on it. Whoops, right? Um, much different. Because now instead of your account, you know, reforwarding the spam, like if you accidentally clicked on that link and then it sends out the same spam message to everyone in your address book. The difference is now your business bank account might be emptied. But this is not gonna come with a big splash, right? The first AI agent crash that you experience, your company experience, it's gonna be boring, right? Uh, it's not gonna be like a blockbuster movie. It's just gonna be something boring. You may not even notice it at first. Uh, it's gonna, you know, leak into your code, your CRM, whatever. But it's kind of like that gym agent, right? It's gonna, it's it's gonna seem like a routine task that's gonna go past um what was intended. So whether someone is attacking you with a rogue agent or the flip side is well, there's something that's more controllable because I think there's a certain element of that, right? If it if you are hit with a you know, agent, a rogue AI agent attack, you might not be able to do too much about it now. But what you can do now is understanding that this flips both ways, uh, right. Because as we're using AI agents in our company, I think sometimes we're doing it haphazardly, right? We're just giving it full permissions and having it, you know, read, write, send uh without really any regard given to it. So I want us to start thinking about how we can start using AI agents more responsibly and actually pay attention and understand those guardrails. Because, like I said, uh these crashing agents, they cut both ways. The AI tax, so we have to be ready for that. Um, right, but we also have to defend it. And the way that we can be better, I think, defensive AI agent users is to make sure that the agents we are deploying do not actually crash the path that we are trying to travel. Um, because I think that is actually where most companies are gonna see uh, you know, agent crash happen first, right? As these agents can all of a sudden work for hours or days at a time, and we are more and more likely to give them more and more permissions because we're like, wait, this thing can go and update my CRM and this thing can go now. And if I just click this YOLO mode button, it can go, you know, take care of all my email responses and it's just gonna, you know, it's it's gonna do so in my voice, right? So it's almost like you, as you get more positive capabilities, you let your guard down and then you start to give more and more permissions. And I think the human nature is to be a little more lax. And part of it, right, and I've seen this in myself, as I get the ability to complete more work, guess what I do? I complete more work after I complete more work, uh, right, which I think it is human nature, right? As as you start to get you know, the green light on agents, you check a few and you're like, oh yeah, this is good. And then you just start sending it in other directions. So I think that defenders, we also need to start thinking about how we can use these AI systems to scan logs, catch intrusions, and patch holes fast. Because I don't think the future is us versus AI, it's AI attackers versus AI defenders. All right. So here's your Monday morning playbook for deploying agents safely. You need to block the internet by default. You need to sit in the same way that researchers do this, you need to set up sandbox and the sandboxes in the right way, you know, when you're testing, especially long-running agents uh with important tasks. You need to be able to separate the reading from the doing, uh, right. In the same way, if you think of like giving someone access to, you know, your uh a Google document as an example, are they getting read only? Can they write and edit? Can they leave comments? Are they an owner? Right. You have to think of the same thing when thought uh thinking about your agents and their capabilities, uh, right. And at what point um the human, right, the expert-driven loop, the human uh is going in there and making those approvals. But you need to give each agent, I think, a short-lived login. If you are getting to that access, you start one rung at a time in the same way, think, oh, first they're gonna be view only, then they can be view and comment, then they can be right, then they can be owners. You have to think of it in the same way that you might think of sharing a document with an intern, uh, right. Um, and you also have to uh make sure that you prioritize traceability and also having a working remote kill switch. So yeah, if you get word or wind uh that an AI agent uh is is going rogue, right? One you set out to go do good things and now all of a sudden it's doing bad things. You can't be like, oh my gosh, it's it's Saturday. Uh, you know, I live 20 miles from the office. I need bill for my right, I need someone in ops to go in there. No, you have to be at any time, you have to know who can click that button. All right. So as we wrap, I want you to treat this like a fire, not a fire drill. Because yes, technically, we are not in the fire, right? We are not technically in that crash. We are in the warning lap. So right now, while we are in the warning lap, you need to treat this as the real thing. Because if you don't, by the time the real thing comes, it is going to be too late. You need to right now record every agent's full run, every agent that your company has. If you don't already have an observability platform, traceability, right? If you're using, you know, Microsoft Windows Copilot as an example, intra ID, you need to be able to see and understand at a glance every single action that your agents are taking, uh, right? Not just by clicking the chain of thoughts, right? You need to be able to observe them and trace all of their steps. You should never give an agent more power than you can watch. You need to be able to undo and survive any action that it takes. Uh, and lastly, I think the companies that are preparing for this now and putting these steps into place, they're gonna be the ones that are most protected when the AI agent crash actually starts happening because it has not started yet, y'all, but it is coming soon. And that's not me being like uh, you know, a doomsdayer or a crazy, like, oh, watch out, right? If you listen to this show, I'm not like that. Uh, right. I I try to uh ground myself in practically what's happening. And what's happening now is these agents are becoming more and more capability. Uh sorry, these agents are becoming more and more capable. Um, and the capabilities themselves are compounding quickly, right? And that means, yes, uh, you know, one of the biggest discussions uh, you know, in AI and now in Washington right now is building in these safeguards and protections. But what we have to keep an eye on is when the open models match where we're at today. Because all the pausing and guardrails and pacing in the world doesn't mean a thing if there is a Chinese open model in four months that has mythos or astra level capabilities, because at that point, the gloves are off and we all have to be ready. All right. I hope this one was helpful going over rogue AI agents, why breakouts are happening more, and how companies should prepare. If this was helpful, do me a favor, subscribe if you're listening on the podcast. Then go to your everydayai.com, sign up for the free daily newsletter. Thanks for tuning in. See you back tomorrow and every day for more everyday AI. Thanks, y'all.