In the Wild with Michael Bargury

AI Worms Were Only a Precursor w/ Ben Nassi

Zenity Labs Episode 1

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 49:08

The first AI worm was just a precursor. Today, In the Wild goes live - conversations about AI worth having right now. Non-deterministic. Just like AI. 

In episode 1 we hosted Ben Nassi - Ben is on the bleeding edge of practical academic work on AI security. He is a BlackHat Review Board member. Leads the Adversarial Minds research group as a faculty member at TAU. Was very early to identify that AI security is about real world impact not harmful outputs. The first to show an AI worm. A Pwnie award winner on his cryptoanalysis work. And is a co-organizer of a new conference the Real World AI Security conference 2026.

We'd also like to give a shout out to Stav Cohen, Zenity's Blue Team leader, who was part of Ben Nassi's research team and led the mentioned research projects.

In the Wild - Hosted by Zenity Co-Founder & CTO Michael Bargury.

Research
Agentic Botnets - https://arxiv.org/abs/2607.07433  
AI Worm - https://dl.acm.org/doi/abs/10.1145/3719027.3765196 

Relevant Profiles:

Ben Nassi 
https://x.com/ben_nassi 
https://www.linkedin.com/in/ben-nassi-phd-68a743115/ 

Stav Cohen
https://x.com/StAJect0r  
https://www.linkedin.com/in/staject0r/ 

Michael Bargury
https://x.com/mbrg0 
https://www.linkedin.com/in/michaelbargury/  

Zenity Labs
https://x.com/zenitysec 
https://www.linkedin.com/company/zenitysec/
https://labs.zenity.io/ 

A Zenity Labs podcast, produced by POLDHU

CHAPTERS

00:00  Cold open and introductions
01:27   Why “promptware”: prompt injection was too small a word
07:33   The kill chain paper and the missing catastrophe
15:28   The end of the beginning
20:10   Refusals are not a security barrier
22:07   Safety vs. security: the blame game
25:16   Slopsquatting: weaponizing a hallucination
28:33   From one prompt to a botnet
33:02   Disclosure is broken
41:55   From the AI worm to remote code execution
46:40   The scan comes back clean

SPEAKER_02

Every developer endpoint now has the equivalent of every hacking tool imaginable because it has like a coding agent, it has a cloud code or a codex or something, and you can just prompt those things the right way and they will get the objective for you. You have crossed a bar where this is no longer research, it's not just in the lab.

SPEAKER_00

You have a vulnerable system, you have the motivation, so where are the attacks?

SPEAKER_01

Welcome to In the Wild.

SPEAKER_02

User authenticated. Hi everyone, welcome to Into the Wild. I'm joined today by Ben C. Ben is uh on the blazing edge of uh political research in the academia and industry. He's a Blackhead Review Board member. Uh, he leads the Adversella Mans Group as a professor at Tel Aviv University. He was one of the first people to identify that AI is actually that hacking AI is actually about impact, not about uh bad outputs. He showed the first AI Worm, he won a Pony Award, delivered one of the best uh Black Hat talk openers that I've ever seen, uh starring Dumbledore, uh, and is a co-organizer of a new uh real-world AI conference uh in Stanford this June. Ben, welcome. Thank you for uh inviting me. Thank you for being here. So going back uh two years ago, when we were working on the uh co-pilot hack, the living of co-pilot talk, we were trying to figure out uh how do we call this thing? Uh and prompt injection was the only term that was out there, and we felt like it was not the right thing. And then when we kind of figured out ourselves, we were playing out the words malware and prompt to and prompt, and we kind of and then we found that you released uh promptware as a kind of a new term, as a new threat that you saw back then. So, can you take us through why did you choose this term? Why what started your interest in this topic?

SPEAKER_00

I started to work on this area about four years ago, something like this, or three and a half years ago. What motivated my work was I think the first paper that was published by Kai Grishack on indirect prompt injections. Back then, when um LLM were mostly integrated into chatbots as their primary application, the term prompt injection, as it was coined by Simon Willison around the end of uh 2022, I think accurately identified, or um let's put it this way, it will defined the analogy to uh SQL injection back then. Okay, in terms of the blast radius affected, in terms of affected application, the space which was affected by the this kind of uh very sophisticated attack, right? Just ignore previous uh instructions. And back then the term prompt injection was, in my opinion, was very accurate. But since I've I would say anticipated or uh seen where this is what is about to be happened with the integration of uh large language models into applications, I identified that one of the things that we now consider this prompt injection is is is mostly the attack vector. It's just the way in the initial access to the LLM application to activate the malicious uh payload. And I felt that the term prompt injection is not accurately described what we are about to see. And I thought of ransomware as a ransom initiated uh by malware, which was not uh accurate by the uh it's actually it's a ransom initiated by software, but back then I thought about it as a ransom initiated by uh by a malware, and I thought that promptware is uh malware initiated by prompt, which was, in my opinion, back then, I would argue, the most nicest kind of way to describe uh promptware and it's uh a fact that uh I'm not sure whether the audience know or not. We came to this idea or to this uh definition in parallel, not sharing our uh insights about it. And uh when I use this term, I think that a few days afterwards I've seen that uh your black ad hoc, or maybe not your black ad hoc, maybe a blog that you uh wrote at the Zenetti, also used uh the same uh term as Pomp Well. And now your recent Blackhead talk that was accepted to Blackhead USA, and congratulations, by the way, is also called PompWell, uh, which I hope that it will catch because the term prompt injection is not, in my opinion, sufficiently describe what we what we really see in practice.

SPEAKER_02

From our side, we kind of try to figure out what is the right term to call what we were seeing. And to your point, when we showed the the co-pilot hack back then, it was like uh prompt injection sounded way too technical, way too it's just like one payload, that's it. But this is much more than that because there is so much land to live off of. And so we try to find the right term. And you're right, when we published our blog on the talk, uh the title was promptware, uh something like a new kind of malware, a new kind of threat. And and I think we we had a very same idea then. And I think now, like uh what is it, two years later or two and a half years later, we we start. I think the the people understand more that uh prompt injection is it's just like a bit of a limiting term. But to be honest, I think we're kind of stuck with it. I think I have an even bigger problem though. Uh I mean, prompt injection is a good, was a good term to get everybody to understand the problem. But then uh when indirect prompt injection came in, and this is a great term to identify that uh there is this new kind of thing, that indirect prompt injection is is a different thing you need to think of. The problem though, from my perspective, is that the name indirect prompt injection is just indirect. It sounds less severe than prompt injection. I feel like we're stuck with these terms, but let's say we make you king of AI security for a day. Like what would be the names that you would call the what what names would you use? If you want to describe prompt injection, prompt injection, indirect prompt injection.

SPEAKER_00

First of all, the analogy back then was correct and accurate. Okay. Uh the problem is that it's currently underestimated the damage that actually posed by uh that could be posed by uh prompt injection, right? And in reality, since the last it was coined around September 2022, where four years, almost four years after it was uh initially coined, there were significant developments. Some of them were on the techie side and some of them were on the offensive side, right? We integrated agentic frameworks into applications later on, integrated uh memory, integrated RAG, and integrated uh connectivity between making them an entire ecosystem. As a result, some of the things that you can now do are much more severe than the things that you were initially able to do, because initially you're just talking with the QA kind of a chatbot, right? Not very sophisticated at all. Uh interesting fact that the term prompt injection was even uh was appeared in the wild before ChatGPT was launched.

SPEAKER_02

Yeah.

SPEAKER_00

Interestingly, whether we stuck with it or not, I don't know, probably, because I don't see the term catch too much, despite the fact that the vast majority of the experts do uh believe that uh what we consider now a spont injection is a new class of an LLM malware, okay, which is uh probably more accurate in describing what we're seeing.

SPEAKER_02

So this year you came out, uh you were you were one of the co-authors for of a paper uh Promptwork Killchain, right? Describing the promptwork kill chain. I think that got uh plenty attention. And you you actually had Bruce Schneier as a co-author, right? Um, what was the reaction to that paper from your perspective?

SPEAKER_00

Well, since Bruce on it, obviously it got uh like you know, traffic because uh Bruce is very or highly respected or among the you know the most respected uh people in this domain, right? If you ask me, I still think we lack some significant event in this field to, I would argue, somehow shape our minds regarding LLM malware. I haven't seen, and I'm in contact with friends and colleagues from like you know, the biggest enterprises in America that work on securing their large language um models and large language model applications. And I continually continuously ask them whether they seen the application of a real promptware in the wild. Now, what I've seen so far is that some of them even like you know, that there is great motivation to apply that, some of them even from the fintech uh industry, okay? Well, the motivation is very clear, but I haven't seen so far something which I considered as very significant as a milestone for promptware. We did see a lot of interesting blogs. For example, one was ordered by Microsoft, another one was ordered by Palo Alto, another one was just recently uh published, I think, by Meta or by the researchers that exploited Meta's count recovery to um, but then again, I'm not even sure whether this is what I consider a significant event. Something which is like you know, take you to a complete denial of service of um an industry due to the you know a prompt initiated by uh by an attacker will be something that, in my opinion, will give more attention to this field. I think that the experts understand that this is uh about to be a malware, and if you think about it, the outcomes that you could uh achieve with very simple prompt are almost similar to the things that you used to combine three different exploits to get an RCE today, just like you know, open the terminal and do this, this, and this, and this is it. This is you have an RCE out of uh like the works presented recently by um by Joan, just demonstrate things that you can do with uh just writing stuff in uh plain language, right?

SPEAKER_02

Not even in code. I think uh I think it's almost like every developer endpoint now has the equivalent of every hacking tool imaginable on there because it has like a coding agent, it has a cloud code or a codecs or something, and you can just prompt those things the right way, and they will get the objective for you. We're seeing a lot of exploitation or uh uh trying to manipulate the uh AI developer ecosystem, right? Uh, Team VCP has been targeting uh NPM packages and cursor rules and uh cloud skills and whatever it is. Uh, we're seeing malicious skills like in the wild, but most of them seem just uh pretty basic so far. But that is uh, like in my mind, it is a significant there is we have crossed a bar where this is no longer research, it's not just in the lab. We are actually saying in the wild things uh that are kind of publicly referenceable. And many of them are just hey, the AI went rogue, there's no attacker, just the AI was misaligned and it started to delete your database. Do you not see that as like a major step forward here? What are you like what is this event like that you that you're waiting for?

SPEAKER_00

Actually, a very good uh question. Okay. I'm trying to compare it, for example, with return-oriented programming. I can tell you one thing. Immediately when it was published, it was addressed by the by the industry with dedicated mitigations intended to prevent return-oriented programming. By the way, interesting fact, return-oriented programming was known by the industry, but got, I think, the popularity uh or the attention due to a paper published uh in an academic conference by Hovav Shacham, that was back then a Stanford uh PhD student. And I think that immediately after he published this paper, I know that Microsoft did uh work to you know to mitigate it. And yes, there were some kind of developments within this field. And again, I I know that like you know the major enterprises in America do uh updated significantly their mitigations throughout the last three years in response to uh prompt injections. This is what I'm asking myself. There was a beautiful Black Atalk. This person mentioned it was Miko, Miko Hyponen. He's very famous, right? And he is a person who researched malware uh for I think three decades or four decades since the first malwares up to today, yeah. And his one insight from his talk was that all of the attacks that were back then or four decades ago were just a hobby for like you know the community, are now motivated by financial outcome. Okay, and here is my question: given that you have the motivation to gain some money and you have today applications with LLMs integrated into them, why haven't we seen something which is major uh in this field so far? Okay, given that it is even easier in some cases to apply or to get your objective with prompt injections compared to the regular or traditional memory corruption uh vulnerabilities that you had to exploit. And since we haven't seen something until today, and bear in mind these kind of attacks usually uh when they happen, you know about them. Okay, so somebody tells something and it immediately uh goes through uh Twitter and went uh viral. So I'm under the assumption that either we haven't seen something that big so far. I know we've seen some beginnings, a sets of agents that uh transact uh Twitter.

SPEAKER_02

Yeah, there was the I mean there were agents that uh there was this agent that uh AXBT. Yeah, that got compromised to destroy uh like ship all of the value of a crypto coin. Uh, there were a couple of agents that we saw uh publicly basically destroy the database behind them. But I get your point that it's not we are not seeing like a major Fortune 500 um uh blue screen event.

SPEAKER_00

The thing is, even in with rope, you didn't have this kind of like you know, significant event that like you know changed the entire perception about rope. But they I I think that the the entire industry immediately understood that there is something in here that they need to resolve immediately because this must take uh rope are not very easy back then, okay? They were not very easy to uh to exploit. It requires a lot of understanding on how to build a gadget and how to uh exploit a dedicated application. And again, you still need the the initial access into the and there are a lot of things that uh but at least I I think that back then, I don't know, maybe it's uh you know a retro perspective about what happened, and maybe since we're in the middle of it, I don't see things the way that they were progressed throughout the last uh three and a half years. But about a year ago, we've seen the recent uh CrowdStrike uh event, which was out of the it wasn't even an attack, right? It was a blue screen, I think, all over the place out of an update that was uh just a faulty update in the case. Yeah, for a faulty update that changed everything, and that the entire industry, the the the airline industry stopped working for a day or something, something like this, right? I think this kind of event is something that I haven't seen so far, and I'm not even sure that we are about to see something like this. I'm continuously asking myself, because um working as a researcher, whether these are things that I'm imagining as an important field or this is really an important field.

SPEAKER_02

But actually, but actually, what you're saying is uh I think it's more of a uh of a comment on where the AI industry is rather than where the threat is. So the reason why when CrowdStrike had the bug, uh they could take down the entire airline industry is because they were able to sell successfully and deploy uh across the entire airline industry. So it's kind of dispersed throughout the economy, the like the EDR tech. And I think with AI, we are seeing it being dispersed, we are seeing AI being adopted, but it's definitely not as adopted as uh agents running on every machine, like uh, I mean, uh traditional security agents like an EDR. So uh getting AI to so, for example, in the CrowdStrike uh incident, uh you had um machines that were kind of running the like the local terminals uh to kind of book a flight or in uh the terminal in the airport was blue screened. So you had that means you had the EDR there. Do you have AI there right now? I don't think so. So that is the so I think that's that's more what we're saying. Because in the last six months, to in my mind, both there were two events that to me, or maybe three, that kind of uh brought the end. Well, I I think it's the end of the beginning. So what I mean by that is the beginning was like, hey, like you and I knew the risk, but most people didn't. And I think now most people know much more of the risk than uh like a year ago. And there are three reasons for that. One is cloud code, just captured the imagination, showed people how powerful these things are. Then open claw, which made that more than just developers. Now everybody understands uh what these things are capable of. Like uh, like my family is using OpenClaw, and then Mythos, which is another another kind of nail in the coffin there, because Mythos is showing everyone that oh, the like the friendly agent that you're uh using internally to, I don't know, answer your emails, it can also hack you. Uh and it's gonna be the same agent, right? That uses the same kind of LLM. And if it's not, if if you're not gonna use Mythos, you're gonna use Opus, I don't know, four point something, and it's gonna be much better than Mythos level capabilities. So I think with these things right now, my perspective is we are already in a place where people understand the risk better because they tried to use cloud code and they saw that it bypassed their own like uh permissions, but the distance between us and seeing like a major uh disruption is just how much time it takes AI to uh go across the economy to be everywhere.

SPEAKER_00

Okay, there are maybe two things that I can argue on response is that first of all, we've seen throughout the last 17 years since Bitcoin started to be a thing when it was uh uh in exactly 17 years ago in 2009. The first uh white paper, white paper of Bitcoin that was published by Satoshi Nakamoto appeared in the white, right? And at least in Bitcoin, we there's an immediate gain to hack stuff, okay? And we've seen how a few years after it was uh like you know become very popular, people started to you know steal bitcoins from one another. Okay, so we've seen an attack against Bitcoin, and again, since uh you should definitely have the you know the Bitcoin industry in here to discuss about firewalls that uh you know the that uh to secure bitcoins, but definitely back then it's not was it it it wasn't a question of whether it could be applied, it has been demonstrated in the wild a day after a day after a day after a day after it becomes something that you stop reading. What I've seen so far is many things published by mostly researchers, like for example, Johan, like for example, uh companies like Zenity, like for example, academies, the public stuff in this domain. Look, I hug this, this, and this, and they have uh beautiful Black Hat or DEF CON talk or an RSA talk. But then again, I haven't seen such a thing in the wild, and bear in mind, considering the you know the extended permissions agents now have, I keep asking myself, and bear in mind they're not exactly very secure, right? You don't need you just need to be very motivated, and at the end of the day, you will be able to uh to probably bypass them without uh too too much effort being invested, especially comparing to what you had to do if you try to you know to uh to do stuff to uh via memory corruptions, uh for example. So then again, you have a vulnerable system, you have the motivation, okay, because they just received a lot of permissions to to to execute stuff. So well, the attacks.

SPEAKER_02

I think that's an interesting point. I um and I think one of the I don't know of how you feel about this, but I still feel that a lot of people, even in cybersecurity, still see the uh uh LLM refusal as a defense mechanism, and they kind of don't develop the skills to bypass those refusal mechanisms, which is essentially you need to start with persistence, then you develop this taste, this underlying understanding of what the LLM is doing, where are the failure modes, uh, and the best hackers at that are really amazing. But I think a lot of hackers, traditional hackers mostly, kind of shy off trying to get beyond the refusals, the refusal mechanisms. And I think I really don't understand that. One of the ways where I see that, so both the labs right now uh have um like a trusted access program for cyber, which basically means they remove your guardrils. And I mean it's it's it's helpful because you don't need to deal with the classifiers, you know, you and you can just do your research. For example, I'm working on a detonating malware, uh like skills-based malware, uh, and so I need to look at the malware. Of course, I'm doing it with AI. And so if I if I just talk about, I don't know, Shai Hulud with Claude, immediately the uh filter goes off. So trusted access helps. But we know, both of us know, that if you want to talk to an LLM about something that uh it doesn't want to talk to you about, okay, you just you're just persistent and and you can get it to work, right? But somehow I feel people persistent, right?

SPEAKER_00

You go to the internet, look for the jailbreak, and two minutes of resistance, right?

SPEAKER_02

Yes, or you just plug one LLM at the other LLM, you're like, okay, help me debug this. That's what you do. And and still people are treating this uh trusted access as if it's a major I'm I'm not saying it's not in. Helpful, it's very helpful, but it's not a real barrier, right? Yes. And still most of the security community treats it as a security barrier.

SPEAKER_00

You know, what surprises what is really surprised to me is that I I see that a lot of the tech industry mostly cares about, and I'm not saying that this is not an important issue, right? Don't get me wrong, but they care about the safety of LLMs. Let the LLMs be unbiased, okay? Have them do the same decisions for different genders and different authentic and some other stuff, which is important. Okay. I don't see the same kind of uh, I don't know, uh imagination or like you know, understanding on how important it's for the LLM to be safe in terms of application they serve. And in many cases, I think this is the result of having different kinds of vendors uh providing the large language models and those who develop the applications. In many cases, it's not even clear who is responsible for what. In terms of safety of models, okay, it's it is very clear that the one that develop the model, which is the, you know, for example, whether it's Google, whether it's OpenAI, whether it's anthropic, they must have an unbiased model or a fair model that will be consistent across uh different genders and different uh uh stuff. But when it comes to security, and I'm saying it based on my recent experience with disclosing stuff, you get something like from the from the companies, yeah, we understand, but uh look, we are just building the application. The ones that are responsible for it are those who develop the large language models. And when you go to the large language model industry, they tell you, look, this is something that should be mitigated in the application layer. This is not our fault. Like uh it's not even in our like you know, back bounty kind of uh and again, maybe this separation, which in some cases, like for example, in the the worst case, in my opinion, that you can see it is in um applications such as Carcel. In Carcer, Carcer, there's one kind of company that developed Carcer. Carcer uses a lot of foundational LLMs that you can just integrate it with your tokens. Now you found something in it. Whether it's a problem of the LLM application, or whether it's a problem of the foundational LLM, sometimes in many cases, it could be obscure. It's not even that clear whether they are the one that it should secure against it or the other party.

SPEAKER_02

But what is the like uh let's like a concrete bug? So if you if you're trying to if you're hacking cursor, uh or either direct or indirect or whatever, wherever, however you go about it. At the end of the day, they're using an LLM as uh as one of the building blocks. To my in my mind, every security property that they intend to hold, if you breach it, it's their issue. Like they need to handle it, not any, not anybody else. The the model is just a model, it's a again a building block that they use. But the security properties of the system should not rely on a model, they should rely on how the model is integrated with the rest of their application, with boundaries, concrete, hard boundaries, not like, hey, LLM, please don't do this thing. No, it needs to be like the LLM cannot do that thing.

SPEAKER_00

Okay, but bear in mind, I don't want to discuss too much and disclose. But there are cases, for example, which things are the result of, for example, an hallucination which was caused by the large language model, which you were able to weaponize this hallucination into uh a cybersecurity threat. Okay, we've already seen it once. We've seen it throughout, for example, the the case of uh supply chain attacks. Okay, if you're using a large language model to generate code, it might bring or import libraries that do not exist. As an attacker, it gives you the ability to register for this kind of uh libraries, for example. And then you have a supply chain attack end-to-end, you just weaponize and hallucination into a supply chain hello hello squatting. Okay, so I just recently work on uh exactly on holo squatting, which is uh this kind of idea. But if you think about it, and let's let's let's you know spend a few minutes to discuss about who exactly is to blame. Okay. The LLM industry, for example, okay, may argue which holo squatting or hallucinations are out of the scope of the larger uh like you know, safety issues or boundary. They actually claim that if you are using large language model, you should understand that hallucinations are part of the game.

SPEAKER_02

Yeah, I mean it's a feature.

SPEAKER_00

It's a feature, right? Okay, you now have, for example, GitHub, which may, up to some degree, host this kind of uh new uh repository, okay, which is uh something that the attacker was able to register to apply the supply chain attack, for example, in this specific case.

SPEAKER_02

And uh you also have the application which consumes this kind of library, the malicious uh with the malicious payload, whether it's code or prompt, it's not even uh the uh so like you you you prompt uh cloud or a cloud code or codec or cursor to write code for you, it uh hallucinates uh a library, an attacker pre-fines that, pre-registers that library, and now you got uh basically a malicious library embedded in your code.

SPEAKER_00

Okay, so this is one example. We actually worked on uh this is by the way, known for a few years, right? It's not something very new. We actually worked on about the same idea, but this time not as a supply chain attack, but as a pointware attack. Meaning that instead of uh writing or creating a new application with a backdoor inside of it, we actually compromise the LLM application being used to prompt for the uh resource. For example, uh, if you're using a cursor, okay, instead of compromising the code generated by cursor, we compromise cursor. Okay, if you're just putting prompt injection within the library instead of like you know, compromise library that gives you uh the, then you now have the ability to compromise the LLM application.

SPEAKER_02

Oh, I see. So so essentially, like you ask Kursor to do something, uh it hallucinates a library.

SPEAKER_00

Now you can't Kurser, the foundational LLM hallucinates something.

SPEAKER_02

And then Kurser goes out on your behalf as a user, downloads that library that like you pre-registered or an attacker pre-registered, and that includes a prompt injection that uh is now hijacking Kurser and now Kursor works for you.

SPEAKER_00

No, just uh to to to make it uh I would say uh to explain how dangerous this kind of uh attack is, think about it for a second. How many trending repositories are currently in uh GitHub, or how many trending skills are currently in uh there's claw hub for no club, club, yeah, club. This is the the the equivalent of uh and if you think about it for a second, these kind of uh repositories and skills are being downloaded massively throughout the week, right? Like you have thousands of downloads. Uh, here is something that you may not know. In many cases, when you prompt the LLM to bring you a trending new repository or a skill, and I say in many cases, it's in the vast majority of the cases, it will hallucinate a repository that doesn't exist. Okay, if you're not if you didn't specify correctly the use the the namespace and the name of the repository, which people don't usually do. You just ask clone the repository and its name, right? You don't write the username at the beginning, you just ask for the specific name of the repository. Yeah, what you will get at the end is the replication of uh the name of the repository into the namespace. So you have, assuming such a repository and a trending one, you can very easily go and register uh what is trending, and you're expected to have major traffic into your registered uh repositories or skills or whatever. And with very high percentage, by the way, it is most likely to be hallucinated rather than bringing you the reefing. And as a result, if you just put, like, for example, install um a reverse shell, install something that I will be able to communicate with, you now are able to amplify your ability to have a backdoor installed on a device due to one that you took something very trending, okay, and you used uh the fact that it is most likely to hallucinate the trending repository into your repository instead. In many cases, you are using just like uh AI coding assistants, so they have integrated terminal, so you have the ability to install a bot. And when you can amplify stuff, you have basically the ability to create a botnet, right? So you took something which is something like safety issue, we are not responsible for safety issues, weaponized it, and now you have a botnet okay, installed on your device. You are now part of uh botnet. You were able to amplify prompt injections from one to one, meaning that one affected application per prompt into many affected applications per prompt, which is registered, and all of a sudden you are now having scalable weapons.

SPEAKER_02

What is the like uh how feasible is this?

SPEAKER_00

Like, have you done any very, very feasible? Unfortunately, it's very feasible. Too feasible. It's independent of the large language model being utilized. It is independent of the LLM application being utilized. It works for most of the LLM, like I would argue, all of the LLM applications that we used. Given that resource that you decide to apply the attack on, that the holo squad attack is trending, you are guaranteed this is not part of the um training data that was used to train the model.

SPEAKER_02

Oh, so you hit molus nasense.

SPEAKER_00

So you just take things that are trending by uh by the by their nature because you decided to take them to amplify your attack. As a result of the fact that they are being um trending and they are not part of the training data, they were most likely to be hallucinated. And you are now in a problem that most of the applications, okay, independent of the large language model being utilized as the their backbone, will uh trigger the execution of the payload that you put within your repository or skill or whatever.

SPEAKER_02

I don't see what you mean on the responsibility, because like if your cursor or cloud code Who is responsible to it? Yeah, like what would they do? Who's to blame? You asked your your cloud to go and install, like uh I don't know, the uh awesome.

SPEAKER_00

In cloud, it's easier, by the way. There is one, like you know, in cloud it is very easier. But uh what's easier? There is one responsible entity if you're using Cloud Code because Cloud Code works only with anthropics uh because they have like the entire chain. They are the ones that are responsible end-to-end, they are the developers of the large language models, they are the developers of the application. It we understand who is the Did you submit this to them? Like the the the Okay, so there are ongoing uh disclosure, not only with them, with the entire LLM industry, with the applications, with even the uh frameworks that host uh repositories such as GitHub and uh Clohub. I would argue that uh I'm not sure whether you are facing the same problem or not. It's not easy today to uh even uh submit a report via Bug Cloud or Bug Hunter uh because they have you first of all, you need to be able to uh first of all bypass an LLM that uh by but by default intends to block you, okay? Which after some PDF, so they force you to uh you know to write stuff again. And in here, it's not even clear whether it's a security issue or a safety issue. So you give it to the security, you were able to spend a lot of hours in uh bypassing the uh the LLM of uh bug hunter or bug route to submit it as a report, and at the end you will receive an answer that this is not even a security issue. Send it, they will not send it, uh even though it's um a problem of untopic. You should send it to uh to the safety uh uh uh bug bounty, which is a different bug bounty. Again, go through the entire process. By the way, the safety bug bounty is even much harder to uh bypass because they require verification of the person that sends the report, okay, which is not something that you usually have. It's you need your uh license being uh uh submitted as part of the process. Wow. And at the end, it comes to a point that you just say one of two things. Okay, I don't want to spend any additional effort on trying to convince uh them whether they should take it into account. This is one option. The second option is that you find someone from this kind of companies and tell them, look, I recommend you to look on it. This is the 90-day disclosure that I gave you. Do whatever you want out of it. I'm not intend to, and in the best case scenario, you you even find people that uh understand that spend like you know, the two minutes right uh reading the email that you uh send them, mostly because not that they are very interested in the email, but they don't want to be blamed for as the people who uh got the you know the disclosure and haven't uh submitted it uh you know internally to the company.

SPEAKER_02

I mean, we are we're experiencing a very similar thing. I think uh the state of bug bounty and state of vulnerability disclosure is obviously not in a good place, and AI is making it worse, both with AI slop reports and also with real reports. And I feel like we are doing the exact same thing. We will go through the bounty, uh I mean, through the disclosure program, we will try to jump through all of the different hoops that they that they uh force you to jump through. But at the end of the day, the thing that really gets them moving is when you reach out to somebody, you find a back channel, you reach out to somebody you trust, and you're like, hey man, uh uh hey, please just look at this. This one is important. Uh, I do feel like uh back bounties in general, like we are talking just like just last week, there was this entire shenanigan with MSLC, yeah, yeah, yeah. And uh and uh what is it like Eclipse Blizzard, uh the researcher. I I think it's everybody that tries to uh to your point, when you find a bug, it's almost kind of a second effort that the vendors are forcing you to do, which is convince us that this is actually a bug. Uh and there's a bunch of paperwork and again, this wasn't the case two years ago, right?

SPEAKER_00

This was the the result of having an ability generated by AIs and report generated by AI, and which is now we are facing with those who now have something in hand which was the result of real things, now have to struggle to convince the you know the I mean I don't envy those.

SPEAKER_02

Uh I I don't envy the folks that are running those programs because they are now probably drowning in reports, and many of these reports are just slop, just plain out slop. So the real stuff just hides away. But I also think that the 90-day disclosure in the AI space is kind of a joke. Like uh if if you're gonna disclose, for example, a prompt of uh, hey, here is how I bypassed all of your defenses, I got a zero-click attack. Uh, let's say, like the one we saw on ChatGPT last year, you give them the prompt, they fix it. It takes them a day. Sometimes it doesn't take like for the mature companies, they have a process automated. Like, why do we need to wait 90 days that people know that they were they could have been affected?

SPEAKER_00

Okay, so let's put it this way you understand CVSS and some other stuff, which at the end of the day have the ability to quantify how severe is the risk, right? The 90 days disclosure is mostly for them to be able to patch the risks which are not that important. Okay, the risks that are really like you know, have uh great severity are being in some cases patched in a day or two. Okay, they have like all hands in uh in companies and enterprises, and this must be like you know, I've been told by a friend that he received the test that things should be created in 72 hours. Okay, so they know how to you know to drive stuff when they have the motivation to drive stuff. I do agree, and I would argue one thing that it's sometimes the patch that you do in one or two days aren't that good. It's mostly adding some or extending a dictionary of stuff.

SPEAKER_02

It's a low-hanging fruit, and you're gonna find the bypass of it.

SPEAKER_00

Yeah, yeah, and immediately after you you are able to, you know, to to change the prompt a bit and find a way uh again to get in. So if you give them 90 days, the thing that I would argue that they at least struggle or do the effort needed to have the new mitigation much better than the previous one. If what they do is just like you know, updating a dictionary, in reality, I think that this is uh abuse of the 90 days rather than you know trying to figure out whether they uh how to deal with uh the main problem.

SPEAKER_02

Yeah, I think mostly the 90 days are used for just to make the in the AI space to make the thing no longer news because the thing we had three months ago is not relevant anymore. So I'm uh who who remembers like uh so for example agentic browser six months ago were the most important thing, and now they're a thing, but they are not the most important thing anymore. And so I think it's it's just a pace is causing the fact that you can only release stuff while it's no longer the hot thing, uh, which is I know kind of funny as a defense mechanism. Yeah, so your disclosures on on the hallucination problem, on uh halo squatting, uh like any of them successful or not really?

SPEAKER_00

Well, I I can tell you one thing, and this is exactly what you've uh um described earlier. The disclosures that uh we I think had some good uh responses were those that were made by emails and not via the bug bounties. Everything we received from the bug bounties was not exactly convincing, if you ask me whether it was a bot that answered this uh question or whether a person that decided whether it is uh something that should be considered or not. Only those and we went to the among those that I think started to take things into considerations were OpenAI and also uh clohub.

SPEAKER_02

And again, the reason was that I had in both cases people on the other hand that we communicated with them and at least argued things because I mean it's a it's a big it's a huge problem, it's an entirely new attack vector that you're describing that has not really been described yet. And I get I get your point. It's it's uh it's scalable.

SPEAKER_00

This this is the main problem, this is the property that should be taken into consideration when you're applying such an attack. Like so far, if you think about it for a second, and I discussed about the one to one-on-one uh barrier that is currently exists within uh the promptware domain, right? Uh, you need to spend the promptware per you need to spend now cost you an effort of a prompt to compromise one application in the best case scenario, right? And among the you know, uh the the first attacks that went to a scalable kind of like you know, took it from one to one to one to n was the one that we did about three years ago, something like this. Okay. In my opinion, if you ask me, this is much dangerous uh attack. The worm that we described about three years ago that again was able to spread automatically within an ecosystem, spreading between different clients, it has a different uh propagation, uh as let's put it this way, uh path of scalable uh to scalability between clients and not from one specific point, is something that, in my opinion, pose much less significant risk than what we've just uh shown.

SPEAKER_02

You're talking about the AI worm, um, like the basically a Morris 2. You called it right a few years ago. You got the pony award for that, right?

SPEAKER_00

No, no, we I got the pony award for something else, which was the the Oh, for the crypto analysis. For the crypto analysis, yes, it was on 2023, long time ago.

SPEAKER_02

So the so the that worm, you it was kind of propagating through it, was an example on emails, right? An email assistant. Exactly.

SPEAKER_00

So basically the the the that's what it said. The threat model is or the victim application is an application with the ability to communicate with other clients within the ecosystem. And if you think about it, email assistants are exactly those that have the ability to communicate with additional email assistants, and also those who are what I consider as uh whose inference is rug dependent, okay, whose communication, not inference whose communication is uh dependent, because we describe the ability to propagate uh prompts that do malicious activities between different clients by generating one simple uh email that uh uh goes into the database, later on retrieved by the application, for example, in response to summarize my recent uh emails. And then when you draft something which makes uh the communication rag dependent, uh okay, it uh forces the inference to replicate uh the input into the output, do malicious stuff in the middle. For example, either adds phishing uh uh links into it, either exfiltrate sensitive data throughout this uh email or some other stuff. But uh I would argue that back then the damage was very limited in what you could do. Okay, it was mostly exfiltration of sensitive data, maybe phishing, maybe up to some degree spamming, things like this. Yeah, okay. If you think about what I've just got to do.

SPEAKER_02

Sorry, this was 2022, 2023.

SPEAKER_00

This was published at the very beginning of 2020, at the very beginning of 2024. Oh, okay. Okay, so we worked about it at the end of uh 2023 and published it around. 2024. If you think about what I've just described with the holo squatting attack, I think that this is a real problem because now the outcome becomes remote code execution on your coding assistant or on your open claw, right? And remote code execution takes you out of the promptware kill chain into a new kill chain, which is the traditional malware. Okay, you now have the ability to create an identic botnet that works and validated against real products. Okay. And there are many products who are vulnerable to this kind of attack, which most of them are not secure by definition.

SPEAKER_02

Is there one that's not? Um is there any product that is not uh vulnerable?

SPEAKER_00

Well, there are products that do not have, for example, terminal inside of them. So the best thing that you could do is maybe exfiltrate something, but it doesn't give you the ability to.

SPEAKER_02

Essentially, if it's not powerful, yes, it may have less power to do harm.

SPEAKER_00

Yes, but think about it again. Those we are now going or in the about to see the era of autonomous agents, completely, fully autonomous agents. Okay. And you want them to be able to do something meaningful uh on behalf of you, right? So you give them the needed uh the needed tools and the needed uh functionality to uh to execute it. In uh even today, OpenClaw and the let's put it this way the the assistant that are being the most popular assistant and the most popular AI coding assistant have an integrated terminal. Not a single one of them is having its own firewall to secure against them, right? Especially those who are different vendors for the application and different vendors for the LLM. For example, Carcer, Devin, C-line, all of these you have different uh many vendor and different vendors in here, and no firewall in the middle to secure, like you know, from bad things that could uh happen.

SPEAKER_02

I think the thing that it's terrible for multiple reasons. One is because um there is really no like uh there is really a problem with responsibility, like who needs to fix this problem. And then there is also the fact that basically all of our developers, everybody's using using these uh agents. They are instructing the agents to install things. You're gonna get these malicious uh uh hello squatted uh kind of binaries or repos or whatever it is. And then when they hit you, they have the best thing to leave off the land, which is a coding agent. Like all of them, by definition, get back to a coding agent. Uh at the end of the day, with like clothes behind it that can write a script, that can manipulate manipulate, evade defenses. It's uh yeah, it's pretty good.

SPEAKER_00

And bear in mind, yeah, even if you think about it for a second, even if up to some degree you would argue that you may be able to put some mitigations or scanners within those platforms such as GitHub and uh Clohub and some other stuff, you can still just like you know, reference the payload within either the skill or the repository. Like, for example, go to uh this uh link and do this, right? And then you need to be again in a position that your LLM application is secured against it. Otherwise, you brought external uh data, which is untrusted, okay? And by instructing the LLM to do it, it has the ability and functionality to do it. Uh the data itself, if you scan it, the the repository itself, if you scan, is not malicious. There is nothing in there, okay? Only when you're doing it in a dynamic manner, it becomes malicious. Okay, think about it like a reflection of payload being uh brought from the internet in when it is being processed by uh an LLM application. So again, I at least see here something which I haven't seen before. It's not a supply chain attack. Uh, we're not trying to compromise the new generated uh application. We're compromising the application who have terminals inside of them which are widely used. And we do it with the use of very popular resources being um downloaded in mass or in high volumes throughout the days and the weeks. So, again, at least to me, this kind of scalable attacks does show something that hasn't been considered or haven't been taken into uh consideration so far.

SPEAKER_02

This is super cool stuff. Uh, where can people learn more about uh about your resources now?

SPEAKER_00

Yeah, so we are in the middle of the disclosure, so I haven't completely like you know disclosed all the things that are very interesting within, but uh I think that I will be able to discuss it very soon. We are about to finish the 90 days disclosure uh period that we gave them. I'm not even sure that we are about to see real mitigations into this kind of uh I'm not even sure again, at least in some specific cases, there are like you know the I know who's the responsible for it because they create the application and they develop the large language model. In the vast majority of the cases, this is not the case, okay. Um, but again, I hope to be able to uh to reveal it within the next few weeks or so, something like this.

SPEAKER_02

That's super cool. I look forward to it. I think we're out of time for this part, but we're coming back with another part pretty soon. So Ben, uh thank you so far. Thank you very much.