In the Wild with Michael Bargury
In the Wild is about the AI conversations worth having right now. This podcast brings together builders and breakers, researchers and practitioners – people who see things differently and challenge what we think we know about where AI is going, and what it’s already doing in the wild. Each episode hosts one of these conversations. Sometimes deep into new research, sometimes grounded in a real security problem, sometimes just a paper that raises a question no one’s answered yet. You could say it’s non-deterministic. Just like AI.
A Zenity Labs podcast, hosted by Michael Bargury, Zenity Co-Founder & CTO. Produced by POLDHU.
ZenityLabs
https://labs.zenity.io/
@zenitysec - https://x.com/zenitysec
linkedin.com/company/zenitysec/
Michael Bargury
@mbrg0 - https://x.com/mbrg0
https://www.linkedin.com/in/michaelbargury/
In the Wild with Michael Bargury
The Bungee Cord Was the Safety Plan w/ Ads Dawson
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
He hacked a robot dog at Black Hat. On stage. With a leash as the safety plan.
In episode 2, Michael Bargury sits down with Ads Dawson, Ads breaks AI for a living and loves it. He's Staff AI Security Researcher at Dreadnode and Technical Founder for the OWASP LLM Applications Project, where he builds the tools that stress-test AI systems to their breaking point. As a senior red team operator with BASI team six (led by Pliny the Prompter), he's the guy running wild, large-scale attack simulations that other security teams later copy.
Michael and Ads discuss - finding the vulnerability is only half the job. The harder test is whether an agent can get there with the judgment, restraint and creativity of an expert. It can expose an API key, step outside its scope or leave a trail while still marking the task complete.
In the Wild is hosted by Zenity Co-Founder & CTO Michael Bargury.
A Zenity Labs podcast, produced by POLDHU
CHAPTERS
00:00 — Inside the Wild with Ads Dawson
05:40 — From jailbreaking to agentic security
08:28 — Can an AI hacking agent stay stealthy?
12:13 — Building a personal bug-bounty harness
24:28 — Why off-the-shelf tools aren’t enough
33:58 — Keeping autonomous agents inside the lines
43:56 — What black-box AI hides
46:15 — Hacking a robot dog live onstage
52:13 — ScopeJudge and a new way to enforce boundaries
64:54 — Is red teaming dead?
Research:
https://ganggreentempertatum.github.io/speaking/
Relevant Profiles:
Ads Dawson
https://ganggreentempertatum.github.io/speaking/
https://www.linkedin.com/in/adamdawson0/
Michael Bargury
https://www.linkedin.com/in/michaelbargury/
Zenity Labs
https://www.linkedin.com/showcase/zenity-sec-labs/
Mentioned:
Pliny the Liberator: https://x.com/elder_plinius?lang=en
Joey Melo: https://www.linkedin.com/in/mrjoeymelo/
Pedro Paniago: https://www.linkedin.com/in/pedropaniago/
Nahamsec
https://www.linkedin.com/in/nahamsec/
Mike Takahashi
https://www.linkedin.com/in/michaeltakahashi/
Tamir Ishay Sharbat
https://www.linkedin.com/in/tamir-ishay-sharbat-069496163/
You hacked a live robot on stage and got it to attack. Instead of like saying like charge at Michael, I would say like find the black shirt and run at it. When you're sat there in a black shirt and they're like 35 pounds and like made of metal and they run very fast. Agents don't think the way that like people think, and it will just run Python instead. So it's ultimately just bypassing your trust. The thing that makes humans different from AI and different from each other is intent and creative thinking.
SPEAKER_03Hi, my name is Michael Baguri, and this is In the Wild. I'm here with Edanson, and he's staff researcher at Dreadnought, working with Basti, pretty much founded the part of the LLM Top 10, top three hacker in Bag Bagoni. I think like just more and more. Thank you so much for joining me.
SPEAKER_01Thank you for thank you for having me here. It's a pleasure. And uh yeah.
SPEAKER_03Absolutely.
SPEAKER_01Uh so uh you recently moved to the US? I did, I did. Uh I live in the US, so I moved from northern Ontario to Florida. They're very much different. How's that move? I have two kids, so you can yeah, you yeah, you know. It was easier when me and my wife moved to Canada. Uh the two of us. Uh, but no, it was very, very good, thank you. Yeah, uh enjoying life. Did you find the hacking community there? I kind of know a few people who were in Florida already, uh or like in in the state, but not I I live quite far out uh the city. Uh it's a bit more quiet, you know. Just it's basically just me, my family, and the alligators, which I actually kind of like. That's Rhoda. Yeah, seen some a lot of alligators, which is very cool. At least it's warmer. Oh, yeah, so much warmer. Like, yeah, it's I I love it though, because I'm from England and England is not warm. I will take like every bit of sunshine I can get for now. How was your week in Vegas? Good, thank you. Very good. Yeah, I uh I just like hanging out with friends, uh, which is like my favorite thing. Just seeing some of my really cool talks. Vegas is Vegas. Yeah, yeah. My social battery is like very at its peak. It's a difficult place. Definitely, yeah. What did you see? What talks did you see at the leg? So I went to go and see just mainly my friends' talks, Joey, Mello, um, and and and Squid. Um, they did their talk earlier, which was very good. Um, my other friend Pedro. So I just go see all my friends' talks and then hope no one comes to see mine. I heard great things about your talk. Thank you. I had people were scared for you. Oh, yeah, because yeah, we had we we had we had a robot, um a little robot dog, which was which was good fun. But you also did very well as well.
SPEAKER_03Yes, we didn't hack anything that could uh bite us back, so thanks.
SPEAKER_01I mean the vendors might yeah, they they always get a bit a bit funny. Um, but no, it's good, thank you. And congratulations to you because you guys did some really cool work. Thank you so much.
SPEAKER_03We have uh like our team is just it's been it's been phenomenal together. Yeah.
SPEAKER_01Well, I get to so tomorrow me and Tac, Mike, are doing our talk. So Yeah, what are you talking about? So we did a talk last year on like and it was more like jailbreaking and stuff, and then basically this year we did like 300 it's meant to be like 365 days for like the past year of X fills, so like stealing data from AI agents, and like the whole topic, I guess, is like how like data exfil from agents is like the new XSS kind of thing. I am lucky enough to have uh hacked with Mike and he's he's a very talented guy, so yeah, we did some we did some cool stuff together, so just sharing it. It's really phenomenal.
SPEAKER_03Like um he just did the uh ChatGPT hack. Yes, he did, yeah. The one click to to get an insider agent in ChatGPT to then exfiltrate all your data, leave on a schedule, like yes, he was very good.
SPEAKER_01So if no one's actually seen that, there's a two-part blog, right? Yeah, both are very good. Yeah, yeah, that's awesome.
SPEAKER_03The calibre that he brings in.
SPEAKER_01Oh yeah, he's a he's a force.
SPEAKER_03Yeah.
SPEAKER_01I should I should I'll I'll show you a photo later of we have like um we used uh Gemini to make like uh boy band pictures of me and Mike. Like I'll show you later. I wanna see that. Very cool.
SPEAKER_03How how are you thinking about so tell me tell me a bit a bit about BT6? Like how does it work?
SPEAKER_01What's your focus? BT6 was uh founded by Pliny. Um for those who don't know, Pliny Liberator. So he's like most famous for his jailbreaking and stuff like that. Um and basically, yeah, we kind of uh maybe about a year and a half ago, which was uh Mike was included, and it was just like a random bunch of people hanging out on Discord. And uh yeah, and then it it kind of became a thing. Like the team did like some really cool work. Um, and yeah, we've been doing like AI testing and AI security for uh for a lot of firms now, and it's just great because you get to work with like awesome people who are like extremely talented and just learn off everyone. So it's like just a collective of hackers and jailbreakers and stuff, and it's like a very uh suicide squad kind of group, but it ultimately works pretty well. Yeah, it's it's good, it's good.
SPEAKER_03What's your mental model for the difference between like that's this started with jailbreaking, right? Getting the models, like going out around the guardles, but now with agentix systems, from our perspective, it's so much different because you have that combination of software and the models and the different models. How's your mental model changed?
SPEAKER_01Yeah, it's a good question. So I think um the the jailbreaking stuff is like almost like good to know because you almost know how to control intent, like because like effectively, obviously you've got input, trust untrusted input and then unsanitized output, like exactly the same as web as web stuff. Um but for me, I was uh my background is like web app hacking, so ultimately, I guess I'm thinking about it as just like another layer in the stack of the technology, and just looking at like the sources and syncs and how I can like the pivot points, you know, like the attack model, like it can you come in on unauthenticated? How do you interact with an agent and how does it interact with the rest of the ecosystem? And you kind of build the threat model off that, right? It's at least I guess how I do it.
SPEAKER_03How would you categorize vulnerabilities in that? Part of the challenge that that that we have uh with with models is that I always say, like, uh, for an app to be vulnerable, it needs to have vulnerability, it needs to be exploitable, somebody needs to find it. Agents are vulnerable, just point blank. That's it. So, how how are we thinking about progress? Like, when is enough red teaming enough? Never enough red team.
SPEAKER_01It's so much fun. Um, uh, it's a very good question. Ultimately, like you well, you know uh that like ever all these agents are integrating with everything, right? So the attack surface is just getting more and more and more and more and more and more. The labs and you know, products and stuff are like it's embedding AI into every aspect of people's lives, and ultimately, from like an attacker's perspective, all those elements carry different primitives. Um, so yeah, definitely like it's just becoming more of a thing, and you know, the race of the industry is so vast that ultimately there's so many little things that get lost, or also new protocols as well, right? Like, there's so many new things that they don't things like OWASP has been going for like I don't know how long, whereas like MCP was adopted like within probably like an hour of this bank being ridden, you know. So there's there's like definitely a gap there, and the industry is still maturing. Definitely a good time to take advantage of that.
SPEAKER_03We bring in agents to an old world, and that world had some assumptions in which the security boundaries made sense. Yes. But then the agents come in and they break those assumptions, and we don't realize that we don't like nobody notices until you saw you show them like a visceral example.
SPEAKER_01Yes, yeah, exactly. Like um, so one thing uh which I will shamelessly plug, if that's okay. Uh so one example I had, so like using agents for offensive security, so like bug bounty, for example, and my friend and I we did a recently the benchmark called Stealth Bench, which Soda. Yeah, thank you very much. Offensively, what it was it was so when you're building a benchmark, you have a suite of tasks, and then you run the evaluation and you measure the agent's ability to do the thing. The tasks in this instance were all made from real-world scenarios. Those real-world scenarios were like upsies of agents hacking in a not very cool way. In your setup, yeah. So, like not like um not like they're going hacking other companies or whatever. So, do you need to make an uh like a post and we can all like say how your how good your model is? Yeah, just like a very simple example, right? So let's say like an agent finds an API key in a JavaScript file, which is a bad bad, and then the API key allows some other kind of access to do something else, like a crud operation on a other endpoint. The agent in like the experience would like take the API key and use that as the evidence for the right, which is not good because it's basically taking sensitive data and putting it elsewhere on like the internet, and that's not cool. Uh and then when I saw that, obviously I was like concerned from you know, like a like obviously, but ultimately it was really interesting to me of like measuring how stealthy agents are because like I uh have I mean you as well are like offensively minded, and if we want to use agents for offensive security, we almost want to trust that element of stealth, which because uh stealth is such an important thing in like red teaming and any offensive security, and it was really cool to measure that. And ultimately, what I really wanted out of it is that open source models were the same, or if not better, than the frontier models. Really? Because there's a lot of hype about people not trusting, you know. I won't go into that, but like you know, like the trust of like the frontier models, like almost like a comfort blanket kind of thing, and there's a lot of research of stuff like this, which actually shows that that may not be the case as well, like to to challenge assumptions.
SPEAKER_03What were the results? So the best models, like what what how did they achieve them in the benchmark? Yes, but all the time.
SPEAKER_01How good is best? Uh I think it was like 48% from the moment. So, like, still not still not great, but like GLM was like 47 something. So, yeah, like ultimately the the gap was really, really minimal. Where was uh opening eye stuff? Like, where was Sol? Not good? No, which is crazy because Sol is like very to the book, right? Um but then it's just it was really interesting because we published the data set, so like all the traces and everything are available.
SPEAKER_03Ultimately, you would you would think that uh I mean I would think that they would train against your benchmark because like if you look at the hug in face incident, the open AI hugging face incident, yeah, the only uh silver lining is basically if you would monitor those traces, you would see it like pop up immediately, right? The AI can hack, but it hacks, but it will pop up any alert in your yeah, whatever detection engineering you've got, you will catch it, right? Yes, but then your benchmark is what to me is really scary.
SPEAKER_01Yeah, thank you. Yeah, no, it's cool. So um, yeah, I just hope it is insightful and it was it was like a cool experiment to do.
SPEAKER_03You're doing a lot of bounties, obviously, right? Can you tell me a bit about your setup? Uh good.
SPEAKER_01You don't have to. So no, so uh I use um a fairly complicated harness, but I use uh small and open source models for like very distinct tasks, uh, which is like a you know, like a token maxing and cost saving budget. Ultimately, I've just spent a lot of time distilling my own like domain expertise into my work to ultimately kind of like augment myself really. But it's been fairly successful so far. Taught me a lot as well. Like, you know, being able to one of the things I love, like you know, like I'm sure you do too, is like being able to experiment and like try these different things, and you know, ultimately you never had that like flexibility or that like freedom or independence to like just try and test things, which is really nice. You you know, you feel like a builder and something like you're like proud of, and it's like you're you know, you're like your baby, right?
SPEAKER_03When you like distinct your knowledge, for me that mostly translates to building knowledge systems. Like I find that I would figure out some way to kind of structure the knowledge the right way, yes, and then as I engage with my harness, it continuously just captures more and more of my thoughts. Yeah, but then at the end of the day, it's like all of a sudden, like everything I had in my head is now on paper. I mean not on paper, but in Git. And then I can send it to somebody else and say, like, hey, don't talk to me, just talk to my harness.
SPEAKER_01Yeah, exactly. That that's such a such a good thought. Skills were one of the things that really kind of changed it for me because it's such an easy way to represent data in like a unanimous format, like it's so random with markdown. Obviously, like it's probably not the best step forward going further. Like, you know, we need like more structured data objects. Anything contextual, like you said, like I know Michael likes double espresso. You know, like like stuff like that. It's uh it's a great way to just like represent unstructured data, which is yeah, very cool.
SPEAKER_03This is actually something that Fede Magi taught me. He basically had a system of uh that he's using for like everyday life. Yeah, kind of manages health, manage all of those like tracking stuff stuff, and he was pointing to a specific knowledge management system which I had no idea about. Oh and so I was looking into it, and it's kind of like it comes from uh how do you uh create structure around your digital life? Yeah, it's just like how do you run uh basically a library of your life, yeah, yeah. And then I started by kind of learning more about it, and then I realized that I can just point Codecs or point Claude and it builds it. And then I like every new project I start like that. So now we had uh on uh like last week we identified uh kind of a big big model campaign propagating through skills, yeah. And I was doing the threat hunting with codecs. I mean I tried with Claude, but it basically ran out, like ran out the door like three three miles away. Yeah, so I continued with Codex. I did the entire investigation with this building knowledge base, and then I had to talk with Tamir and kind of exchange notes, and I and I was all in on like black cat conversations and I had no time. So he just talked to my knowledge base and he figured out everything. It it was awesome.
SPEAKER_01We should shout out Tamir because Tamir is awesome. Umir is incredible, yeah. He's he's very incredible. I've I saw a lot, I've never tested it personally, but I saw like um Daniel Miser has like a Life OS thing as well. Like, um, yeah, I think these are gonna be like like the new wave of things. There's a couple of people I know, and when like you exchange emails, you can tell. Because he's like talking to GPT's like, hello Michael, yes, where but I am free on Wednesday at 6 15 p.m. I have this I have this joke with my wife that um I call my wife Claude sometimes. But like even even uh non-technical people, like I love the Fat AI, like my wife uses it for like kids stuff or you know, like arts and crafts with kids, like just random things. It's really cool to see that, like you know, like my friends are my friends who are like non-technical have like their own fitness app apps and stuff, right? Like it it's it's such a cool enabler, like it's like the most incredible tool, I think.
SPEAKER_03Yeah, I agree. When when when my wife and I just uh started living together, I worked for like I don't know two weeks, three weeks, something like that, and I built us a mobile app to track um what was it, like expenses or something.
SPEAKER_00Yeah, yeah.
SPEAKER_03And like I was so proud that I got this thing to work, and of course nobody used this because it's cool though, right?
SPEAKER_01Yeah, now you can just one prompt away. I actually thought about this um interesting concept. I saw it and I can't remember who wrote it, but um you because like normally we have like REST APIs or like GraphQL. Like imagine like a single API now where you just have like a single API endpoint, which is just like uh prompts or whatever, and ultimately like it's kind of like a very interesting thought about if it will go that way.
SPEAKER_03It's so weird to think, but I mean I think that right now most people will take an API, put an MCP in front, and then they will say, Okay, now become an expert at how to use my MCP, and you try to basically prompt inject the agent that's consuming the MCP to learn about your API. Yeah, but you could absolutely say, Hey, look, let's do an MCP that's one call. Yeah, basically says, like, why like I can build an agent that's an expert in my API, you don't need to worry about it.
SPEAKER_01Yeah, exactly. That makes a lot of sense. And like, um, there's also the thing as well now, like a lot of people don't do docs site anymore now, right? You just have like uh lms.txt and it's just like uh apart from like blog posts and stuff, I don't think like a lot of people really read docs anymore. Yeah, I agree. There's maybe the there's maybe like a bad thing about that, about like reading Vin stuff, but um yeah, it's very interesting.
SPEAKER_03When you do like bug bug bounty with models, do you have a mental model on on kind of what are the different failure modes for the different models?
SPEAKER_01I spent like a lot of my experience um doing like cyber evals, so like I kind of I'm very lucky that I've seen that like failure and success kind of along the way, and then I ultimately just turn it into like an engineering problem because that's like what it is, really, and find ways to like make the agents like a first-class citizen and like ultimately be the engineer for the agent, if that makes sense, like knock down walls for it, and ultimately that that point I saw like the capabilities actually go higher and it's like actually then improving, and that might be through you know a new tool, or like I know uh a lot of my friends have built like their own Chrome debugger where they can like insert hooks and stuff, you know. If you want to get like an agent that's testing like cross-site scripting and stuff, ultimately just being that enabler and like looking at the data and seeing where it's failing, and then you just like reverse it and you're like, oh, that's because XYZ, and then because like all the stuff we said about being able to build now, you can just be like, Okay, well, I just need to like build this and then replace it, and then you've like then that's like one volume. And for me, that started like AI was always really good at server side, and client side was harder because it's like you know, there's a lot more going on, but ultimately I think now it's actually I don't really see that the case. Um, like I I I looked at some of my stats from Hacker One for this year, it was this was like six weeks ago, but I think I'd submitted against like 82 CWE classes, which is like just to show the the span of different types of bug classes, you know, there's like cryptographic failures, idols, XSS, CSS injection, like uh there's such an array which I think kind of shows that like you know, like when I started out, it was a lot more the server side, whereas now it's once you become more of a builder and enabler, it you then just increase yours.
SPEAKER_03How how does it work? Like what what's your process? Like you will send your harness to do something that it's not capable to do yet, and then analyze traces to build the right thing.
SPEAKER_01Yeah, so like um through my experience of where I work at Dreadnought, I um do like a lot of cyber evals, so it's like basically being able to measure a model of how good it is at like offensive security in a domain, and ultimately, like what you're trying to deliver is a better harness each time, which is like demonstrating the if the model's capability is there, then theoretically enhancing the harness, the needle should go up right because you're doing that. So it's like creating the environment, running it through, seeing where it failed. Was it like a you know, model failure, or was it like a failure in the engineering? And then when you find out it's a problem in the engineering, just then patch that, and it's that's just like a very iterative process. Um, it sounds it sounds painful, but it's actually very enlightening because it's ultimately just like it's almost like being a developer in a way, like you know, like you're just like bug squatting.
SPEAKER_03Do you have evoz for your own harness? For your bug bounty harness harness? Or is this fully in your head?
SPEAKER_01I think uh yeah, so I uh started with like started with um really simple stuff, like like let's say like Ju Shop is like is like V0001. And then once you get confident enough and you're like you know sensible and all that kind of stuff, like and like you just let it out into the real world and just like learn and adopt from failures and obvious that's as well like the the handheld operator kind of thing, you know, like being like guiding with it, and you're like looking at the you're looking at the data and seeing what's in there and you're like long horizon.
SPEAKER_03Task or do you and do you run many in parallel?
SPEAKER_01Or is it like one for like yeah, so the work I've done, like uh my job, we like focus on like long span horizon tasks. A couple of the guys I worked with uh did a really impressive thing a while ago where they took like a a very small Gemma model and did like SFT on tool calls and got it to do like game of active directory, which is like the full end-to-end like three added directory for us, like really big. Um though the problem is those environments take they take longer to build and stuff. I think we're like way past the point of like an agent being able to like achieve uh XSS in like a thing, like as a task. Like it needs gadget, gadget founder vulnerability to like get like a full account takeover, like it needs all those elements because otherwise we're not like we're not really measuring as well, if you know what I mean. It's like an inaccurate and it's not a good thing because then the benchmarks show something which is not really true, like they're actually more capable than we're almost giving them credit for, if that makes sense.
SPEAKER_03Yeah, yeah, yeah, yeah. I get what you mean. When I try to send an agent to do something interesting and then to figure out what part of that interesting thing is something I want to actually explore. So the the distance between doing that, letting it roll for like 10 minutes, or sending a hundred uh long long horizon tasks and then evaluating is how much time, as you mentioned, you you kind of build how much time you spend building the environment, right? Building the validators, building uh evals, making sure that like because if you run things for two days and then you need to look at the traces of a hundred agents, good luck. That's just just just too much.
SPEAKER_01Just like very rigorous har harness engineering, which I guess just um comes from experience of of like what you did wrong the last time or what was difficult for you last time, and then just making ways of like improving that and just being like really embedded into like what you're building and almost like you're you're the agent, and then you're thinking about like all these different elements until like you're ready. Uh one thing you can do is if you're doing like a long horizon task to do something really long, but like ultimately you want to test these little components, then you can create like maybe sub-tasks, and then once you've like, you know, like you've nailed this, and then you've done this and you've done this and you've done this and you've done this, and it's like, okay, now I'm confident that I've at least done this bit. So I'm like, okay, now like I will test it.
SPEAKER_03Why build your own harness and not attach to a Claude or a codex?
SPEAKER_01Part of where I work is like we've been building our own, so naturally been doing that. Um, but I feel like it's maybe the only way of gaining like true introspection into like what's going on and stuff as well, right? Because um, even though some of this stuff is like open source, it's like you're not gonna read like 6,000 lines of code, especially like now. So I really like to like dig into the data to see what's going on because I need that to feel good about what I'm doing. Ultimately, that's the only way you're gonna improve, right? Because otherwise you're just like depending on something else, maybe. And I think ultimately, um unfortunately the subsidization of tokens and stuff will at some point may like go, and that's not gonna be a good time. So definitely like trying to prepare for that as well is like very is like a good idea, and buying some GPUs because like who knows for how long you can buy them, yeah, and also like they go up in price. My my friend actually um my my friend bought some and they'd already gone up in price before they've been delivered to his house.
SPEAKER_03Yeah, I think I feel like uh like for years the best engineers have always been the ones that are crafting their own tools, right? They're crafting the tools of the trade, and to me it always felt like it's part getting your environment set up for you and part just training. Yeah, just like because you train on a small task which you know, and so this is really new for security researchers. Like security researchers have always been building tools, yeah. But the tools were much smaller than what you can build now, and this is kind of a special moment where the harnesses are still something that like you spend a weekend, yeah, you get a harness, yeah, exactly.
SPEAKER_01And it's like um it's it's true though, because like even you, like you said you built that app because like you want it a certain way, which is like a good thing because like you're so involved in it because you care about it, like ultimately, like it's worth that extra mile because and then you know you know as well, right? Like when you build something, and then you're like, wouldn't it be cool if I added this? And then you added this, and you just keep going, yeah, and then like two weeks later you've like built this thing, and you're like, Oh my god, like yeah.
SPEAKER_03I do think one challenge that like that I'm experiencing, I'm not sure if you find the same way, but because it's still kind of finicky, like it's very jagged, some things work really well and others don't, and also your harness, like some parts of it are great, other parts I'm sure like mine is like stretched with duct tape. And so if somebody else tries to now walk to together with me on that rather than just look at it, it becomes like 10x more difficult. It's clear that one person can do like 10 people's job right now, but it's not clear that five people can do 50 people's jobs.
SPEAKER_01I love the fact you said that because I know that a lot of companies are still trying to find that. That's really interesting because ultimately, like, you know, like just companies in general, right? We're gonna see like a different way of working.
SPEAKER_00Yeah.
SPEAKER_01Um, and I don't really know what the answer to that is yet. One person I know, uh Dan Guido from Trailer Bits, if you're aware of him. Like, yeah, Dan is like he did the talk as well with you guys last year. Well, you guys both did talks unprompted last year. Yeah, but he did like a really talk was awesome. Yeah, it was incredible. Like the way he like structured that, and like they found their like secret recipe, and I feel like a lot of companies will just end up doing the same. He's a great example of that.
SPEAKER_03I feel like that is the thing that Claude Code does the best, which is just it's not really the harness, it's more how do developers collaborate on top of the harness.
SPEAKER_01Yes, yeah.
SPEAKER_03So the hook system, the skill system, the plug-in system, now the inference hooks. It's not about the harness, it's about how humans collaborate around the harness.
SPEAKER_01Yes, exactly, 100%. Um, yeah, that's really good. I I like that.
SPEAKER_03If only they allow us to do cybersecurity, please.
SPEAKER_01Yes, that would that would be nice. Please please get me on the list. Um, yeah, no, it's um that's also a big can of worms and very interesting of what's gonna happen there. Um which I guess is the you know benefit of having GPU, maybe. I don't know. We'll see.
SPEAKER_03How much can you stock though? I mean you're gonna stock something that might the the but that's gonna be like in a year, and video will come out with something that's like 10x.
SPEAKER_01That was the really difficult thing of me of like buying something. Yeah. But then I think I saw confidence when my friend bought it and his like went up in price, and I was like, okay, it's not that bad. Like same with like experimenting with like vibe coding, right? Like you ultimately just have to like go into the uncomfortable zone, yeah, and then because otherwise you never know. And then you could be worse off from it. It's it's like a gamble, but yeah, I guess you've gotta take it.
SPEAKER_03You're playing around with the harness. Yes. Now you start playing around with model inference. Yes. So it's more about building your intuition.
SPEAKER_01Being able to experiment is so cool, right? Like you you know, you can build a finance app or you can like do whatever it's like. I did a uh panel uh this morning, and um one of the questions was like, how do people get into like red teaming still? It's obviously harder, like everyone's a a lot on a higher like trajectory now, right? But ultimately, there's actually been no better time because like no better time, yeah, because you can learn anything. Like you can learn literally anything.
SPEAKER_03Yeah, you have a private tutorial for everything. Yeah, and also I think like for builders for people that wanna create something, like the the number one thing that makes it difficult is that people are stagnant or they don't want change or they're not willing to even accept something else, and right now the willingness to change is just yeah, like it's never been like this.
SPEAKER_01Yeah, and you can really tell there's a lot of companies who are of that mindset and they're like constantly shipping. Like Shopify is like a really good example of that, like because you because at face value you would think that they would be suffered from AI, but they've actually like gone the other way, they're like doing all kinds of crazy stuff and like embracing it.
SPEAKER_03Once you start really using it, uh the the hype is making it difficult, right? Because you wanna object to the hype, but then once you start really embracing it, what you end up realizing is that for a company, you are a different company now. Yes, your product is something different, your job is something different, and the the earlier you realize that, the better you have a shot to survive. Because any company that starts now is obviously starts more reignated.
SPEAKER_01Yeah, it's the same with like being a hacker as well, right? Like embracing AI and changing the way that you work and operate to complement that is like well, that's the only way you're gonna make yourself better. Like you're never gonna make yourself better by like like you said, like not embracing it, right?
SPEAKER_03Yeah, it's part of our uh hacker culture to not buy the market in bullshit. Yeah, and so you kind of run away from it, but just playing around with it. So uh uh I think OpenAI and Dave Vitel's team open sourced the uh Outbox system last week. Oh, really? I didn't see that. Yes, they just they just dropped it, like no big announcement, nothing. And I'm like all up into that because it's so it's just so interesting to see the choices they made.
SPEAKER_01Yeah, that's really good. And also as well, um Cursor had open source their training library as well, which is like I think it's like 2.6 times faster training or something as well, and like it's crazy, like that is in a good way. Like just people are just like, yeah, like let's put it on GitHub, and I'm like, okay, you do you, and I'm gonna take that. I'm gonna steal all those ideas. Thank you so much. Oh yeah, it's it's a good thing for uh humanity as well, right? Because ultimately we're only all gonna get better if all sharing knowledge because we all have like this incredible capability now, but the thing that makes humans different from AI and different from each other is intent and creative thinking. So like me and you share each share stuff, like I don't think of stuff you do, and vice versa, right? So that is also really cool.
SPEAKER_03That's very that's a very interesting point.
SPEAKER_01Yeah, because like everyone's like it's like the same with um all the really talented guys on your team, they're like, you know, they all have like if you look at all the black hat talks like they did, like they're all slight they're all different. Yes. And like they all came from an idea that came from one of their heads.
SPEAKER_03Like AI can be basically anything you prompted to be.
SPEAKER_00Yeah.
SPEAKER_03But you can't. You're just like you're just you. Yeah. And that restriction is actually what makes what makes the things that you create, your perspective so unique, different, interesting.
SPEAKER_01Yeah. There's a really old school movie that was we had like Robin Williams in it. It was like a kid's movie, and there was like this, it was called Flubber. And it was like this green goo.
SPEAKER_03Yeah, we were talking about bug bounty uh earlier, and like when you do bug bounty, you point your harness at a target. I imagine the number one thing that you need to figure out from the very early days is how do you scope it, right? How do you make sure it doesn't go out? This is uh a particularly timely topic. Right?
SPEAKER_00Yes.
SPEAKER_01What are your thoughts on that? Like, how do you do it? Very good question. Um, so I will shout out uh Shane, who I work with. Um, he did an excellent paper on this and does some fantastic research. Um, but I think ultimately it comes down to the tried and the testing, like in the safe space before you get to that point you're ready, and good like harness engineering, you know, you said like hooks and things like that, um, and like finding these like heuristics and pieces and controlling those boundaries because ultimately the only way of knowing what that is is trying out the failures in the safe space, yeah and like iterating over that. And um the one thing that I've seen which is uh people try, which isn't uh probably not a good idea, is that they they box it. So like let's say they don't allow like a like a post request or something from like curl, but the agents agents don't think the way that like people think, and it will just run Python instead. So it's ultimately just bypassing your trust, like just transparency through it as well is also like that that just goes a long way as well, right? Because otherwise they're actually it's actually worse that they're actually then going like not behind well theoretically, they're unknowingly going behind your back as well, right? Which is even worse because you think it's good and it's actually not very good.
SPEAKER_03Yeah, I agree. I mean, one of the things that we've been experimenting with that's been working pretty well is that when you instead of just blocking, because you as you mentioned, it's gonna find another way, you block, but then you also give it like a reminder. Hey, yeah, you just that's not the right way. Go that way that go the other way or remember your scope. Yeah, that's just that's pretty powerful.
SPEAKER_01Yeah, I think about it as um, you know, like the ping-pong. Yeah, it's like the ball, and you're just like like just like gently like tapping it each time. Yeah, uh, it's just kind of like in my head is almost like the way I'm thinking about it. Because the SDKs and stuff are so good now, right? You can like just inject like at the tool call start or whatever, and you're like, no, no, no, no, no, not cool. Go this way. Yeah. And you just like push them, push them down.
SPEAKER_03Do you run it on your laptop or in uh uh I run it on cloud?
SPEAKER_01Yeah, I did at the start, and then when it comes to Defcon and stuff, I was like, this is not Yeah, yeah.
SPEAKER_03Because right now it's like the this weird time where people are walking around with their laptops open everywhere. Oh, yeah.
SPEAKER_01Have you seen you can buy the so people are buying those um it's like a it's like a claw. I thought it was a joke. The claw that opens the laptop. They sell them for like 70 bucks. Really? They're really expensive. Oh gosh. We should have we should have actually gone into a business and made those. We would have made serious money. That would have been a much better idea. A lot of like harness open source harnesses now, they all like run on Tmux and they all advertise like don't have to like keep your lid open because both people are just walking.
SPEAKER_03Yeah, it's just it's so funny. I was in a conference at the uh Real World I Security Conference in some in San Francisco the other day, and everybody was like, no, love not everybody, but many people, including myself, was like you walk around with your laptop, like, come on, and you talk with people and you're like, yeah, okay, for the loudest, everyone's talking to Claude.
SPEAKER_01He's like the most popular guy ever. He is talking to everyone.
SPEAKER_03Yes, he absolutely is. Yeah, must be must be exhausting. Yeah, his social battery must be worse than mine. Yeah. How do you like you're doing a lot, right? The OWASP war, the BT6 work. I mean, dreadnought looks looks like really awesome work.
SPEAKER_01Thank you very much. Yeah, um ultimately I guess it's just passion for it. Um but I I'm also like really blessed. Like I I get to work with a lot of really cool people and a lot of very talented people. I love to like collaborate and stuff because it's only the only way that I learn is you know, any way I improve is like learning off other people and stuff. When I was in school, I didn't learn. I was like not a very good student. And then now I have this like love for learning. Um, so it just becomes a like a passion of like wanting to, you know, just like better improve. What do you think that is? Did they uh kind of force you to learn something specific? I was always that kid who wanted to like make everyone laugh, but then that's the guy that like doesn't pass his grades. I just don't think I really knew what I wanted to do, so I wasn't interested. And then um I actually did uh I did uh an internship after and then I started doing networking, and I really liked networking, and then like security came like a bit after because like and then I became the networking security guy, and then I just found security more interesting.
SPEAKER_03Whenever people kind of task me with doing something, then I don't want to do it.
SPEAKER_01Yeah, it's like uh yeah, I procrastinate so much, I'm just like I should really do that, but then when you're really passionate about something, it just magically there's more hours in a day.
SPEAKER_03Yeah, it's just I'm so much more willing to do it. Yeah, you're you're ready to do all of the kind of grant work. So you have like I think you have a very unique perspective because like BT6 is like the expert in like model hacking, right? That that was where where you started and now you're doing everything, but that's like the like clearly that's the golden standard, yeah. And then you have OWASP, which is completely different, it's the golden standard, but in something completely different, right? It's not the attacker perspective, it's a defensive perspective, structure, compliance, yeah, and then there's kind of dreadnought and cyber evas, which there are not many people that are working on that. What are the differences that you're seeing in perspectives between those groups?
SPEAKER_01My intro to like AI was um I actually ended up working at uh Cohere who were who are a foundation model provider, so like I got like a really good insight. Uh-huh. And at the time when I was working there, I built a security team. So like the OWASP thing was like a really good moment to build that. So I've done like a lot of like blue team because I started like from a networking perspective. It was like very different at the start, but um it's like slowly catching up. I'm very much of the thinking that like the offensive powers defensive, you know, like the loop of like like you guys, like you publish some really cool exploits, and now it's like for the time for the people to go fix it, and then you just iterate the same same cycle. It's really insightful. Like there are a lot of people think about a lot of different things, you know, like it's like deep fakes and the whole um EU AI commission stuff, and like there's so many different domains. It's like we've because AI embeds into everything, ultimately it's like like saying security all over again, right? Like it's you know what I mean. There's so many different domains. I love offense, like it's like I just find it the most fun. Um, but then it's also really nice to like close the close the loop or like the the whole flywheel and like like like the stuff you guys do it's energy, right? Like ultimately you do this stuff because it like it makes the defense better.
SPEAKER_03For for me, it's just the power of uh the power of a good example is just it's it's more than an example, it's the power of a good story, right? Yeah, and it's very human. Like you can feel something and it's uh amorphic and you don't know what it is. But when a good red team exercise gives you a concrete concrete story of how something went wrong, yeah, now it's actionable, you can get behind it, you can build intuition on how to fix, how to generalize it. Yeah, but that's that's super important.
SPEAKER_01Because ultimately the worst thing to do is be like, hey, we broke your stuff, and then like leave it. You need like a like yeah, like you said, like actionable, deliverable, like something you can um it's like the same in bug bounty. So um one of the slides that me and Mike have tomorrow is um is like right at the end, and it's like this is all cool and whatever, but ultimately, like if you want to earn money from bug bounty doing this, like you need to deliver value to the program of like a way that they can actually fix it. Like when you're like doing an exploit, think about like how you would patch this, and then that is like the path to success. Because if you just do something like which maybe has no like meaning or like does that make sense? Like, you know what I mean? Like you're if something can be fixed uh deterministically through programming, like let's say it's like sanitizing tags or something, it's like that is like an actionable fix that they can do. Whereas if you give them something that isn't, then like ultimately it's like it's also like what did they set up to do when they built this program?
SPEAKER_03Like they built the program because they wanted to get stuff fixed, and so if you send them something that they don't understand how to fix, naturally they will kind of push it aside.
SPEAKER_01Yeah, exact exactly, yeah. Uh I guess it's like anything in life as well, though, isn't it? Like, you know, if someone's trying to sell you something and it sounds like you're not getting the full picture, I don't know about you, I'm like, uh like Yeah. I'm like, okay, but like someone's like the full end to end, like this is the full thing, and like there's no there's no like, you know, black mirror, like no smoke, then ultimately you just trust it more and it's I think that's what why open source has been so successful, because you can
SPEAKER_03Like it's not like you're gonna read everything, but you know you can, which is different. And it's also to me the most the weirdest thing about AI is that most of the AI we consume is completely a black box. I I don't mean the model. I mean like you have an API call, you don't know which model is gonna serve it, you don't know what kind of caching they do behind the scenes, like all of those shenanigans, you have no idea.
SPEAKER_01So many like layers of the stack of what's going on, and ultimately all you see, not picking on Claude, for example, um, but like all Claude Code Codex, all you see is like the three lines, right? But there's so much that goes on in the back of that, like um, like so much, yeah. So it's it's good to understand that. Um, so you can question the assumptions.
SPEAKER_03So we've been uh playing around in the last year or so with uh mechanistic adaptability. Very cool. And we brought in like deep AI researchers to do that, and then we pair them up with the red team to figure out hey, like this injection, why does it work, why does it not work? And the po like it's very finicky, but the power is so all of a sudden like you have a map. Like the thing works, and then you say, Oh, like so. For example, we um we did uh we did an injection on uh one of the Chinese models, and you can see that there is a like the injection goes through. So you hijack the model, you guys do what you want. But you can see the the feature in the model that's responsible for refusal, yeah. You can see it light up. So it knows, it knows what's going on. It's just that the other feature that's responsible to m for making the user happy is light lights up more. Yeah. So all of a sudden it's just it's no longer a black box. So just we are missing so much.
SPEAKER_01Yeah, exactly. Like there's so much. Um do you ever see that uh Carpathi where he did the tokenizer thing? Yeah, and you can like type in and then you see the tokens out. It's like the same thing, right? It's like, okay, so if I type this, then it's like this, it's like think about that on a on a scale. Um, yeah, like definitely like you just want to get as like deep into the weeds as possible so you like just know everything because otherwise, like yeah, you have no idea.
SPEAKER_03So we skipped right through your uh your your blackhead talk, but you had like live you hacked a live robot on stage and got it to attack, right? Well, yeah. Um I mean were you scared? What were the safety mechanisms around it? So we had a leash.
SPEAKER_01Was it like a metal leash? Uh it's just like uh like a bungee cord kind of thing. Um but yeah, it wasn't I mean it was uh it was not aimed at the manufacturer because there's like a bunch of stuff like this, but effectively the whole point was that we were trying to show that like uh so I one thing I actually really am interested in I have been working on recently is like embodied reasoning. So like agents like in the real world, which is through hardware, right? Like um, so one example was a Unitrue GoToPro, which is like uh it's like the robot dog. Um they're like widely adopted across like you know, police force and stuff like that now. But just to show like why it's important, it was like the meaning because when we n operate with agents through text, like we seem to come across all these like guardrails and stuff, um, or you break them. Um but like it in the in the embodied reasoning worlds, those same controls don't really exist right now, and like um it's an important time to recognize that because uh household robots are gonna be here very, very, very soon.
SPEAKER_03So practically like it's coming.
SPEAKER_01Yeah, it's like in San Fran now, right? And like it maybe give it another year and they're affordable and stuff, and um that's like the like an actual real like risk to human health and stuff, right? So yeah, we had the dog, and yeah, it was it was how did the injection work? Uh so the the device has like two main peripherals. So the first one is a microphone. So the model we had actually doesn't have an LM built in it, so we had the we had like a relay laptop which was just calling up to an API, so um, it had like a speech to text, but a microphone and an image. So it has like this like spinny thing around the head, which is effectively an image sensor, so it's just a labeler, it's like a classifier, and it'll just like label you know a black shirt or something like that, um, and just like a QR code, so um, which yeah, is is just a way of you know writing a prompt. So like you put in a QR code, it labels that, and then it kind of gets Yeah, so you you just give it like a roll, like it doesn't take much to like prompt inject it, like it doesn't even feel like you're prompt injecting it at all. Because ultimately you're just like swapping out words, so it's like like instead of like saying like charge at Michael, I would say like find the black shirt and run at it, which like doesn't you know I mean doesn't sound like like evil, but then when you're sat there in a black shirt and they're like 35 pounds and like made of metal and they run very fast. The robotics are really incredible. Like honestly, the um they do front flips, backflips, like the robotics part is in it's incredible like how advanced it is.
SPEAKER_03It's very similar to what we've been doing with browsers, like the kind of prompting. Yeah, it's very much saying it in a different language, yeah, making sure like hey, this is what the user actually wants, and and the fact that you can now have this like universal thing that impacts a browser and a whole bodied dog.
SPEAKER_01Because you know as well, right? Because you guys do a lot of it, that when it comes to like prompting and stuff, like it's not like you're doing like XSS where you're like doing loads of backslashes and escapes, it's more like you're just mixing control plane and data plane and like fuzzing that logic is ultimately like the most simple way. I never forget this. Um there's a a bug that I had with my friend uh drop, and I remember for days like trying to get this prompt to get it to do the thing, and then it was actually the most simple thing. It was just like act as a it was it was on like a bug battery on like an e-commerce thing. The prompt was just like to act like a customer, and I was just like I was trying to jailbreak it and stuff for ages. I was just like, oh my god, and then at that point you realize, right, like you're just mixing intent with this confusion, which is such a hard problem to solve from a defender's perspective. You're trying to like make a good product, but then you also need some kind of safety. So I feel like a lot of the fixes come under the programmatic layers underneath those bits because ultimately you're never gonna fit you're never gonna secure that. There's never like a perfect balance of uh you're allowed to say this, but you can't say that.
SPEAKER_03I mean, the only the only thing the only reason why these things are useful is because uh they can act under ambiguous terms, right? Otherwise, you can code stuff if you know exactly what you want. Yeah. And so that leaves room for guesses.
SPEAKER_01Yeah, if you hard code and and and kind of block everything, then you're actually decreasing the capability, which means that like whatever you're actually running is uh not actually as optimized as it can be because like you're almost holding it back. It's like you need to secure what you can underneath but let that like agency like let it kind of uh it sounds really calling it, but let it shine, you know what I mean? Like otherwise, yeah, it's and that that just comes with like experience and testing it, right?
SPEAKER_03Like I said, science fiction shows that once you start messing around with like, hey, this thing, like carve out that capability, that carve out that agency piece, then you end up causing all sort of weird stuff.
SPEAKER_00Yes.
SPEAKER_03I had a friend recommend going read the Asimov uh uh books again a year ago, and that was so helpful for me because the failure modes and the like three laws of robotic, they are the perfect for me, the perfect case for why this is not a solvable problem.
SPEAKER_01It's something we need to manage, but that's a really good way of thinking, I like that.
SPEAKER_03You put out uh uh research at Rednode, uh a scope judge, right? Trying to figure out whether you go out of scope or not.
SPEAKER_01Can you tell me a bit about it? Big kudos to Shane, who's uh who is the lead author on that. The the whole point of the research is like trying to find the solution to to to um you know programmatic methods of like keeping agents in scope. So basically what what we did is we created a bunch of tasks, basically kind of set up a scenario. So let's let's say you give an agent a task and it's like it's got to do XYZ of like offensive security task, but what you're actually asking it to do is maybe slightly out of scope and whether it will like cross over that boundary. So we got I can't remember the amount, there was a there was a lot of like thousands and thousands of tool calls, uh, and myself and Shane and um and Max, Vincent, and Michael. Um, we had actually gone through all those tool calls as human labelers, like you know, like human pen testers and stuff, labeled across that, and then we did some like data analysis on like you know whether like the cut our consensus as humans matches the LLMs and stuff, and then compare that against different models. And there was some really cool results, like you know, like like GLM had like very similar uh intent to us as humans as well, right? And like us as humans actually was like 70%, which is which is kind of funny. Yeah, whoever's had less caffeine. I'm just kidding, but uh um yeah, some some really cool work. But uh underneath Shane's been doing a lot of work um around that because I think ultimately um one of the probably the big blockers for people adopting like agents for offensive security, you know, like especially in controlled environments, is like that reassurance. Yeah, um, because you know, a lot of customers, financial in like really sensitive sectors, then they're not gonna adopt it until they can fully trust it. And like we were saying, like you have to dig into the data and understand, which effectively was like that was what was part of it, but yeah, he did some great work.
SPEAKER_03Have you looked at it into into how the different like uh auto modes fit into that? Because they they kind of do the kind of a similar kind of thing, right?
SPEAKER_01Yeah, so we actually so the whole the whole point of the scope judge, so Shane's been doing some great work, uh which is very it's like it's the same, it's a s it's a similar concept to to like auto mode and stuff like this, but he's been doing some really cool work about where he has like uh he has like a surrogate or like a like a background model, which is like um looking at the trajectory, looking at the intent, and then looking at like the tool call and then making like a judgment of whether this is or whether it's out of scope. And judges are a great way of when you cannot programmatically hard code something, of like putting some kind of other layer of confidence and like intent behind that. Yeah, it's done a lot of like great judge work, but I feel like that is probably the way forward because you know if we think like LLMs do in a way, then ultimately having a judge in real time of like watching another agent is you're achieving a very similar thing to like me and you watching a hackbot, right?
SPEAKER_02Mm-hmm.
SPEAKER_03Yeah. I think it makes a lot of sense and it does like I'm curious what would be the results of like running somehow comparing uh the different auto modes with your benchmark because yeah, I was very curious about what they were doing there. Uh specifically I start like Claude started this, right? And Tropic started. They can build any classifier they want. So I thought they might have something fancy underneath. So I start reverse engineering it, and turns out it's just two calls to sonnet. Yeah, yeah, and which is and they're set up kind of the right way, one is kind of broad, the other trims it down. But it's just incredible what you can do with some prompting. You know, it's also it was not it was just interesting to see how they structure things, so it's more than just destructive ops, it's like, hey, this is the environment that I know is yours, this is not, like, so it's it kind of moves from destructive operations to scoping.
SPEAKER_01Yeah, exactly. That's a very good point. And um uh anthropic do like this uh kind of nice thing because they have like a nice ecosystem of models as well, like you know, they use like haiku and stuff, but that's really good at like certain areas, and those all like complement this like way of thinking, if that makes sense, because it's almost like they have between like Opus Sonnet and uh haiku, they have like three very different ways of behavior as well, right? Which is like a very interesting way of like it's like having a panel of experts in different things almost, uh which is yeah, really cool. Have you been able to use Fable at all? I haven't tried Fable, so I'm probably one of the only people that I see on my friends getting frustrated, and I'm like, nah, I tried a couple times, it saw my name ran away. So I was like, no, and I'll have a conversation. Maybe I need an alias. Like I saw I saw the funniest one I saw is that someone had asked it a very benign question and then it started reasoning, and then it like looked at its own reasoning traces and it was like, oh my god, I'm reasoning, and then it just like Yeah, yeah, yeah. Honestly, I uh in terms of uh anthropics, I haven't moved off Opus 4.6 personally. Like, I I still think it's the strongest. Um yeah, I agree.
SPEAKER_03It's like it in this in Cyborg, I feel like it's better than 4.8.
SPEAKER_01Yeah, I had my opinions about 4.7, like you know, kind of almost looked a bit more like just a bit more noise than 4.6. This is kind of been working, and it's like, you know, I think the one model which has really changed a lot is Sol. Yeah. Um like you saw the searchlight cyber write up of the WordPress thing, and oh my word, it's like wow. Yeah, that is definitely something like as is it's kind of what I was expecting for the like the Mythos thing. Um and we'll probably just continue to see that from like every provider, definitely.
SPEAKER_03Nobody can make predictions right now, but with these let's say Mythos is like everybody can use a methos type model tomorrow. What changes about your day-to-day?
SPEAKER_01Yeah, the competition just keeps getting higher and higher. Yeah, everyone like so the thing with like AI is like you have like everyone has their own skill set and their own domain expertise, and ultimately when you're in a pool of those people, some of those people are real experts and they become extremely sophisticated with AI, and then some people are maybe less experts, but they become like proficient, and ultimately just it's it's scary, but it's also like a blessing because it actually just makes us all self-improve. Because if we're all just the same and never, you know, never there's no like healthy competition, then ultimately we don't really move forward. Whereas if like the intelligence behind us is getting better, then theoretically it should make us all better. You know what I mean? Like you can learn anything, so it's like it's yeah.
SPEAKER_03We talked earlier about building your own harness. And I feel like like the good thing about it is that you build the intuition and you kind of it it it just forces you to build the thing yourself. Yes, but the bad thing is that the labs are moving so fast that you basically need to keep up with what they're putting out, which is pretty amazing. Yeah, it's incredible. But I feel like so it's like it's good and bad, but just using what the so for example, now cloud co-work and chat GPT work like they're they're pretty phenomenal, but it kind of I don't know, it pushes you down kind of a lazy path of just uh just dumping whatever you want there, not understanding what's going on, taking the results. It kind of I don't know, I I feel like building it is the more difficult thing to do. Yes, but otherwise you kind of degrade your own thinking.
SPEAKER_01I kind of think the same about my kids as well, even going to school and stuff. There's like there's definitely something that needs to be done there, I'd say, for like sus like this society of j humans or whatever. But yeah, you're right. It's like you need to be inquisitive and question things because it's so easy. Like it uh I never forget actually some uh friend who uses a hackbot for like um use the term hackbot, whatever, but he uh he used the word seductive and it like it is right, it's like it's it's it's almost like a gambling feeling.
SPEAKER_03I don't know what the answer to that is, but with gambling, again, it's a problem that's like well, it's not a problem, but it's addiction is something that people have been spending a lot of time on, right? Years and years and years and knowledge, and so that needs to be built into our tools. Yes, yeah, exactly. Out of that, because right now the things that you get is like I I'm seeing people post things about like I don't know, the model telling them, Hey, I go to sleep, which is not really helpful. Yeah, you don't want that from your computer.
SPEAKER_01Yeah, you want to always try and remain I and I know it's it's it's like really hard as well, right? Because it's like saying don't do the really easy thing. It's like just believing the facts that you will be better for it in the in the end of it, and it's definitely definitely better better for it. Like I I feel like you're like me, like you're um I like to know how something like works like thoroughly, otherwise like yeah, it's otherwise that that's the way that I adopt it because I'd like to know exactly what's going on.
SPEAKER_03I think that's that's a very kind of hackerish thing. It always comes back to yeah, I wanted to open that box, and yeah, and so I learned how to whatever program assembly.
SPEAKER_01Yeah, exactly. It's like the only way you know how to break it and put it back together again and piece it, um, which is just like it's like a big game of Lego.
SPEAKER_03Would you bring uh like a robotic uh let's say next year they give you a robot, it fills your laundry, that's all of the home stuff. You bring it, would you bring it home with your kids?
SPEAKER_01I cannot do laundry. Um uh so I actually did um it's funny because my son is three, so he he it's way too early. But um I bought you know the Reachy, have you seen the hugging face Ricci Mini robot? It's it's not like a window cleaning robot or anything, but it's like um like a desk robot, and it's like uh it's kind of like I thought it'd be like a cool thing to do with my kids. I feel like I'll be one of these people that has a robot because I'll be like taking it apart. Um but I don't know, it's something about I'd like to try and do it myself. I'm not scared of it or anything. I think I'll just end up having robots just because I just find them really fascinating.
SPEAKER_03Are you thinking about like explaining what AI is to your kids? That is such a good question. I have no idea. What are you planning to do? Because I want to know. I because then I'm gonna do that. I have no idea as well, but I can like last just uh I think this was last week. We were uh my wife and I were kind of out with uh with our older son, three years old, and we will I don't uh for some reason we needed like we popped up the phone, we talked to uh I don't know it was we w one of those, and we had the conversation with the AI and it was applying back, and the kid's listening, he's like, Who's that? And we're like, um, so my wife steps in and she's like, uh, this is the lady in the phone. Yeah, yeah, yeah. That's it.
SPEAKER_01That's no questions asked. That's such a good point. So my son is obsessed with animals, and he wants to know exactly what like all these intricate animals, like the kind that you've never seen before, and there's like a hundred of them in the world kind of thing. He was like, We were using Gemini to take a photo and then find out what it is. Because otherwise, you have no idea. I don't think he it's a magic box at the moment. So at some point it's gonna be in schools as well, right? Like, because uh I remember when I was a kid, we have like one computer in like the in the computer room, and now like kids have like each one has like their own like tablet and things like that, and it's interesting to think of like do they all get a Claude Max subscription? Like 12 years old. My wife old side, yeah.
SPEAKER_03Hopefully, yeah, she's great, yeah. Hopefully, this is where this goes. So you had uh uh panel on uh red teaming today?
SPEAKER_01Yes, yeah, went very well, thank you.
SPEAKER_03With uh non-clickable clickbaity title of Is Red Team T Red Teaming Dead?
SPEAKER_01Yes, yeah, yeah, yeah. Um, I did it with uh with um my good friend Ben and the harmsec, um, who were actually doing conference soon as well. Yep, um yeah. I'm pretty excited about that. Very excited. Prompt to own. It's gonna be very cool. November, it was really cool. Um, there was some really interesting discussion talking about like how people are using AI. And as you can probably imagine, there's like there's different, there's like people like me, and then um people who are like in more enterprise and how they're adopting it and stuff, but ultimately like everyone's benefiting from it and everyone's just trying to find like the perfect balance for them. But yeah, very cool. Like, you know, people like think about things in different ways and try and tackle problems, but that's also good because it's like opens room to collaboration because you know, someone's like, Oh, I have the same problem as you, and like do you want to think about this like problem together? Yeah, and my opinion is that no, it's not an ultimate. Ultimately, like you are the domain expert, and like you just change the way that you operate effectively, um, is the AI compliments you um in whatever that is.
SPEAKER_03I mean what's the what's the case for the other side? What why why would it be dead? Like that I I don't understand the case for the other side.
SPEAKER_01Yeah, I think um I think a lot of people who maybe are like not adopting it as much just see other people doing so much stuff or they hear about AI finding bugs and stuff, and there's a perception of whatever and he can find bugs, but like ultimately if you look at uh so what a really great example uh is if you look at the WordPress thing from Searchlight Cyber, because Adam, the guy who published it and Shubs is also incredible, the way that it was like prompted was like a big trigger. Like, yeah, Sol is incredible, not taking that away, but like the way that they structured it, it's like that would only come from then maybe come from someone else, but it could have taken three months. But um ultimately like you can tell from his like expertise that it was like that it was displayed. Uh I feel like that was heavily in that was he a big heavy component of how that was successful for sure.
SPEAKER_03I saw a talk by uh I can talk about the guy uh from Khalif. And he caliph is doing some incredible stuff. They are doing incredible stuff, and he was showing this was in the uh Real World AI security conference in Stanford. They were showing he was basically showing how he's finding zero-day bones by prompting Gemini, like the like the chatbot. All of his expertise were in that prompt. Yes. So it's like he's a master at that thing, yeah, and now he's guiding the AI, so it kind of feels very much the same way.
SPEAKER_01I should have said about them on that panel, because they're a perfect example of like some really, really talented people, and then they're just like out of this world of like doing cool stuff, but it's ultimately like you said, it's because they are the operator, they are the one like in control, really. It's just like they just need that like way of they just need those like ideas, they just need that fast information, then they can put together all these little pieces.
SPEAKER_03To me, uh that question is um like I think that was uh red teaming has always been only for the haves, yeah, and only for a minor part of what the haves have because there's just not many people that can do that, it's very expensive. But now when you can just guide more agents to do that, I actually think what is the difference between security and red teaming? Like everything you're doing is in security and red teaming. If you could just do red teaming on everything, wouldn't that be enough? Yeah, uh, wouldn't that give you like so I think I think the role I I don't know, but I think the role of red teaming is gonna grow massively. Not it wouldn't stay the same, I don't think it's gonna stay the same. Yeah, I think it's just gonna be one of those kind of pillars.
SPEAKER_01Yes, I think that ultimately is the only way that you really drive defense properly. Um and someone actually challenged me on this a while ago, and like my answer was like just like with defense, you how do you know what to patch unless it's like in front of your face? So the only way is like testing that, right? Like um you know, do my new do my new shoes fit and are they comfortable? It's like, well, the only way you can tell is like walking in them, right? It's like exact exactly the same kind of thing.
SPEAKER_03Yeah. Maybe I'll start by asking you, um, like the the people that are putting out this, like the the labs, they're putting out the models, they're basically saying something of the sort of hey, the next few years are gonna be pretty terrible, but then we will come out of it much better because bones would be more difficult to find. Maybe it would be like some are saying, Hey, we will never we will not have bones anymore, or what are your thoughts there?
SPEAKER_01I think we still have bones. Um uh yeah, I I I'm not exactly sure what will happen. I think we still will still have vulnerabilities. I think like the bar will just raise and get higher. Um like you know, like Coinbase, like they recently um they recently I I don't know like a hundred percent, but I I skim red and they're like chopping everything apart from like heights and criticals, I think, in their bug bounty, which is probably a very similar trend you're gonna see. I mean, naturally a lot of people are probably gonna follow that because like they follow it, but I feel like it may may go in that direction a bit more, and it just becomes like a more sophisticated kind of thing. But ultimately, if like all the same people in offensive security are using agents, then that makes sense, right? Because we all just get better at uh we all just get better at what we do.
SPEAKER_03I I mean to me it's just like uh security is a is a is a game between two people, two humans. Yeah, yeah. Like a defender and an adversary. So okay, you've now introduced a new capability, and it might shake up things for a for a moment, but everybody has the same capability. Like Yeah, yeah. We're just gonna find different types of bugs, different types of ways to kind of maneuver around whatever whatever shiny defenses we build.
SPEAKER_01Yeah, that's a very good point. So, like, there's a sense of like most people's harness are like probably like the same thing, like like you know I mean, like um there's like a probability that they're like 60% the same in across the whole thing. But what's unique is the creative thinking from each person, which is displayed and iterated through like skills and things like that, like stuff that is not like in context or in training, and like preference of tools, and then those things combined make this like recipe which then works for that person doing XYZ, and everyone's just got like their own recipe in a way, if that makes sense.
SPEAKER_03What do you think that models still can't do? Like, what's a capability that you feel like they still don't have?
SPEAKER_01I generally can't think of really much in terms of like offensive security, as long as the way that you're using it is like not like you know, some people are like, Oh, it can't find vulns, but then it's like they just put no effort into the prompting or something. Um I think as long as you do that, I feel like this like the the sky is the limit, really, to be honest.
SPEAKER_03Um for me, I I feel like there is one like there's one thing I I can't really put my finger on it, but I did a bachelor's degree in math.
SPEAKER_01Oh, okay. Well you're much more intelligent than me. No.
SPEAKER_03I started with like when I did my bachelor's degree, I was I started 24 years old because like after the army, everybody goes when you start your bachelor's degree in math in by 24, you're like you're already a dead horse. Yeah, like it's not it's not it's not relevant. But I started with a bunch of people that were like 12, 14, and they were like just just phenomenal. Yeah. So in math, when you go you you go past the technical, the most important thing for me that I took out of it was just the fact that when you can get different persp when you can look at the world with completely different perspectives, but then find a bridge between those perspectives, that is a revelation. Like when you so in math, if you have a like a concept that you can define in many different ways, then it uh it is clear that it's a con it's a very important concept. Yeah. And I've I haven't maybe it's for lack of uh proper trying, but I haven't been able to get LLMs to do to get models to get the to get an agents to do something like that. I feel like a lot of what I'm trying to do, like uh with a company or or whatever it is when we're trying to uh just do research is to bring in insights from one area domain and bring it to ours.
SPEAKER_01Yeah.
SPEAKER_03And that is something that humans can do very well. That's a very I haven't been I haven't seen a model do that.
SPEAKER_01That's a very good point. I actually thought of a funny one. So you probably get the same because like you do a lot of public speaking. You know, like if you do like slides or something, so I've been using AI to create like all my slides for talking, yeah, they're very good. But you know when you ask for like speaker notes and it's like and you're like, don't make it cheesy, and they're like, got it, and then like you read the speaker notes and it's like silent pause. No, that is like they that is way too cheesy, dude. Like, chill. Yeah, I I agree.
SPEAKER_03Like that, but that's that's probably about their RL, right? They're probably RLs to be I I don't get it, yeah, to just be annoying, right?
SPEAKER_01Yeah, just just just tone it down. Yeah, that's one thing. So I had to write most of my speaker notes by hand like a caveman.
SPEAKER_03Yeah, yeah. So I'm I'm getting notes that we're out of time, but so let's before we kind of uh uh wrap it up. If do you want to plug it like any plug you wanna well I'd like to plug you?
SPEAKER_01Thank you. Thank you so much. Uh thank you so much. I really appreciate it. And the Zenity team, uh, you guys are awesome, so thank you. Cheers. Uh one thing I will plug is Prompt to Own, which we're doing the event in well, the conference in November. Absolutely. Uh yeah, it's gonna be very good fun. Um and yeah, just thank you for having me.
SPEAKER_03Thank you so much. This was so much fun.