mnemonic security podcast
Hosted by Robby Peralta from mnemonic, one of Europe’s leading cybersecurity companies, the show features conversations with researchers, founders, operators, and security leaders working across the cybersecurity landscape.
Each episode explores a specific topic within cybersecurity: from incident response, threat intelligence, AI, and geopolitics, to leadership, resilience, and the changing role of security leaders.
The podcast is tailored to cybersecurity practitioners and decision-makers who want grounded conversations about where cybersecurity is going, what organisations should prepare for, and what experienced people are seeing.
mnemonic security podcast
State of the Union: Agentic AI
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Everyone’s talking about Agentic AI, but beyond the buzz, what’s actually happening on the ground? The mnemonic security podcast is continuing to dive into the world of Agentic AI in our latest episode, recorded live at Sikkerhetsfestivalen.
For this episode, Robby is joined by fellow podcaster (CloudFirst Podcast and KI til Kaffen) Marius Sandbu. They look at real-world implementations of agent-based systems, particularly what’s been done in Norway. And try to answer the question: are we ahead, behind, or just cautious?
They also discuss lessons learned from local projects, the current state of the global ecosystem, and what it really takes to make Agentic AI useful; diving into integration concepts like MCP and RAG, and how they’re being applied in practice.
Welcome to the mnemonic security podcast. Every country with a sizable security community has its conference. Everyone gathers in the same place and does whatever it is that cyber people do amongst friends. And in Norway, this conference is called the Seeked Eats Festival. 15 different tracks, 1400 or so attendees, the usual vendor booth, and a podcast studio conveniently located in a microbrewery. And that's where you found me this year. Alongside a new friend of ours, which is someone you'll probably hear more from on the podcast in the future. A hands-on agentic AI system expert, Marius Sandbu.
Speaker 5Welcome to the podcast. Thank you all for joining. And good afternoon. You have two podcasts, Cloud First and KI til Kaffen. But you've never had a live podcast.
Marius SandbuNo, never. So I've never been this nervous nervous in my life.
Robby PeraltaYeah, I was just saying, uh usually I just edit everything out, all the dumb stuff I say, but I can't do that this time. Though at least for two. So you are an expert system in the flesh. You have very many awards. You're a Microsoft MVP, an Nvidia grid community advisor, a Veeam Vanguard, awards from Nutanix, Hitrix VMware, and my personal favorite, A Dell Rockstar.
Marius SandbuYeah.
Robby PeraltaA lot of old titles, but uh you have an about me page on your blog and you have to scroll. Yeah, yeah, yeah. It's pretty long. Very impressive. I have to update that one. In addition to that, you are a certified instructor, you're a blogger. Like and subscribe. You're the author of five books. Is that correct? Yeah. What are those books?
Marius SandbuWell, mostly about networking. And then I did one last year, which was mostly on ransomware detection and protection.
Robby PeraltaCool. So no books on AI, Agentic AI?
Marius SandbuNothing yet. Probably gonna ask ChatGPT to make one a bit later, but I I feel like that time is by now writing books because it takes so long time and effort, and a lot of late nights actually write the different content. So I think I'm uh I think my authoring days are over.
Robby PeraltaYou've heard about that podcast app, or you just put a PDF and it makes a podcast. You haven't used that yet? Yeah, Notebook LM. Yeah. Yeah. So we're we're soon being replaced. It's the last podcast track of the Sikkerhetsfestivalen. What do you do in Sopra Steria?
Marius SandbuWell, if my official title is Cloud Evangelist, but I do a lot of stuff. Consultant work, advisory, marketing stuff. So seminars, events, and go and talk about things that I find interesting at the time. Uh, but also do a lot of like service development, looking at okay, what kind of new products should we be building? So it's uh role with many different responsibilities, hats, and that's the way I like to work.
Robby PeraltaAnd you said 70% of your time is with clients. Yeah, correct. So that means all the other things that I just spent five minutes introducing you, that's done in your in the evenings.
Marius SandbuYeah, so I think a lot of my friends know that I don't have a lot of social life.
Robby PeraltaI was literally gonna say either have either you don't sleep, you have no friends, or you automate your life using AI. Yeah. Well, yeah, it might be a lot of AI. Yeah. Who here has uh built an agent or tried? A couple of yeah. Can you explain what a gentic yeah just at a basic level before we go further?
Marius SandbuYeah, so as we say in Norwegian, we have like the "kjært barn har mange navn". There's so many different names when you look at it, but uh I try to define it as okay, we an agent is a language model, it's a set of instructions, so a system prompt. Okay, what should this specific agent do? And then you have an orchestrator that is responsible for taking that prompt and and the language model and actually doing stuff with it. And then we have a set of integrations, which can be okay, you're allowed to talk to this specific API or this specific system or a set of data. So there's these four basic components that you have. So one agent could be user-based or user-triggered. So let's say that okay, I want to ask my agent about something that I have, like unstructured data, I want to get answers, so simple chat bot. Or you can have autonomous agents that are running in the background, working 24-7 to actually solve a specific problem or work on specific tasks. So I like to say that okay, agents, orchestrator, language model, instructions, iterations, and then you have them either user-based or autonomous.
Robby PeraltaAnd of the 70% of your time is with clients. What is actually being created or what's being asked of you these days?
Marius SandbuWell, in most cases, there's a lot of okay, we need to get a better understanding that technology actually does. Because I feel like a lot of customers that I'm talking to is like, oh, we need to use AI. Okay, why? What do you want to solve? And they're trying, well, we have some different use cases, and then it's looking into okay, can we actually solve these from a technical standpoint, or is there something that we need to wait further down the line until the technology actually gets there? And then we have those that have done their research, they know what the technology can do, and then they see that okay, we need to look uh building an agent that can do different tasks for us, like line of business actions. Either if it's automatically scanning invoices, automatically scanning incoming emails and being able to categorize those and send them through the right destination. But right now we are looking more into agents that are trying to help those working in IT operations. How can they solve issues faster with using agents? Has anybody done it yet? Well, we have some that are in the early phases, but most of them have only gotten to the point where they have like a chatbot that can answer questions related to data. And right now we get more and more of these integrations that makes it easier for those agents to actually take action. So let's say that okay, we have a known issue appearing in the monitoring system. Instead of someone at the operation team actually looking into that incident, okay, knows what the issue is and he knows what he needs to actually solve this, he can do an set up an agent to actually do that instead. As long as everything is documented, the steps involved, it's a lot easier now than it was six months ago.
Robby PeraltaSo right now, agentic, the the state of the union in Norway is it's basically like internal use cases.
Marius SandbuWell, we also have a lot of external use cases. So um let's say public facing chatbots from my point of view and what customers that I'm working on. But I know there's a lot of other use cases and a lot of a lot of other uh projects as well being used using agency or agents from a end user or public facing perspective.
Robby PeraltaSo the chatbot part, isn't that just an LLM? Like if you have a you said you had certain clients that were having like chatbots externally facing. Yep. But is that agentic then at that point, or is it just an LLM?
Marius SandbuUh well, it is agentic in the sense that it's it has a task. So when you ask it a question, there's an orchestrator that takes the prompt or the question that comes in and he needs to maybe slice it up into different tasks. Okay, might be that I need to get some data from an internal source, might be that you have different data sources that it needs to look into and find data.
Robby PeraltaOkay.
Marius SandbuSo uh one customer that I were uh working with, they have many different data sources. So when the chatbot gets a prompt or question, he needs to look into 10 or 15 different data sources and try to find okay, what's the most relevant data for this question and generate a reply. But it's really simple to set up.
Robby PeraltaThat's knowing what I heard yesterday. Uh, I'm not sure who was in the application security track yesterday, but it's basically all about how was it called the missing S in MCP? That sounds extremely scary to be exposing that to the outside world. But that's happening.
Marius SandbuYeah, yeah. I think that a lot of organizations are they they want to try it out. Okay, how can we solve this? How can we get an LM to integrate with all our different third-party systems and and services? And MCP is uh really easy way to get started building those integrations. Fortunately, MCP has gotten a lot of upgrades since it was initially released because MCP is still like uh 10 months old protocol, because it was released in November. And we got a new specification update in March, and we got a new specification update now in August. So there's the S is soon there. Yeah, but it's coming. Um the problem that I see, especially with MCP servers, is that there's so many community-based MCP servers. So people seeing that, okay, well, I want to build an MCP server for a specific third-party tool. There's no one available, fine, then I'll make one. A lot of these also don't have any security mechanisms built in, so you need to post in your API key, it needs to run locally on your machine. And I've also seen scenarios where these MCP servers, which are just pasted on GitHub, where someone adds some malicious code into it so it can run specific actions to collect environment variables and send it to a third-party service instead. So a lot of things you need to consider there from a security perspective when you're using MCP.
Robby PeraltaWe're gonna come back to that in a second. But I just want to uh what's probably like one use case that everybody and all the companies and organizations in the room could actually should be looking into and can implement that using that's smart? Like where do they start? What should they start?
Marius SandbuWell, most companies suffer from one thing, and that is they have so many unstructured, so much unstructured data. And making information available but inside a company, because one of the things that I see so many spend a lot of time in is okay, do we have data on SharePoint or internet or is it on file server? Is it in my email? Spend a lot of time searching for content. So I always look into okay, how can we solve this? Make that information easier accessible and making it understandable so that anyone can get access to the type of data that they have inside of a business. But of course, there's no like magic pills. You can't just magically set on AI and make everything work, but it's uh probably an important use case that I see most companies work on. How long does that take? Well, how much data do you have? Well, the problem is that we spend like 90% of the time preparing the data, making it available, and the last 10% setting up the agent. Yeah.
Robby PeraltaWhat are some of like the common pitfalls you've seen in the projects you've been a part of?
Marius SandbuWell, it's again one of the things that we typically see is that companies think that AI is going to solve everything. It's not going to solve 20 years of neglect when it comes to no management or data governance. So we spent a lot of time on preparing the data, making it available and making it usable in a rag scenario. So we want to have an AI to actually find relevant data. Then it's also understanding the technology, okay, what is actually possible with the technology and the frameworks that we have today, and also the language models. Because one of the things that I've also seen is that a lot of these frontier models from Google, OpenAI, even though they support Norwegian, they don't have the in-depth understanding of different dialects and well, book more than the other stuff we have here. So sometimes we also need to use Norwegian language models to get a better understanding if we have a lot of data on Norwegian. And of course, this ecosystem is still extremely young, right? Chat GPT was released under three years ago. And now we have all these agent frameworks coming out constantly, new updates, new versions to models. And the problem is that when you're starting on a project and you're building something, uh the framework might get an update and the documentation is missing. So we spend a lot of time actually troubleshooting, okay, why isn't this working? Uh but now we've gotten to the point where these products are becoming more stable, documentation is fairly available. Um, so it's a lot easier now to build these agents than it was for like eight months ago.
Robby PeraltaYou said the word ecosystem. What does that look like in your eyes these days? Because uh we didn't talk about there's that podcast that's called like This Week in AI or whatever.
Marius SandbuYep.
Robby PeraltaAnd it's like a four-hour episode, and they cover you know everything from what NVIDIA and Google are doing to yeah, what is what's the state of the union on the ecosystem right now?
Marius SandbuWell, it's chaotic to say the least. I think I spent like one and a half hour, you know, two hours each day just to figure out okay what's actually going on. And uh during the summer, like if we look at June and July, we had like 150 updates to different models available in market. And if we start by looking at the language models, we have like OpenAI, we have Google, we have Meta, Facebook, we have Microsoft, they have a partnership with OpenAI, but they're also now building their own models. And of course Scrock from XAI. And we have a company in France, Mistral, they also have a lot of new updates targeting mostly the EU, and of course, they have a much richer data set when it comes to new Scandinavian languages. And then we, of course, have China. So they have Deep Seek and Alibubble Cloud, they have a lot of different language models. And the thing that is a little bit interesting to see now is that we have all these large frontier models, which are only accessible from the cloud. But more and more now are actually making their models open source, meaning that okay, you can download a model and run it on your local computer as long as you have enough GPU memory, or you can run it with any of the infrastructure. And even I think it was two days ago, Elon Musk released Scrock 2, their older version as open source. So we see that more and more of these language models now are becoming open source and easily available, so you can download them and run on anywhere, any infrastructure that we want. But then we of course have the agent frameworks as well. I think I see like a new agent framework coming up every day and being released on GitHub. What does that mean, by the way, an agent framework? So it's the orchestrator, the engine in between that connects the language model and the instructions and the integrations that you want to have. So if I today want to create a chatbot, which is going to be accessible in Teams or in the Microsoft ecosystem. The easiest way is to get started, okay, we use uh they have a platform called Copilot Studio. Easy way to get started in building um your own chatbots. Google has a similar called Agent Space. Salesforce as one ServiceNow has a framework. But of course, then in most cases, then you're locked to a specific language model if you go into one of those vendors. But then you can also use SDKs or frameworks. So we the most commonly used framework now is uh Langgraph, which is an open source uh framework. Also also provide a cloud service, but uh that way you have much more control of okay, which language model do I want to use, what kind of parameters do I want to integrate, what kind of data source do I have? So it's a fairly complex ecosystem. There's so many different frameworks that you can choose from, so it's kind of difficult to understand, okay, which one should I choose, depending on what kind of systems that I have, what kind of language model do I need. Uh, because let's say that I okay, I want to create an agent that's running completely off on-prem and no access to internet, okay, then I need to use one of those agent frameworks that are available.
Robby PeraltaMaybe a dumb question, but when you're like actually building an agent, is it like, do you need to code anything, or is it kind of like vibe code? You just tell it what you want to do and it does it?
Marius SandbuWell, you can do the different agent platforms from like Microsoft, Google, and Amazon, which also has a service called Bedrock, is pretty no-code, low-code approach. You can specify using prompt saying, okay, I want to have an agent that does this and that, and it automatically creates an agent for me. And if I want to add a data source, it's just two clicks and you specify data source. Uh, if you want to use the let's say, development approach using Langraph, then you use writing Python or JavaScript. So, no, you don't need to actually do any coding. Uh you can do it using the UI, or but of course, if you want to go into nitty-gritty details and tweak it a lot, you need to uh use the development approach.
Robby PeraltaYeah. So you should have like a I was gonna say data scientist, but you don't need like developers to be doing these projects.
Marius SandbuNo, you don't need to, but of course, when you're doing these lot more complex agents, you need to have some programming understanding or development understanding.
Robby PeraltaYeah. MCP. Uh first of all, what is that for the people that haven't heard about that? And uh I will start there.
Marius SandbuYeah. So uh MCP is a protocol standard that was released last year by Anthropic, which is one of the vendors in the LLM space. They have uh uh Claude as uh one of their uh language models, but uh MCP stands for Model Context Protocol, which is well, the easiest way to put it is like it's like USB, the USB standard for connecting AI to data. Now, uh before this protocol was released, then let's say that okay, I want to create an LLM agent that's talking to some third-party service. Then I would need to create something called GPT functions to say that this language model is going to trigger the specific action uh if the user asks for something specific, so like keywords. So then I would need to specify all these integrations using GPT functions and specify, okay, this is the endpoints and API that I'm talking to. Now with MCP, one thing is that more and more vendors now are supporting this, even though it's a fairly young protocol. And they create these different MCP servers, which are preset of integrations that you can use. So that it allows me to easily set up an integration between a language model and a third-party service. So one thing is, for instance, GitHub has its own MCP server. So that means that I don't need to develop all those integrations if I want to have a service that talks to GitHub. I just specify the MCP server that I want to use, and then I can use natural language. Talk to the language model, say that, okay, what are the top 10 repositories in GitHub? And based upon that information and the MCP connection, the language model knows that, okay, uh he's asking a question about GitHub, then I need to talk to that service. So uh the good thing now is that we have probably about close to 5,000 different MCP servers available. A lot of community-based, a lot of official MCP servers from different vendors. So it's fairly easy for me to just set up a bunch of integrations and get started just to talk to these different services. Um previously, these MCP servers needed to be run locally on your client. So if someone, and of course, some of these MCP servers also had access to your local operating system. So they could run tasks and scripts locally. So of course, if you had an MCP server running locally and someone added like uh format C or run some PowerShell commands locally, they could take control of your machine. But now in the latest release of MCP, they now support uh one thing is that they now support OAuth authentication, so it can have better authentication mechanisms that have like an API key. And the second part is that these MCP servers now can run remote. So they don't need to run locally on your machine, they can be hosted in the cloud service. And these MCP servers then acting like an API gateway, so communicating with different backend services that you uh need to the different backend servers you connect them to.
Robby PeraltaSo based on the feeling I got from all the presentations in the app sec yesterday, that was like it's impossible to lock down MCP. It's just a there's all the OWASP top 10 for LLMs, like they were just like this is gonna be a huge problem. Do you share that?
Marius SandbuWell, uh I wasn't at the track yesterday, so I'm I I don't know all the details of it, but most of the setups that I have now is okay, I'm setting up an MCP server hosting um in a cloud server somewhere. A lot of these days I'm using Cloudflare. Um and I have I have a set of uh well to actually be able to use that MCP server, they need to authenticate using their own credentials. Um and of course, I'm only using official MCP servers or the ones that I'm building on my own. So of course, as long as I have control of the source code and the runtime environment, or that I know that I can trust those third-party vendors that I'm actually using, then I don't think there's not a big risk at the moment. So you can lock it down if you just do things properly from architecture point of view. Yeah. I think the problem is that a lot of these MCP servers can be well, doesn't have any authentication methods. They haven't been updated to use delay specifications. So they can run locally, they don't have an authentication thing, so they can do a lot of executable or stuff on your on your computer. Have you heard of any like war stories from coming out of well it's it's still in the early phases, this uh ecosystem, so I don't have any I have I have a couple of examples that I can talk about, yeah, even though some of these are going a few years back, but just to give a little bit of context. Initially, when when ChatGPT was released, it didn't have any good integrations to actually parse content within a PDF file, for instance. So if you wanted to parse PDF files with ChatGPT, you needed to use third-party uh services. So before this was called plugins. So one customer that I was working with, um, he was using a plugin to parse PDF files. So it was uploading a lot of reports on PDF. Um, and then this third-party service was parsing the or generating understanding the content using some form of OCR and then presenting the content back to ChatGPT. So it was really proud because it saved a lot of time to actually understand the content. Um, and then we're a little bit curious okay, what does this third party service do? Where is it hosted? And eventually we found out that uh this third party Service was hosted somewhere in China. And he has been uploading documents for the last two months. So we have no information about the cloud service that was running underneath. We have no way if they're storing the data or what they're doing with the data. So of course he was uploading data. It was sent to China, processed, sent to Chat GPT in the United States, then presenting a result back to the user. So of course, one thing is to be aware of those different integrations. We also had another issue where we were another case, I wouldn't say it's a war story, but we were creating a chatbot, public-facing chatbot for an organization. And this was in the early days where those language models didn't have a lot of security mechanisms or content filters built in. So what we found out after we added monitoring to the chatbot is that people were using it to create recipes for creating bombs. So we needed to add a lot of security method or security mechanisms into the chatbot so it couldn't be abused by others that were able to do prompt injection to it. Because there were so many weird uh recipes and instructions being shared on that chatbot. We just just there to actually answer commonly asked questions. But underneath we still have the language model, so the knowledge is there, but you need to buy, you need they needed to bypass that security mechanism before they're able to get that answer.
Robby PeraltaSo by the way, so you so that that means that in practice you're building guardrails around it, right?
Marius SandbuYeah, yeah.
Robby PeraltaCorrect. What does that actually mean?
Marius SandbuSo it's essentially a filter, right? Because right now, there of course there's a lot of difference between the different language models. But I send, let's say that I ask a language model, okay, how what's the recipe for creating a bomb? Or how do you hijack a car? And these different filters are built in into the language model to try and see that, okay, this is something that is uh trying to abuse or malicious intent, then it's and then it automatically tries to block that um that request. Now, of course, uh some of these uh say, let's say safety or guardrails built into the language model, sometimes they can bypass those using prompt injection, which is trying to hide context within another context and ask asking a question to bypass those. Now, when we're using guard rails, of course, there's many different tools there as well, but those guardrails also are uh they're kind of like a proxy. Okay, we're sending a new prompt. That proxy is intercepting the uh the text that is coming in and looking at okay, is there any malicious intent there? And then it blocks the request. So it is essentially just a proxy looking at incoming requests and blocking uh those with uh based upon category or bad words or coming from a known source, those types of things. Like kind of like a firewall.
Robby PeraltaYeah, but you have to know what you want them to block. Yeah. So you have to specify everything. Correct. Yeah. Luckily you don't have to do all that because the models, most of the frontier ones, have those guardrails built in.
Marius SandbuWell, most do, but uh there are some that are not as good or not as focused on it. JPT is fairly good at it, but if you look at uh Grok from Twitter, they don't have design. That's by design. Yeah. Yeah, it's yeah. Um something about free speech, I think. But it's not uh well when we look at the technical reports on the uh language models, they're not as focused on that compared to uh JPT. And same, of course, with DeepSeek. They don't have a lot of the same levels of guardrails as well. Unless you talk about President G.
Robby PeraltaYeah, correct. Yes. Agents being used in like the security industry. You wouldn't say that you work in security, would you? Well, I I used to work a lot more in security before, but these days it's mostly AI. Yeah.
Marius SandbuYeah.
Robby PeraltaSo I was at RSA, and there it's like the word agentic was every every company had agentec AI. What are your thoughts on that? Just on a high level.
Marius SandbuWell, we we see all the security vendors now adding agents to their security tools. But in most cases, those agents now are fairly simple. Some of them just okay, we there's an incident, we analyze it. Uh the problem is that when you're using agents built on generated AI, they don't have the ability to handle a lot of large data. So most cases these agents are just looking at specific incidents and trying to make a summarization. Maybe they talk with your internal knowledge base and try to pinpoint okay, what's actually going on in a specific agent. But right now we also see that there's new agents coming that can not just summarize content and find information, but also do actions. Just as, okay, here we have someone that's trying to break in using some form of vulnerability or trying to bypass the firewall rules. And then we can have an agent actually see that okay, there's something going on. I see the IP address, I have an integration with the firewall, so I can automatically update the firewall rules to actually block out that specific connection. So it's it's still in the early days, but it's kind of hard because a lot of these tools, first off, what I see is that some vendors are now creating their own language models, which are much more security focused. For instance, Google, I don't remember the name, but they've created their own security language model, which is trained on their own content that they have from their own security tools. So that way it's of course much more understanding on how to exploit vulnerabilities, how to navigate through the network, how to analyze logs compared to the let's say general purpose models. But what what what we've been missing now is these integrations, right? Because, okay, how can I make sure that if someone's trying to break in, that I can block out the user from Active Directory, or block out the user from the network perspective, or make sure that I isolate the device that's being compromised. So what I see now is that most vendors now looking into okay, how can we create these different integrations and make them work properly across the different products?
Robby PeraltaBut that's the scary part, right? Once you have products that have access to all those important systems and they're built on an agent or an LM, that's just it's it's a risk.
Marius SandbuYeah, it's Skynet in the making, right? Yeah. No, but I I think that we'll get there, but it's gonna take a long time. So right now I see that a lot of these initial agents that are coming in, it's not an autopilot, it's like they have the co-pilot part where they're just okay, here's a recommendation on what you can do. So the agent is giving you a summarization and things you can do, but you as a person need to take the action. And I think we're gonna have that for a long, long time before we get those more autopilots working. Yeah, so that's like the human in the loop. Yeah, exactly. Can you elaborate on that human in the loop thing when it comes to gen tick? Well, I think it's more about if um we need to make a decision on something. The a human needs to be involved. So instead of the agent actually making the decision, it gives, okay, this is my conclusion or summarization of what's going on. It's up to you now to take the uh decision on this specific task. It's kind of like when my kids ask me for ice cream. I'm the pre parent in the loop. That always says yes.
Robby PeraltaYes, yes. Correct. Last week I interviewed uh this researcher, PhD candidate from Carnegie Mellon at the University of the States, and his friend worked for Anthropic. And so, long story short, they got together and they made an a system. It did like an anonymous, not anonymous, autonomous cyber attack. Like it did vulnerability scanning, found exploits, data exfil like the whole nine yards. And they did that in like three weeks.
Marius SandbuYeah.
Robby PeraltaAre you worried about those sort of things?
Marius SandbuWell, you know, as you said, it's fairly easy to set up. Just to give an example, uh, Shodan, uh, for instance, fairly, fairly just to do external tech surface management, they have their own NCP server they can use. So I can ask it as an agent, say that okay, I want you to scan this company, find as much information as possible, and then I can connect it to, let's say, a new GitHub repository that contains information about a new vulnerability. So it's fairly easy to set up. But again, it's the same type of attack vectors that we've seen before, right? So uh even though now it's a little bit more automatic that you can actually use this set of integrations, and it's fairly easy for anyone to actually do it. But I think that what concerns me more is it's not that well, it's not that because I think that these are the same attack patterns that we've seen so many times before. But what concerns me more now is that when more and more people are using these tools and MCP servers and as part of their daily workflow, as I said initially, like I have a lot of these different workflows, and we have all these different integrations now where attacks can be based on natural language. So we have seen some scenarios already where um, or I if I remember correctly, there was uh integration or a plugin assistant for Amazon Q. Amazon Q from James Bond, right? They have their ancient uh where someone managed to uh add some malicious commands to their plugin, which essentially said, okay, wipe all on your own computer. So, and that was there was no malicious script or obscure command running, it was just natural language. Okay, wipe this computer, even though it wasn't successful, but it just shows, okay. Down the line, we'll have all of these different agents and integrations, and a lot of these instructions are just natural language. So I think that we'll have like more cyber attacks just trying to inject or uh take over these different workflows and integrations.
Robby PeraltaSo but if you were to like follow the OWASP top 10 for LLMs or agents that have like new frameworks, then then you should be pretty pretty good.
Marius SandbuYeah, yeah. I uh and the the problem is also with the OWASP LLM is that like one thing they're concerned about is data poisoning. So uh poisoning the LLM. But of course, if you as long as you use LLMs from, let's say, known sources, the larger models from OpenAI or Google, I don't think that's gonna be uh something that's gonna impact their language model. But OWASP for LLMs also has a lot of guidance also on okay how to build agents, what kind of security mechanisms that you should have in there. So I think that's a good approach. But also, and one thing is that this, of course, OWASP LM is focusing, trying to be broad and general, right, regardless of which framework that you use. But in most cases, you might be using a closed ecosystem where you don't have the same ability to impact or integrate with security tools. So we also need to understand, okay, what kind of security mechanisms is available from that platform that we're using. And if you're using Microsoft, Google, or ServiceNow or whatever, you need to understand okay, what kind of security mechanisms is included as part of this ecosystem because in most cases you don't have any way to integrate with third-party tools. And one example is uh I in some when we're building some agents, we tend to use a tool called LLM Guard. It's an open source tool, really good at handling prompt injection and also to obscure personal information. But since this is an open source tool, it can't be used in any of those cloud ecosystems because you have no way of integrating it essentially, unless you want to build it on your own.
Robby PeraltaUh on the topic of tools. What are what are tools that you think are worth mentioning that are be useful for people that want to play more with Agentic AI?
Marius SandbuWell, I have some frameworks. Uh one is N A N N8N Nate. Nate, I think so. I think it's what it's called. But they have a they have a really good way of building agents. They have a UI-based run locally. They also have a cloud-based service, fairly easy to use. Uh I also tend to use LM Guard for security mechanisms. And then we have um Nemo agent from Nvidia, also a really good framework, and also has a lot of built-in guard rails. And I was also thinking about Langgraph. Or um, there's another one called Langstacks, which is a company acquired by IBM now, but they also have a really good framework to create agents. So you're not bound to any specific vendor. You can specify what kind of language model you want, and there's a lot of pre-built integrations that you can use.
Robby PeraltaInteresting.
Marius SandbuBut of course, now we we we have so many agents now these days. So we have GitHub, we have codecs, we have clawed code. So there's like a lot of uh agents you can run locally from your own computer if you want, and it can actually do anything. Well, a lot, a lot do a lot of stuff with it. And it's changing all the time. Yeah. Yeah. That's it, that's the problem trying to keep up to date on what's going on because I I noticed that uh I think it's later today or tomorrow, we'll most likely get a new model from uh Google, Gemini Tree. So there's there's new updates coming all the time. And like the capabilities and how much information these models can handle is also something that scales all the time.
Robby PeraltaDon't you get worried that whatever you're building for your clients now is gonna get like just become irrelevant because something new comes?
Marius SandbuYeah, yeah, all the all the time. But I think that when working as a consultant and trying to advise customers, okay, try to find a framework that can adapt, use any new language model that comes out. Because okay, before the summer, uh, the best language model for coding was uh one from Anthropic or Cloud. Now we have GPT-5, which is fairly a lot better in some scenarios. And of course, next week it's gonna be Google, most likely. So having like a framework that allows you to switch between different models is really important. But other than that, we well, as I said, there's a lot of changes going on, so it's kind of hard to predict what's what's gonna look like downline in two to three months. Any questions?
Speaker 1Does the customer think about security, or do you need to like tell them, okay, let's introduce a security earlier in the process? Yeah.
Marius SandbuSo uh in most cases, no. They don't consider that or they don't focus on that, they focus more on the technical capabilities that it has. In most cases, they have like one person in the HR department or the leader or management team that says, oh, really useful to have this chatbot to answer all these commonly asked questions that we get. And they in some cases they've set up something on their own. Just plug it in the data and just make it available. So they don't consider the security aspects. Okay, it's this information that should be available. And one scenario I encountered, they actually created a chatbot that was publicly available for anyone, even outside the organization, that was directly getting uh HR the personal handbook information publicly available. So that was really useful. What to go wrong? But luckily we we saw that okay, I've we didn't think that anyone was actually trying it out. It's kind of hard to find it, but it was it was publicly available. So we needed to go back to drawing board and say that okay, we need to consider the security aspects. Okay, what kind of data does that is available? Should it be available for everyone? Who should be able to access it? But I think that in most cases people are, or well, most of the companies that I work on are more focused on what kind of capabilities can this provide. Doesn't that always happen and security comes afterwards?
Robby PeraltaYeah, nothing's changed.
Speaker 7Yeah.
Marius SandbuBut I think that we have more of this FOMO effect now that people are, oh, we need to use AI. And then, okay, but what about security? I screw that. We want to just have a chatbot. And then uh six months later they say, Oh, this has access to way much more information than we thought it would have.
Robby PeraltaAnd needs to have. Yeah. Have you had a uh project yet where you had to deal with some security like a client security team? Yes. How was that interaction?
Marius SandbuWell, I think that a lot of them are from like a traditional security background, right? So, okay, how do we integrate it? How does it talk? An agent across different uh and also the the problem I think that agents sometimes they don't actually trigger that specific action that you want it to trigger. And even though with newer models this is improving, but I would say that in the early days we had like, okay, uh ancient or an uh LLM could trigger an agent like seven out of ten times, even though the set of instructions are the same. So there's no built-in mechanism to actually handle that. And I think that was a also a trouble for the security team to understand that it's not working, it's not going to run the same set of uh commands every time. But again, it's about building understanding of the technology, how it works, because language model essentially is just trying to predict what's the most likely next given word in any given sentence. So, but we also had some scenarios, okay, we need to create an offline service, which is not not using any of the frontier cloud models, but running on a local set of hardware. So, of course, that's something that security people tend to like a lot because there's no access to the internet, everything is running locally, but these frameworks are even younger than running from the cloud. So then we also need to use a lot of these security tools to make sure that we okay, want to make sure that we don't have any prompt injection, want to make sure that we uh anonymize any personal information that's being sent to the model, because it might be that a user is asking the model for a specific person or individual, and we want to make sure that that is not coming to the server. So we use tools there to actually anonymize or remove that uh content before reaching the language model.
Robby PeraltaYou said that that seven out of ten example that you used. Uh shouldn't you just not use an LM for that because you want it to do only the things that you want it to do?
Marius SandbuYeah, yeah, and that's another thing. Uh should it be done using AI or is it just automation?
Robby PeraltaYeah, right.
Marius SandbuSo a lot of the scenarios that I've been talking with customers as well is that okay, this is okay, they pitch me an idea that they have. Okay, this is a use case. You want to use AI. And then they talk to you, they give me the pitch of the idea, and then they say that isn't this just a script? Well, yeah, but it sounds cooler when you use AI. But in most cases, okay, uh, this is condemned by a script. They should just lie to their management then. Like you have AI now. Yeah. This is this is AI. Doing business automation. But uh they try to say that okay, if there's something that needs to repeat it, the pattern is the same, okay, then you can use automation. But if you need to have someone to do some form of analysis of the content that's being sent in or being uh generated, then you need to have AI on top. Yeah.
Robby PeraltaI never thought about that. Thank you for clarifying that for me. Um hidden costs of agents. I'm just thinking like if you have data like going here and there, and like are there what is the cost of having agents working for you?
Marius SandbuWell, uh it's really difficult to say. When you look at some of the cloud providers and their pricing, it's uh well, you need to have your own agent to actually understand the billing terms that they have. But we have we have some uh some experience or some examples. Uh, we have one agent that is like an internal chatbot to answer questions, and I think it's about 28 to 35 crones a month to actually use that use that specific agent. Because firstly, it it takes questions, it needs to find the different data sources and generate a reply. Sometimes it also needs to trigger actions, so it's uh fairly cheap. The problem is if you have really complex agents that need to handle a lot of data, because let's say the largest models today can handle up towards to one million words. So if you like upload a lot of documents which are fairly large and needs to do a lot of analysis, and I think that with the latest O tree model from GPT, it costs $15 to upload to generate all that type of content, so like 150 Norwegian crones just to analyze all that content. So of course it can be quite expensive, but at the same time, since since GPT was released three years ago, the cost of using LMs has gone down with 99%. So it's becoming a lot, lot cheaper. But uh if it's kind of difficult to calculate, okay, what's the actual cost for using this?
Robby PeraltaYeah, but one thing is like the uh the money too, the LM or agent money, that company, but I'm thinking about like the hidden cost of like everything else that you needed to support it.
Marius SandbuYeah. So first off, there's a lot of well uh building knowledge and expertise in a company, right? To actually understand, okay, how do we build this, how do we maintain it, how do we upgrade it when the time comes. And then we have, of course, also integrations and dependencies that the model has, which of course talks to third party. And of course, there's also sometimes you need to have licenses on top. But I think that the biggest cost is building expertise in the organization to actually be able to maintain and develop it. Because as I said, there's so much happening all the time with new updates. And so we're trying to keep it up to date also, it's a lot of cost. Yeah.
Robby PeraltaHave uh have you met, I guess it's so early, but have you met like a company that had really good, uh really good way of like organizing that sort of process?
Marius SandbuWell, some are looking into building the well, what they call it, uh, AI Center of Excellence.
unknownYeah.
Marius SandbuIt's kind of like a separate. You wrote a blog about that. What was your answer? Well, depending on the size of the company, I think, uh, which was my answer because uh some are really small. They have one person working in AI. So, of course, do they have their own center of excellence working on finding AI use cases in the company? Probably not. But if you're fairly large, we have we see now that more and more companies are looking into finding different AI initiatives within a company. Okay, then you probably should have a team that is responsible for taking in all those use cases, evaluating if you can use AI for this or what kind of approach you should do, instead of having different teams within a company doing their own approach. Because that's the thing that I've been seeing most now is that within one company they have so many different tools and frameworks that they use instead of having one standardized approach and also being able to share the expertise across. So one company I was talking about, they didn't know what the other part of the company was doing. So they had two or they had many different tools they were using, even though they were essentially doing the same. I was like, okay, have you been have you been talking to those guys there? What they're using AI as well, oh didn't know that. So there's uh there's a lot of that going on. Any questions from yeah?
SpeakerI want uh so you talked earlier about uh the person in the loop uh and how you use it to kind of summarize what you make decisions and what the people make decisions. How do you identify whether that summary is actually correct or if there's any like mistakes in it? How do you go about verifying that?
Marius SandbuSo when I'm building those, let's say, but decision or business decision-making summarizations that I want to be able to understand the context. Sometimes I use uh set of multi-agents, so I have one agent to try and verify the content from the first agent. That's one way to approach it. But in some cases, I'm making sure that the prompt of the agent is trying to, well, I do I spend a lot of time on prompting to making sure that okay, this is the description. I need to also adjust the uh some parameters and the language models as to make sure that it doesn't do any creative thinking. I want to make sure that it sticks to the facts. But I also want to make sure that, okay, the agent is generating a reply, and it also links to the source. Okay, this is the content where I found, or this is this is where I found the content, and this is the analysis that it did based upon that content. So it's uh well, it's uh it's a couple of steps. Adjusting the parameters of the model, adding multiple agents to actually verify the content, and then I have links to the source to make sure that okay, this is the correct reply. But there's there's no way to make uh well make it 100% foolproof, right? I still need to have a person in the look to actually understand, okay, is this bollocks or is it actually the facts?
Robby PeraltaI would assume that uh that building that trans, what's the word for it, um, to be able to audit it afterwards? Is that how because you gotta log everything, right? Yeah. So that that that has a cost.
Marius SandbuYeah, yeah. And that's another issue because we haven't had many tools that are available to actually do that type of auditing process, right? Because let's say you're working with an insurance or bank, and then you have an agent that generates a reply. Okay, you can apply a loan to that user. And uh, this is based upon an agent that looks into the person's history and finance history, credit score, and and other stuff that it might find. So, of course, having that audit log of what the agent actually did to actually generate that reply if something comes up a little bit later. But we haven't had good tools to actually do that uh in most cases. We're getting more and more. Uh but we've I've been using some tools. One of them is called the Langsmit, which is from the creators of Langchain, and we had another one called Century, which is uh connecting to the orchestrator so it can see okay, what kind of decisions does this agent do? What kind of API calls does it do? What kind of data sources does it talk to? So that way we have like traceability on what's been going on. But again, it's back to the same that it's a fairly young ecosystem. So it's uh difficult to try and uh map all the missing pieces when they're building an agent together. So much to remember. Yeah.
Speaker 5Any any questions? Yes?
Speaker 8One might say that uh there's a lot of hype going around about identification AI. What is the biggest uh myths being sold to the public right now about this whole thing?
Speaker 2The biggest the biggest myth being sold.
Speaker 8If you went to arendalsuka for example, all the companies, all everything happened, all the talks were about like, yeah, we should do more AI. We need to have a tongue slipped.
Speaker 2Yeah, yeah, yeah. Yeah, yeah. Yeah, and I think so.
Speaker 8What is the biggest mess around this? People don't understand.
Marius SandbuWell, uh it's gonna be a complex answer to just to start with that. But I think that first and foremost is that uh we see that many are looking into okay, we need to use AI, we need to use AI now. And I think the bigger question is, okay, how should we how can we use AI correctly? Because people overestimate how much this type of technology can actually do. They think they can just, okay, we have a business issue here, okay. We just put AI on it, it's gonna solve everything. And we can go back drinking coffee and sitting on our chairs. And the second part is that um, well, what one one of the things that I see is especially when it comes to data governance, companies have been neglecting the way that they manage the data for the last 20 years, and they think that okay, just add AI to it, it's gonna solve everything. No, it's not. Need to do a lot of work there, and there's a lot of boring work that needs to be done. Um people are overestimating how much uh how what the technology can do.
Speaker 7Shit and shit out.
Speaker 2Yes, correct. Now in agents as well. Shit in and shit out agent. Any other questions?
Speaker 1I've heard some people say that uh AI is the solution looking for a problem, kind of. Uh so what I'm very interested in in is uh do you see any good use cases in security for uh particularly LLMs and agents in general? Like what can we do in the next few years until we get like that singularity or something?
Speaker 2Yeah, the singularity is still a far, I think it's far far away. Yeah. But okay, just a couple of things. Uh just code, just if we look at the security perspective before we go into agents, just code scanning. You have this obscure script, which is 250 or thousand lines long. Okay, you probably have a really good understanding of code, or you can just add ask an LM to make a summarization. What does this code actually do? And I've used it for a lot of different use cases. I see a lot of value, saves a lot of time, takes five seconds, then have a good explanation of what the script does, uh, if there's any malicious intent in it, or what it actually does. The second part is also there's a lot of reporting going on in security, right? You need to write, okay, this is a root cost analysis of a security incident, or you need to give a deeper, deeper understanding on it. That's also a great way to use genitive AI agents, right? To give a good text summarization of a specific instance. Because I don't think any in security likes to write documentation or long text uh summarizations, correct? Yeah, everybody's nodding. Personally, I love it. Yeah. So of course, using genitive AI for those purposes, I think is a really good use case. And also having those agents that are running in the background continuously and trying to find new things that you can do to actually tighten your security even more, right? So not using like uh like an agent that's going berserk and just locking everything down, but look at your entire environment as a whole and look at input from different systems and give feedback saying that, okay, I've looked at your organization from a technical point of view. This is the things that you should be doing. And just to give an example, and one of the things that I did for one customer because they were doing an audit of their uh cloud platform. So I had one dedicated virtual machine, I had two MCP servers that are used. One of the MCP servers was talking directly to the cloud platform, and the other one was talking to the ITSM system. So I said to the MCP server, okay, get a list of all security configuration policies that you have, write uh in a full text report on it and post it into the knowledge base of the ITSM system. So it took me five minutes and ten minutes if you include a coffee break, but then I had like a full report on what's been configured, what's been going on, and then can get to work on actually tightening those screws or tightening the uh security risks that was there. And of course, when these models now will be able to support even more context further down the line. So we have a new model from Meta which supports up to three million words. Okay, suddenly we can just paste an entire application system code in there and ask it to look at find any vulnerabilities or compare it to the OWASP recommendations. So there's a lot of different new use cases coming in now and using from a looking at it from a security perspective.
Speaker 5There's at least 650 companies at RSA that said that it's gonna revolutionalize their uh your life.
Speaker 2Yeah, that means you're gonna save a lot of time writing text. That's revolutionizing, right? Yeah. Yeah. Unless you like to write. As I do. Any closing thoughts? No. I just want to say that uh this is a is a really interesting space, but it's also important that you stay curious on what's going on, because even though we have, well, most cases they overestimate how how much of what it can do, but sometimes we also see a lot of companies are haven't seen how far it's gotten in a fairly short amount of time. And I think that the next six months, if we look at this, uh having this podcast, maybe next year this agent is gonna be all over the place. So also want to make sure that you have governance on those agents so not gonna pop up everywhere.
Speaker 5But maybe it's smart to start very something small just to get a grasp of what it is and what it does, and then you know that whole concept of digitalize or die that was like 10 10 years ago? Like companies wouldn't exist if they didn't completely, you know, digitalize a business. Do you think that's that's the same with agents and yeah?
Speaker 2Well, I think that's an important factor when it comes to if you want to be able to compete in a market that you okay need to understand at least okay, how can we use AI? Uh not for AI's sake, but how it can support our business. But again, I uh it's important to at least try and stay up to date on what's going on for be able to understand how it can impact our business.
Speaker 5Well, if you ever need help building or thinking about an agentec AI system, Marius would uh love to hear from you, I'm sure. And uh with that said, thank you all for taking time out of your day to be here with us. And thank you so much for joining. Thank you. And uh enjoy the rest of your Sea Cats Festival.
Speaker 3Well, that's all for today, folks. Thank you for tuning in to the mnemonic security podcast. If you have any concepts or ideas that you'd like us to discuss on future episodes, please feel free to hit me up on LinkedIn or to send us a mail at the podcast mnemonic.m. Thank you for listening. We'll see you next time.