CX Today

CX Leaders Can’t Ignore This Agentic AI Lesson From the Pocket OS Outage

CXToday.com

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 26:44

A production database deletion and 30-hour outage at Pocket OS has become a cautionary tale for any organization bringing agentic AI into critical systems.

In this interview, Nicole Willing speaks with Alex Gallego, CEO of Redpanda Data, about what went wrong and why blaming the model misses the point. They unpack the enterprise basics that still get skipped: separating development from production, tightening authentication and authorization and scoping agent permissions away from dangerous root access.

Gallego also explains why governing agents is harder than governing humans. Agents can be more capable, less predictable and execute at machine speed. That combination makes traditional controls feel too slow. Gallego outlines practical guardrails, the need to reduce the “search space” of tools an agent can access, and why many organizations will need proxy layers to enforce fine-grained permissions when legacy systems cannot.

The conversation closes with the bigger stakes, from customer trust and regulatory exposure to the risk of catastrophic failure in human-critical systems. For CX and contact center leaders evaluating autonomous agents, the takeaway is clear: speed and automation are only valuable if governance is designed in from day one.

For more Customer Experience tech news visit https://www.cxtoday.com

SPEAKER_00

Hello and welcome to CX Today. I'm Nicole Willing. If you're in any way following Ingenic AI or customer experience systems, you're likely aware of a recent incident involving Pocket OS, where a Cursor AI coding agent based on Anthropics Claude deleted a production database and its backups in a matter of seconds, causing a 30-hour outage. It was a dramatic failure, and it may not be the last time that happens as more organizations deploy autonomous agents into critical systems. To unpack what it means for enterprise deployments of AgenTech AI, I'm joined by Alex Galago, CEO of Red Panda Data. Alex, it's great to have you here.

SPEAKER_01

Thank you for having me. Excited to be here.

SPEAKER_00

Great. So what was your initial reaction, you know, when you saw that an AI agent deleted an entire company's database?

SPEAKER_01

My first was like, I wish it was an anomaly. And I wish, you know, these agents are so like psycho psychophantic. What they're trying to do, they would always prioritize like pleasing users over, you know, sometimes doing the right things. And I think maybe that's just like ultimately the difference between enterprise agents and what tends to be consumer agents. So especially for customer experience and you know, the the people that tune in and listen to your audience here. It's like the the main space is like, how do you govern, how you reduce the chaos, how do you tame these agents uh into or coerce them into doing the right things as opposed to going out there and exploring things. And so my first reaction was like, wow, this this is terrible, right? Like it sucks to be a CEO responsible working with these small businesses, and then losing your data, like that, that's just terrible. But then the follow-up thought is there's like the world doesn't really know yet how to govern these agents. And so I think we're all trying to figure this out together.

SPEAKER_00

Yeah, absolutely. I think that's a common theme when I'm speaking to people, is because it's so new, everybody's trying to figure it out, and things um, you know, like this are happening because it's still a learning process.

SPEAKER_01

Totally. And and so what's interesting, well, there's multiple parts, I think, to the the Pocket OS outage. But in particular, we can just sort of like uh maybe take a step back and try to like analyze what were the, I don't know, perhaps the systemic problems that, you know, whether it's Pocket OS or any other enterprise um are gonna face. And and and so one, you know, I guess in Europe we'll also have a third dimension, which will be um governance. But, anyways, the the core problem is actually needing to separate your production from your development environment. And so um cursor is an application that effectively is a way to write code um faster. Um and so it's used by engineers, right? So there was this whole evolution from like uh completion of code into now you give this coding agents like cursor or open uh agent. There's like, you know, there are many examples here. And instead, you're now able to give these tools, this high-level task, like build me a marketing dashboard, or build me this like subset of accounting agents. And so you give this like end-to-end um goals to codex from OpenAI is another good example. But then you have to give it access sometimes to critical systems. And so what is missing is really this like big separation between the way that you develop code and the way that you run code. That has to be just general hygiene for really enterprise or any company doing enterprise agents. And then the second one, which the CEO points out during his post, and you know, he was super frustrated on Twitter. He's like, can I please just get some help? Um is that by definition, most systems were built before uh agents, right? Because agents are this new thing that probably became popular over the last you know six to eight months, really. And and so um, and so they were built in an era before agents, and therefore sometimes the authentication and the authorization and the fine access controls to resources like Deline in the database aren't scope. And so you have this like wide permission systems, you know, typically called root permissions, where a token could do anything. And so, anyway, so there's multiple parts of this, and happy to dive into whether is like, you know, what are recommendations that we tend to give to enterprises on like the separation of concerns from development into production, and then perhaps once you take an agent into production, like how do you ensure that the permissions that these agents have access to are like audited, are constrained, scoped down in, you know, in other words, if the end party system, I think in this um way, like Rails, Railsway, which is uh basically their hosting application, I believe, um doesn't have the the permissions are too wide, then you need to introduce another system that scopes the permissions so that you don't run the risk of deleting your customers' data.

SPEAKER_00

Absolutely. Yeah, I was gonna ask exactly that. That, you know, um people tend to blame the tool or the AI model when someone or the agent when this happens. But like you say, it's really about the permissions that can access, isn't it? So when companies are deploying an AR agent, how would you um advise them to approach making sure that the access isn't there for an agent to go and do something completely unexpected?

SPEAKER_01

I we and I wish the answer was as easy as like here's a painkiller. And and and the reason is like, you know, it's it's a little bit more nuanced. So let me give you a detail because I think it's important, especially for those listeners that are running technology companies that are thinking of introducing agents or like agentic workloads into their applications, right? There's basically this massive demand from users. Like, you know, when I was on my way over to Milan, I was like, oh, I wish um, and like my A was sleeping, it was like one in the morning. I was like, I wish I could be talking to an agent to like do some travel last-minute adjustments. There's this unlimited demand from users to be able to do more, to be more flexible, et cetera. There's also this massive pressure, whether it's from Wall Street or investors or whatever, to understand like what are the, I don't know, perhaps at the baseline, this workforce implication of optimization, right? I think like it often is is like uh blended as whatever, some lame marketing messages, but like CIOs are very much thinking about cost optimization. And the first level is, you know, are we could we optimize the way that we're operating our baseline costs? Are we too fat? Is I think is a question that people are asking. So, anyways, so to the root question, which is how do we help people govern access to private data, right? That that really is the key. Like public data, okay, fine. You you could do that too. But the the like what makes a bank a bank is not that you have two check-in products. What makes a bank a bank is like the customer relationships. Every bank has a check-in product, most banks do. Um and and so uh, but but those specific customer relationships, that is this is what makes unique. And that translates to effectively your private data. You know, what user is in your particular Salesforce instance, what user is in your particular accounting system, what is the balance sheet, and so on. And so when you're trying to give agents access to private data, you really have to think from first principles, which is what permissions do this agent have access to? And how am I going to govern access to this data? And so, like, yes, it is obviously carrying the right uh uh authentication, like, are you a valid system and then the right authorization? Is this user therefore allowed to see this database stable, you know, this this credit? And so the classical example I like to give is I want to go into my credit card application and I want to up my credit limit. Yeah. And so the number of systems that an enterprise agent has to touch is, you know, maybe 10 or 20. You have to go to the credit bureaus, you have to look at your account, uh, you have to look at your spending balance and the spending trajectory and whether like your customer true reference table is in Salesforce or a particular customer data set, right? There's all these systems. And so you're handing off this piece of data that, by the way, weren't really invoked by me. It got started by Alex. True, right? You have to carry what is known in the industry as like OBO on behalf of permissions of Alex, you carry all through those agents, and yet that's not enough. Because in the case of Pocket OS, like some of these systems have root permissions. And so even carrying, like, I'm the CEO of Red Pendant, so I have a little bit more access than other people at the company, but I want my agents to be doing tasks on my behalf that are scoped to a much lesser like authentication scope. And so then really lies the problem with uh enterprise access control to these agents, um, which is gosh, like, how do I ensure that both my legacy systems and my new systems are working in a way that makes you ensures that your agents are accessing your private data responsibly?

SPEAKER_00

Yeah, that makes a lot of sense. Because obviously with an employee, you manage their permissions, right? But then yes, with agents, they m you might want them to do tasks across multiple things that maybe a human wouldn't have permission across those multiple systems, but you want the AI agent too. So how do you manage that? How do you get the balance right there?

SPEAKER_01

Yeah, well, it we we wrote a paper, um, I think it's about to be published, which is called like we we we all hired agents but forgot to onboard them. And so it's exactly that thesis. It's like it makes sense. And you know what's interesting is like the the crux of agents and the problem with governance, and I want to give you a little bit more intuition, which is like the ability to understand what's happening with these agents, is that uh they hallucinate. But by the way, humans also hallucinate. So it's like that's not enough to say, like, well, agents are hallucinating and are telling you to put glue on your pizza, which is a real thing. It happened, right? So there's like there was this massive uh, you know, X thread where somebody was asking, like, what is the best topic for my pizza? And I think ChatGPT responded, like, you should put glue on your pizza, right? And so um, there's a hallucination piece, a problematic, especially for enterprise agents. Uh, you want to reduce the number of hallucinations. Humans do it too. The way humans do it is when a peer or a colleague or someone says, Hey, can we work on this? Often when people say, I think I heard you say X. And sometimes it's actually, by the way, totally separate and different. You're like, what? I didn't say that. Or maybe that's to me every now and then, you know, like I don't know if if that's the case. So there's a hallucination, so so there's that. Um, but then uh, and so agents are more capable and less predictable. So that and the speed at which they execute this is at machine speed. Sure. And so it is the combination that um that you have to contain them, and that the fact that they just simply, you know, AI behaves like a watermark intelligence is I I think my my mental analogy of how I tend to reason about it, AI. And so it's like if you have a coding agent, if you have an AI that has, I don't know, let's say uh workforce intelligence of an accountant level two, it is also likely, or you could, you know, build an argument that it could also be likely that that uh AI or that model may also be a level two coding intelligence and so on. So you have this watermark level intelligence, it means you have like a multi-domain expert with the ability to write code and the ability to execute code with root permissions. And so you're like, it's a lot of very dangerous things. Yeah. And so what we tend when we work with CIOs, and so, you know, like as a CEO, I tend to work with some of the fortune, you know, 2000s or some of the largest companies in the world. And um, when I go on, I advise CIOs, like, hey, how do you reason about bringing governance to uh to agents? And so, of course, is the authorization and the authentication piece, but largely the architecture is about reducing the search space of tools. In other words, it's like it's about increasing predictability, reducing hallucinations, ensuring that you're minimizing the entropy of the systems, right? Predictable tools, like, and so that is very different from what tends to be agents with consumers, where like the search space of tools is very large. Like, you know, give it access to like open claw would be a user productivity boost. So it has access to, I don't know, Google Drive and my Gmail and Calendar, maybe even Salesforce. Enterprise agents are the opposite. You're like, agent, you have access to three tools. And by doing it, we're gonna audit and govern every single access to those particular tools. And so, in summary, it's really um reducing technically the search space of tools, so end tools rather than user-specific tools. And then for each of those tool access, ensuring that every agent has the minimum amount of access to those private data at that particular point in time. And in the absence of the back-end system having fine-grain access, then you need to introduce a third-party tool like a proxy that then enforces a much uh narrower permission and system. And so those that's kind of like the high level and the high-level takeaways of how I would reason about bringing governance to agents for access controls.

SPEAKER_00

Yeah, absolutely. And you mentioned that it all happens at machine speed with an agent, right? So then that kind of requires a real-time control. So, what would that look like in practice, you know, on the data layer?

SPEAKER_01

Yeah, so I would say there's two parts to like um what I tend to call the agent killed switch. Okay, so so let's talk about a customer experience. So I wanted to rebook a flight and what I remove some hotel, standard travel agent thing, nothing sophisticated here. And so there's the synchronous piece, which is like let's say an agent is misbehaving, and instead of like uh canceling the previous reservation and moving to the next reservation, it just like books five hotels because it's the easiest thing. You're like, for sure, Alex will be covered, right? Remember that their goal is to please the user, not necessarily do the right thing under like what a human, the human judgment part, the onboarding, the career training, it's it's all missing, right? As we talked about this, it's like we hired agents but forgot to onboard them. Right. Um, and and so there's this synchronous part, which is you need to build both positive assertions and negative assertions into your agents. And so this tend to be called hardnesses. This could also be called uh guardrails. Um, and so we'll talk about both explicit guardrails like toxicity rules, making sure that agents don't curse at customers, be like, hey, whatever, like you're dumb. You probably don't want a customer success, you know, agent to be saying that to the CEO that's buying, like, you know, whatever, trying to use your your travel agent. So um, okay, that's so that's the synchronous piece. You're keeping taps of the positive assertions and the negative assertions. And so, in other words, like yes, the agent is being helpful, and no, the agent is not being helpful, like, or is simply deleting all of your traveling tickets, right? There's like you encode this per agent. This is the job of the engineer and the product ultimately is to set up this like um this range, this guardrails. Like within this box, you could be creative, but only within this box. And depending on how critical you may like reduce the the size of the box. You're like, you can only do a very small subset of things. Uh, so that's like the synchronous rules. And what's uh easy about that is they're largely well defined. For a credit card application, is you should approve credit card, you know, in America for people that are, let's say, above 800 points, credit points, right? Or people that make certain threshold of income per year or whatever. You could define like a lot of this rules. So that's easy, and you still need to implement it at a scale, which maybe that part is harder. You can partner with someone like Red Panda or others for that part. Now, the asynchronous part, that's the part that I think gives enterprises anxiety. Let me give you an example. In Self, in customer success agent, that is usually a turn-based agent. So either in user-facing agents or customer uh um yeah, or like B2C agents where the user takes a turn, the agent takes a turn, the user takes a turn, the agent takes a turn. It's relatively straightforward to govern access to one agent. That feels easy, that feels like tractable to an engineer. You just basically observe 100% of the things and you can, you know, glue it yourself, and maybe you could do a great job of that. I think I think there's like a high likelihood that an engineer does a great job at that. The problem with enterprises is you have agents talking to agents, talking to agents, talking to agents, and some of those transitive agents are not designed or engineered by your team. They may, but they they don't have to. And so that's when it gets super complicated. That's when you need to introduce like this idea of uh what we call an agent kill switch, which is you need to be constantly monitoring the behavior of the agents such that when like the end-to-end user experience is poor, like it turns out this agent cancel 52 bookings over the last five minutes. You're like, I don't know what's going on. There's probably a prompt injection attack on this thing, is leaking tokens, trying to access whatever, right? You can observe this end-to-end user experience and then you can decommission it. And so for that, the advice to CIOs often or CTOs, it's usually this kind of decision tends to sit in the office of the CIO or the office of the CTO, is to set up effectively a durable log, a way to capture all of these immutable events to reconstruct them and then make a timeline and say, actually, this agent is misbehaving. I'm going to decommission because I tell you what, I would rather not go broke as a company than approve every single credit application that comes through through the banking website.

SPEAKER_00

Exactly. Exactly. And you know, right now these incidents, you know, with AI agents, they might result in an outage, as in this case, but you've warned that it could involve something bigger. So what potentially could happen?

SPEAKER_01

I mean, I think in human critical systems, as this agencies start to move into those obviously in the US was this great debate around the politics and the ethics around using models for weapons, right? Yeah. And and so there was a strong stance from model companies that says you sh you cannot use the models for mass surveillance and you cannot use these models for automatic weapons, right? Like they there's no sophisticated because of this erratic behavior that we're talking about, right? And so, like, as it also feels like inevitable that models are gonna make it part of this decision tree. Maybe they're not, maybe that human is still in the loop to like approve uh, you know, let's say a particular whatever defense attack or whatever you want to do. Um and so so even if that's still the case, it is likely that the information that is fed to the human is still massaged and processed by you know by these models. And so I I think like, you know, probably now in systems that you and I have no access to, um, it, you know, that this this is the case. It just feels like logical that that you know basically these governments are in a race against uh each other that they're all tried to protect. Like somebody's gonna come up with with a set of systems that is gonna be superior. And so I think that's ultimately the most catastrophic decision. Now, for companies, it'll be customer trust, turn, loyalty. It may be like fines. I know, you know, the uh EUAI Act, which I think is going live in in August or something like that, somewhere around that time frame, which says for all of these determined like critical systems like credit card application, utility bills, internet providers, etc., you have to record like all of these agents' interactions just to make sure that you know what's a good example would be like you don't have a racial profile. And so, you know, like if uh if you're a woman or if you're a Latino or like your your particular ethnicity, like you're not getting nightless credit card applications, or at least you can like detect it and make it better. And so uh so I think those those will be like as AI continues to move into a way. From what I think tends to be the dominant use case today, which is internal productivity. That's kind of what more most by by frequency count of agents, the vast majority that I've seen are internal agents to help employees be more productive. Like, you know, it what and so when they start to move to more external facing, where the core of your product offer becomes agentic, becomes is a way of interacting with the world that is different, that is executed in machine speed, that is like, you know, being helpful in particular ways. I do think that the risk is all of the above. It's like it's fine. It could be possibly catastrophic to human lives if you know deployed incorrectly, if you don't have the right guard layers. It could like, you know, basically lick all of your trait secrets. And so the risk, and the main reason is what we talked about is they're less predictable, are more capable and execute and machine speed. It's like it's sort of like the combination of those three pillars that make them so useful, obviously, but also, you know, basically hard to govern and and risky. And part of the reason why we're having this conversation.

SPEAKER_00

Absolutely. So then what needs to change now to kind of prevent these failures from scaling into this broader risk?

SPEAKER_01

So, as we are partnering with um basically engineers at this this company's building these agents, that this is exactly the question that often the board comes and asks a CEO, and the CEO is like, you know, tech team, we need to deploy some agents because of whatever the price of the stock. And and and you know, not not to be facetious, there is obviously intrinsic value in being better, in being like lean and and like, you know, using your resources intelligently. It's like the job of the executive team is like, how do we make the best use of resources? They're largely capital allocators, you know, when you get to a certain stage. So, anyways, at the technical level, it is a mindset shift away from how we used to build applications into a totally different world that requires a new way of thinking and a new way of tooling. And so, if I were to have advice to people, that is not that uh that like you're writing the code. That part is easy, right? So the cost of generating code is the easy part, is that the architecture to run these agents is fundamentally different to the applications that I grew up building. The way that we you grew up building these applications, it is not the way it is. The architecture looks different. And so it's just being open into the fact that that is just, you know, how the world is today. And that um, and and so you need a totally new set of primitives from the way that you deploy code, from the way that you optimize code, so that when you bring these agents to production, they aren't deleting your data. They have the right permissions. And by the way, when something goes wrong, you can actually go back and figure it out. Yeah, and so I think it's only a matter of time until either it is forced on companies by governments, like you need to ensure that all of your decisions are audited across for the next six years or whatever. Like that's still the case, obviously, with the EU AI Act. Uh, or just because, like, how else do you debug a thing when something goes wrong? Like, how could you possibly figure out what went wrong when the agent itself could mutate the state that you know that governed it? Like that is sort of the gnarliest thing. These things are relatively capable. And so um, that would be the advice is like for people to be open to a fundamentally new way of building applications, and I think most are, um, and to rethink what it means to scope access to these agents and ensuring that you have infrastructure uh that allows you to go back and reconstruct all of the decisions that agents took while interacting with customer data.

SPEAKER_00

Yeah, absolutely. That's really important advice. So I want to thank you, Alex, for joining us. Really appreciate your insights.

SPEAKER_01

Thanks for having me. I appreciate it.

SPEAKER_00

And to our viewers for more interviews and articles covering AI infrastructure, security, and customer experience, go to cxda.com, subscribe to our newsletter, and join our community on LinkedIn. Thanks for watching.