This Week in Leading AI
Imagine two mates at the bar. Thirty years of business between them. And all they want to talk about is AI.
That's "This Week in Leading AI". The podcast where Kieron and Neil cut through the hype, share what's really working in the world of Generative AI, and helping people figure out this AI thing without the techno-babble.
Just honest conversation, real stories from the AI coalface, and the kind of straight-talking advice you'd only get from people who've worked together for 30+ years, been there, done that, broken things, gone "Oh S***!, fixed it, and lived to tell the tale. They claim Leading AI is the best job they've ever had and are having a blast doing it. It shows.
Warning: may cause you to actually enjoy learning about AI
Pull up a stool. We'll get the beers in.
This Week in Leading AI
Su Belagodu and The Human in the Loop
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Six legs on the pantomime horse again — and this time we had to explain what a pantomime horse actually is. Su Belagodu joins us from Winchester, Massachusetts (yes, we compared notes on the other Winchester), introduced by our previous guest Nicole Alos. Her first question after the explanation: "Are you the head or the back?" Kieron's answer: Neil's very firmly in the front. He's at the back, but driving.
Su started in computer science until a manager told her she asked too many questions and wasn't a very good coder, so she should go into product. For Su that was a blessing in disguise. She's led healthcare startups, co-founded an AI-native company in 2024 when she didn't yet know what "agentic" meant, and now advises companies on the thing everyone gets wrong.
The warm body problem. Everyone knows you need a human in the loop. So organisations drop a programme manager at the end of a workflow, tell them to approve things, and call it governance. Su's research, built on 40-plus interviews into a Human in the Loop Maturity Model, found what actually happens. After the first 20 or 30 approvals, confidence in the AI rises, fatigue sets in, and people start rubber-stamping. One human at the end of a workflow isn't oversight. It's poor design.
Trust, quantified. Accountability × transparency × accuracy. Accountability means someone answerable who can actually change things. And the answer can never be "it was the AI." Perhaps her most incisive line in the chat was that enterprises aren't buying AI, they're buying confidence that AI won't add risk. That's why they don't adopt. Not cost. Not scepticism. Risk appetite.
One-way doors and two-way doors. Taco Bell's AI took an order for "lots of water" and added 99 bottles. Annoying, reversible. Get a gas boiler enquiry wrong and it could have life-changing consequences. Design the checkpoints around which door you're walking through.
Su also talked about pathology study where humans scored 97%, AI scored 98%, and the two together hit 99.8%. That's co-intelligence in one number.
Safeguarding is a big topic for us and today Su explains why she tells her sixth-grader's AI class don't humanise it. Alexa and Siri trained us into blind trust, and no chatbot should replace a teacher or a counsellor.
Management agents that supervise swarms of other agents (a callback to Kieron's son and his arguing bots). And Su turns the tables at the end to ask what AI adoption actually looks like in the UK, which gets us into the sovereign LLM problem, the Azure Foundry queue where Sweden gets everything first, and Kieron's line that deploying Copilot as your AI strategy is like handing everyone Excel and calling it data analytics.
We've promised her a return at Christmas. In an actual pantomime horse outfit. She says she'll hold us to it.
Two mates and a third stool with a new friend on. A bar. And all they want to talk about is AI.
Pull up a stool — we'll get the beers in. 🍺
Right, well, in that case, let's get this pantomime horse of a podcast underway. And this week's podcast, there aren't four legs on the pantomime horse, there are six legs, because as anybody that will see the um video will be able to see, we have with us a lady called Sue Belagodu. Did I get that right, Sue? Excellent. Sue, welcome. Thank you very much. But um uh we'll we'll we'll talk in a second, but I I need to explain something because when we were talking earlier, you um you live in America and you don't actually know what a pantomime is, much less a pantomime horse.
SPEAKER_03Very wise. Very wise not knowing what it is. It's not uh yeah, I think it's quite quintessentially British, I think, a pantomime.
SPEAKER_00I think it is it's a Christmas thing, and it's like a it's like a jokey play. It's on in lots of the theatres, it's tradition to take the kids, there's lots of uh booing at the baddies and cheering of the goodies, uh lots of dressing up, lots of songs, lots of really bad jokes. But one of the things in these pantomime um uh theatre performances is something called the pantomime horse. And effectively you've got uh somebody in the front who's got the two horse's legs and then the horse's head, and then somebody that has to be the back end, so there's two legs, but they've got to bend over so that their back becomes the horse's back. And I've always thought it's gonna be really quite uncomfortable in either of those because you'll be like a straight back if you're gonna get back, and and I can my back now, because I my back would hurt so much if I was in the back of the pantomime horse. So um so yeah, it is uh if you ever get the opportunity, uh if you're in the UK to go and see a pantomime, I'm sure you won't be disappointed. They're always a lot of fun. Uh but yeah, it's a Christmas thing, and um, I think you know it because it's because it's uh just coming up to August, they're bound to start being prepared for Christmas anyway. So yeah, they'll they'll be getting ready. So if you feel the need to appear in a pantomime horse, if anybody ever asks you to get into a pantomime horse outfit, just so no, just is my outfit for you.
SPEAKER_04That sounds fascinating. Uh, I will check it out. The only thing I've done in London during Christmas is visit Oxford Street. I definitely do more, but this is fascinating. Now I'm going to not stop thinking about how uncomfortable those two people are. And what's the fine horns with six legs? Now I can't even fathom that. Like, oh my dear goodness, like who's sandwiched between them now?
SPEAKER_03That's right, it gets worse. Very good. But welcome, Sue. It's uh light to go on. Sorry, sorry, are you speaking?
SPEAKER_00I can't go ahead, Kieran. Go ahead.
SPEAKER_03Uh so yeah, I'm very key. You work in AI adoption, but I'd love to hear more uh how you position uh what you do. But I'm fascinated, it's one of the biggest challenges we have in what we do is encouraging clients to really make use of tools that we provide to design tools that do things that they need. Um, but yeah, getting people on the journey is really tough. So I'm fascinated to explore that with you in this session. But tell us about your background. How have you ended up where you are now?
SPEAKER_04Thank you, Kierin. Um, I started as a technical person. I have a computer science background, and then I asked a lot of questions. And so my managers all along said, Hey, you know what? I think you belong in product because you care about what the customer wants, you're asking a lot of questions, and frankly, you're not a really good coder. So go to the product. So I I mean, you know, blessing in disguise because I loved doing product and designing how systems should look. I then led company, I led startups in the healthcare space, became a product leader, and essentially that means, you know, I'm the fall man if something goes wrong. And after many years of doing product leadership, I would started advising. I didn't know people pay me for advising, and I was like, I love it. I mean, that's my side hustle now. And so once such advising gig turned into me becoming a co-founder of an AI native startup, and this was in 2024, and I didn't know what agentic meant then. I was just offering product strategy to that company. And my my reason for joining that absolute startup was I wanted to learn more about AI and how to build with it. Fantastic journey. The technology, AI technology blew my mind. What would take us six months took us like one month to build an MVP. We were doing really well with small businesses, and then that's it. There was like a period. We couldn't progress to bigger companies and enterprises. And that's when I started to realize that all right, you know what? Just automation in AI is not great. We need a human in the loop. AI could give you speed and breadth of things, but then the human should be the one guiding. So once we adopted those principles, we were able to crack the enterprise deal. Uh, the company, the startup did excellent. We got funded by a VC company and they moved to San Francisco. I live on the other coast in Boston. And I found my love in helping other companies navigate through the same journey that we did. And I kept saying, listen, building with AI has gotten very easy, but positioning and selling AI is hard. So led to becoming a co-founder to a solo founder. And that's led me down the path of understanding human in the loop more and not just a warm body in an AI workflow. So, what I do now is I do consulting, helping companies understand how to build good AI products and systems with a human in the loop as part of their design. I teach. Um and yeah, I've also built I've built another small product that's going to be something, but you know, it's still in its nascent stages. It helps companies measure how their AI workflows and AI systems are performing. It's more a compliance thing. Yeah. That was a long pentomizing for us.
SPEAKER_03And there's so much in there. I would I'd love to hear more about the kind of measuring how they're working, but let's let's get into that um in a bit. Um I think that um where you talked about sort of product manager role, I guess that you that's the point where you become responsible for kind of interpreting user wishes as well as sort of actually owning the kind of technical, um, well not owning the technical, but liaising with the technical side. So is that in your product leader role? Does that is that what you're spending your time doing? Is it about users and adoption? Is that how you got into that piece?
SPEAKER_04Absolutely. You're right on the money there. I almost look at as bridging three kinds of roles. One is the user, of course. What are they going to use, not what we believe they'll use? Then there's builders who just want clear instructions. These are your techie developers. You give them a clear instruction, they'll build to it. And then there's typically the founders or the owners of the company. These are visionaries, right? These are idea people. Um, and seldom all those three things don't align. The idea people have big grand ideas, and then your users don't really want to use what you're building. And the builders are wondering about what is it you really want to build? Like you said something else last week, and this week you're saying something else. So the product person sort of sits in the middle of these three roles, uh, helping them understand each other and build something solid that users and customers will ultimately uh benefit from.
SPEAKER_03Yeah. Interesting. And that whole kind of leader's view of what is necessary versus what the team will. I mean, that's just played out through software forever, really, hasn't it? Of like, I always think the kind of software sale, at least one of old, would be they'd come in and do a demo on perfect data with perfect use and and maybe look at look at what we could do for you. And the CEO or the C suite are going, that's brilliant, we need that. And of course, that's when the trouble really starts. And yeah, getting those benefits are tough.
SPEAKER_04Yeah, I think you know, um, the gap that I would see was no one's really asking your actual customers and users how they're using their product. And I think that's when the product role kind of started to take shape, because there was C-suite and strategy consultants who were telling you what to build, and then there were the builders who would follow the instructions and just build it. But who's going to go out there and talk to your customers, understand the market, and actually start to build the right things?
SPEAKER_02Yeah.
SPEAKER_04So it was kind of it was fun because I, at the end of the day, love working with people and I also understand technology. So it was the ideal role for me because I would speak to the users, speak to the founder, speak to the developer. It's it's like being uh not bilingual, but more than that, right? Because you're talking different languages to different people, ultimately trying to build the same thing.
SPEAKER_03Yeah, yeah, I can imagine. It's uh it's like just the impossible task of trying to kind of manage two three sides of a of a three sides of a pantomime horse, perhaps.
SPEAKER_04Yes, exactly. The six-legged one.
SPEAKER_00Have you seen any difference in kind of pre-AI and then post-AI? So your post your post-2024. Is it still the same kind of issues or or have things changed? As anything, as anything different?
SPEAKER_04I I think it's significantly different, Neil. Um, that's that's a good question. I think right when AI became a thing, when ChatGPT launched in 2020, the first use case that everybody could think of was uh efficiency gains. Whatever took the marketing team to do research, say one week, would now take them a day. Uh, if a product manager or a developer needed to write an email to a customer, they would type it, erase it, type it again, erase it again because they're like, wait, we're not saying this right. That kind of miscommunication got aligned in minutes because now they would just feed that in their Chat GPT window. It would tell you how to communicate given the type of user. Um, and so the immediate value was efficiency gains. This meant things got done much quicker, more efficiently. It took a while for the market to look at AI applications as beyond just that. What's after you've gained the efficiencies? And now people have more time. Like you had a person to literally write emails to your customers. Now that can get automated. What next? Right? That's when AI started to um, I think, become part of the solutions you were building. What could we not do before that we can do now with AI? But going back to your question, like uh I think it got a lot better, but also a little chaotic. You know how I was talking about there's a visionary founder, and then there are builders who would just want definitive requirements and they'll build to it. What used to take months for these two to align on what to build now takes hours because you know, a visionary now can use a wipe coding tool and tell you exactly what they want to build. And a product person can take that to a user or even create synthetic users using AI and validate that the very same day. So at the end of the day, you have a detailed requirement document for your tech team to build, and then the tech team uses Copilot and builds that out, the first version of it in a day. And so the whole feedback loop has crunched down to from from days to hours now. So that's definitely happened, and that was great because everybody was like, Yay, you know what? We are awesome. We're going to pump out 100 other requirements. And there was kind of an overload of how much can your user actually use? What do they really want? But no one stopped to think that because we were all so excited about look at this shiny new technology, right? And we're a couple months after that, and I keep saying months because with AI, things are changing every week, with every base model updating itself significantly smarter than the previous version. Um, things are moving really fast. So a couple months from then, people stopped to say, Oh my goodness, here are all these goof ups that's happening because of AI. Did you guys hear the story about uh I think McDonald's rolled out an AI order system and someone asked it a coding question? Yeah, okay. I know, like this is food with entertainment and you know, getting your job done. Like, why wouldn't you use it, right? Yeah, but can you imagine the token maxing that happened there? Like McDonald's wasn't prepared for people to ask a full-on Python solution.
SPEAKER_02Yeah.
SPEAKER_04Uh similarly, there was another one. There is another chain here in the US called Taco Bell. I'm not sure if Taco Bell is um in the UK, but in Taco Bell, someone just keyed in in their order, they put in lots of water and it added 99 bottles of water or something to that effect. And then they had to come back and say, That's not what we asked for. But hey, there wasn't a lot of water, exactly, right? But what was missing there was an actual informed human who would say, maybe in a disgruntled voice, they would say, What does a lot of water mean? Yeah, how many bottles do you want?
SPEAKER_02Right.
SPEAKER_04That's not something the chatbot did.
SPEAKER_03And I think that human in the loop thing is really it it interests us a lot. Um, we talk about because obviously, you know, all AI exact caveats, it can be wrong, you need to check. And then I kind of think to how much can you expect a human to check? We uh a number of the tools that we produce are kind of uh inquiry managers, if you like. It's sort of you know, you an email comes in, the knowledge flow, our platform will read the email, look in the policies and data systems, work out what the answer is and draft an answer. It then goes to a human, and that's how we like it. And we want the human to be able to see the data sources, the uh the question and the draft answer, so they can then make a judgment. But I really worry because the you know the promises, this is all collated for you, and there you go, but you still need to check, and I think it's really challenging to work out how do you help a human to check something, which ultimately is going to become the click of a mouse real soon, of like, yeah, I've checked the first one and the second one, and then I got a bit bored, and then I got distracted and just talking to my mate, clicking his mouse for a while. I think it's a real challenge, and uh as you'll have seen, I'm sure there's um a couple of case law uh now that have been out there with the in Germany. The Google summary, the AI Google summary in Germany they have ruled that that is Google's they own that. So if they're wrong, that's Google problem. You can't just go to AI. And there was the famous case of the Canadian Airline that um it's chat bot said to the some chap he could have a refund, and that wasn't really true. And they and he took it to court, and the court in the US said, or Canada, I guess, said no, you have to stand by what it what it said. So I think it's it's a real challenge, isn't it? Of how do we how do we get humans reasonably checking? And and how do we kind of flag up the things they should check? And it's a real tough one. Do you have a I mean have you tackled any of that? It's I mean, it's really difficult, I think.
SPEAKER_04Oh yes, absolutely. Uh, what I was noticing, so I've used I created this framework called Human in the Loop Maturity Model based on about 40 plus interviews and work with different workflows, AI workflows. And what you're describing, Kieran, is exactly what was happening. They were rubber stamping approvals because there's fatigue. And there was the quality of the output was still slightly better than maybe what if you know, if they had done it without the AI. So what happens is after the first 20, 30 AI outputs that they've reviewed and approved, the confidence in AI starts to increase. And so the human starts to slack.
SPEAKER_02Yeah.
SPEAKER_04And so what we've in my work, what I've tried to do is help companies define metrics and KPIs that are based on human and AI collaboration. Typically, companies say, All right, you human, how many, how many files or emails or drafts did you approve? You know, it's uh like a like a mechanical thing. If they say 100, okay, you met your quota for the day. AI, what was the accuracy of that output? It says it was 80%. Okay, the AI met the quota for the day. But no one's measuring the collaboration between AI and human, right? No one's saying, hey, you uh you approved or you auto-approved less than 5% or more than 5% of the AI outputs. Why did you do that? Like, is there is there a reason for that? Or anytime a human approves it, there needs to be a one-liner reason for why this is or this is not. Uh, or even just measuring human fatigue. If you put one human at the end of the entire workflow and call it human in the loop, then that's just poor design.
SPEAKER_03Yeah, interesting.
SPEAKER_04You know, what it should be doing is checking in with the human at the right times and automating the right steps. So the way I've seen this managed is understanding what sort of use cases need to be fully automated, which ones require a human oversight, where they review something, make a few changes, give feedback to the AI, and push it forward. And which of them absolutely require a human? So there is a confidence threshold that you set for each of your AI outputs, and based on that confidence threshold, you then redirect it to different sources.
SPEAKER_02Yeah.
SPEAKER_04Um, and that's all part of your system design. I, you know, I referred to this a little earlier, which is the warm body problem.
SPEAKER_02Yeah.
SPEAKER_04People realize that you need a human in the loop, and so they would just drop in a program manager or a consultant and say, Hey, start approving these things. You are the responsible human. They're not going to be effective in their role if they don't have a way to shape the AI output.
SPEAKER_03Indeed. And I think that's because it's interesting Gartner's sort of view of the world is that we'll we'll have the have businesses that are sort of an individual running a team of agents. And and in that model, I think it it sounds like it what they're getting at is exactly that. You're responsible for the design, the data that's driving it, how it works, and then obviously its output, so that you, in theory, own it like you would if you're hiring a team of people and managing them to do a task and trying to get people's skills to a level where they're able to do that, I think is a huge challenge for everybody. But I it sounds to me like a very sensible future where it's effectively loads and loads of product owners, but they're responsible for this, you know, little team or swarm of agents if they need that to do the work.
SPEAKER_04There's uh there's quite a few mature systems out there where they have a monitoring agent or the management agent, right? There's one product, um, no endorsement here. This is something I've personally used called uh Wayfound. They create these management agents. These agents manage your swarm of agents, right? They're responsible for making sure that they're working as per design, they've routed these requests based on confidence threshold accurately. They're giving feedback to the human appropriately. So this management agent then manages that team of agents. And so the human now is only responsible for that one management agent. So that's also part of your design, right? You can't have you can't say adding agents will increase efficiency and then add one human for every agent you.
SPEAKER_02Right.
SPEAKER_04Right. So how do you do it more efficiently? Uh, that's kind of how I've seen multi-agent systems work. You need these management agents dropped in. Because can you imagine the damage in a multi-agent system? If one agent gets an incorrect output, it triggers 10 other agents.
SPEAKER_03Indeed, yeah.
SPEAKER_04You know, the output is catastrophic.
SPEAKER_03On the podcast some time ago, I was sharing with Neil my son, who's um 20 years old in that university, and has built a load of clawed agent stuff, which I love and I encourage enormously. But he said he burnt they burnt all of his tokens with two of the agents having an argument. And it's exactly that model. And they did the thing just obviously took off at the speed of AI, arguing with each other. And then so he built a project manager agent that now his his agents aren't allowed to talk to each other anymore. They can only talk back to the project manager who judges, decides what yeah, yeah.
SPEAKER_04And that imagine that in an enterprise.
SPEAKER_03Exactly, with you know, I mean his token burn was I think five dollars or something, so it's kind of like you know, he was annoyed, but um, yeah, imagine when it's like Amazon and it's 500 million, which uh happen happened, so I hear allegedly.
SPEAKER_04Yes, there were a lot of memes showing uh you know money down the drain, literally, because of upsetting. And this again is lack of governance as part of design.
SPEAKER_03Yeah, interesting. I really love that. I really love that th thought process of finding the right moments to bring the human in and check the right stuff.
SPEAKER_02Yes.
SPEAKER_03We have we have a we call them our stepped, our stepped runners in Knowledge Flow, which do that. They they will go through and say, Oh, here's the evidence I can see. Check that. Are you happy? Am I right, you know, within reason, and then okay, I can now move on to the next piece. But they yeah, they have sort of moments in the cycle. But I I think we did that more out of well, it wasn't as cleverly thought through, I don't think, as you have done. I think it was more sort of like that's what's necessary for this task with explainability and governance involved. But I I love the I love that as a principle for product design. That's really sensible.
SPEAKER_04Yes, absolutely. Um, you know, what what I started to think about was especially when we were talking to enterprises or SMBs, they weren't just buying AI, they were trying to buy confidence that AI will not add more risk. That's all they want, right? So your AI model could be fancy, could solve a math Olympiad quiz and whatnot. But if they believe adding that will increase their risk, they're not going to adopt it. And so I started to think about all right, so that correlates to trust. And how do you kind of define or quantify trust? It can't be just a feeling, right? Like all these years, trust is something marketing and brand building does because it's still the feeling, it's invoking a feeling in your users and customers. But when it comes to AI, trust needs to be quantified. And the way I was able to kind of define it for my customers and um friends as well, is I look at it as accountability times transparency times accuracy. Accountability is basically saying who's responsible for it. And the answer can never be it was AI. Is who's going to hide behind that, right? It has to be an accountable person who can make meaningful changes if something goes wrong. Transparency is current what you referred to in knowledge flow, where there's explainability, right? It's talking to your users and showing them why it's doing something, um, you know, showing that reasoning. And then there's accuracy, which primarily depends on the model you're using behind the scenes, like the base model, but also your own training data. Have you trained that model on your data?
SPEAKER_03Yeah, exactly.
SPEAKER_04How accurate is it? And if you do all of that, I think for each one of those steps to increase trust, you need an actual human who has the domain knowledge to be able to do that. So when we talk about trust in AI, you can't develop trust with technology without a human involved. And that's kind of what kind of led me as well to develop this maturity model to help people understand how do you define trust? Because that directly correlates with adoption. People are not adopting AI, not because it's expensive or because they don't believe in technology. They don't have the scope to add more risk. Right. Um, yeah, and you asked me earlier on what is that product that I'm building and how does it help? And that's precisely what it does. Like human pulse is something that plugs into workflows and then tells you the health of your AI systems. It tells you, hey, here's your index. It gives you a score. It's potentially measuring your entire workflow against six dimensions. Something like, you know, how often are you intervening and giving feedback? How well are you training the base model? How is the performance of the human involved in it? So it takes all of that and the trust equation and gives you a score and tells you how to improve it. And the way currently people are using it is there's either a product or an engineering lead who uses it to monitor the AI systems they've deployed in their teams. And then the C-suite has a governance overview, which says, what are all the AI workflows we have in our company?
SPEAKER_02Yeah.
SPEAKER_04How are they doing? Are they doing are they performing well? Are they subpar? Um, and then there is a correlation to the EU AI Act and the NIST AI Act here in the US, which basically says, are you compliant or will you not be compliant? Is there a gap in the way you've built it, or are you going, you know, is your board going to laugh at you when you present this to them?
SPEAKER_03Yeah, indeed.
SPEAKER_04So that's kind of what is.
SPEAKER_03I love it. That's really good. So we so one of the things that we're right now adding into our knowledge flow agentics side of things is uh so in it, we got an email handle I've mentioned already. So the reads the email puts the draft answer in front of a user. The part that we're now doing is if when they edit that, if they edit it, we're tracking what they've changed so that you can classify it as was it a tone of voice thing, was it an accuracy thing, was it you know, wildly off, or you know, so we can then start to see which one which work which use cases, I guess, through the workflow are low risk and never get changed. So we've got data on it as opposed to humans telling you they liked it or didn't like it. The idea being that eventually we deal with, for example, social housing providers here, so um uh, and they will deal with a load of inquiries all the time. One of the biggest ones is can I have a pet dog? And the answer is invariably yes, you can. But they'll get some of the large we deal with one of the very largest in the country, and they get that about a hundred times a day. And so it's like, well, actually, can we let's look at how often it gets changed by the humans currently using it? Yeah, it never gets changed, or if it or if it is, it's so minor. And then there's another thing which I think is really interesting, which is uh a what if check. So if we had sent it without the human changing it, what would have been the impact?
SPEAKER_02Right.
SPEAKER_03And so that's what you say. It's an AI runover of like asking to do a risk analysis. So, for example, if if someone says my my gas boiler in my house isn't working, that potentially to giving them wrong information could end up with an a gas explosion or some horrific thing. So, whereas telling them they can't have a dog when they could have had, or the other way around, is kind of like an irritant, but not you know, not exactly life-threatening. And so you know, immediately the risk is a different level, and then when it's even more minor, like what time do I hand my keys into the office? And it's like, well, you can see it from 9:30, or is it 10? You know, again, you're kind of like in an inconvenience, right? Not catastrophic. So it's a kind of just a risk assessment on what would have happened, and all of that data then can come together to help you inform right, these things, let's push them to full automation.
unknownRight.
SPEAKER_03The question we haven't answered yet, though, which is interesting, is when you push a pets to full automation, how often do you then still check anyway? And we've got to kind of we've now got to work on that because you know it could be wrong, it could be silently wrong in three months' time, wrong all the time. And no one's looking.
SPEAKER_04Listen, it it will or it may be, because it again depends on if the question is has a binary response or not. What if the question is I have a pet, but I'm going to add five more in the next exactly.
SPEAKER_03Or it's a flat, or it's a block of its apartment block, and you're not allowed pets in that apartment block. There are there are nuances, but no, exactly. A building in enough of the kind of um, you know, we use rag AI retrieval augmented generation to make sure it's looking at the right things. But yes, and uh but but that drift that happens anyway, I worry about is that this thing's now automated and running, and you're never gonna know it's wrong until something bad happens.
SPEAKER_04Yes, yes, absolutely. I love that you have that as part of design. Um, that's just the human giving feedback to the AI to make it smarter, right? And early on, I'd done this talk where I said, who really is in the loop? Is it AI or the human?
SPEAKER_02Yeah.
SPEAKER_04We often assume that AI is helping us get better, but it should be the other way around as well. Like it's our responsibility to make sure that the output gets better uh for the AI, right?
SPEAKER_03Um we had so the other side of it, the part of the reason that we built this data-based folk, the data focused, evidence-focused. I've had many a time, and the best example, I won't name the client, but it was an HR team. Um, and we had built a HR assistant, a policy assistant, RAG, based on RAG AI, very accurate. I mean, our tools are really good in that space. And they've they kept saying, Well, we don't like it. And they never could give a reason. And I eventually thought it was always kind of mind the things, like, well, we wouldn't have written it like that. And I'm like, okay, well, how'd how would you have written it? And they'd show you, you know, they'd say their response, and it's kind of just a slightly different approach. And you're thinking, this is not any longer about accuracy of the AI, is it? This is about you just not liking it for all the reasons that some people may not. It's like, hang on a minute, what do I do now? If this is if this AI talk's gonna do this, I'll I do all day. So it's kind of some of that is that if we can get actual data, then we can change that argument slightly into there are better things. There's all these difficult things that does need a human, all this easy stuff we can potentially take away from you.
SPEAKER_04One of the ways to look at it is I think you said it quite well, which is is that decision reversible? Is that a one-way door or a two-way door? If your AI says, like in that example of the number of water bottles, yeah, it's not catastrophic. Or even in your example of a pet, if it says no and then they check with you guys and you say yes, nothing's happening. But if it gets the water heater question wrong, then something catastrophic can happen. So one way to look at it is that is that a one-way door or is it a two-way door? Can you come back and reverse? Yeah, no, it's um, and as for you know, people that do not adopt, because it's not just right, I found that it's got a lot to do with education as well, educating them on how to use it and also that it's not a risk to their jobs. Um, we had a lot of pushback in my first startup because people who are ultimately assessing the tool and purchasing it, we were automating their workflows. And so the C-suite would be like, oh my goodness, ROI, let's go. We need this tool. And then they would have the tech person assess it and they would come back and say, Oh, this isn't quite right, you know, X percent of use cases slipped through your system and whatnot, because they're threatened, right? So you start with the leadership, is what I say. Leaders can convince their teams that increased ROI doesn't necessarily mean them losing their jobs, then that can do really well for them. Because the roles are merging. I mean, that's not something we can deny. Before there was a specialist role for everything people did, and now it's moving towards generalists. And so I think it's a good opportunity to tell people well, if you use this, you now have experience using an AI product, you're smarter for it. And you know, there are other responsibilities you can take up. Um just do this task that's going to be automated. Because if not today, then maybe two years from now, that kind of a role will definitely not be around. Um, and so the sooner you get comfortable using AI tools, the sooner you get comfortable working with AI as a teammate, the better off your early.
SPEAKER_03I completely agree. And I see no evidence of job losses in any of the things we do. Um, you see people being able to focus more on the things they want to do. I often talk about in teaching, we work a lot in education, and teachers, 50% of their time is spent outside the classroom doing all the things that they have to do, in a load of reporting, a load of writing stuff, and admin. Um AI, as I always say, can handle a lot of that, or can certainly help with a lot of that. Here's the big thing is uh 25% of teachers in the UK leave within two years of starting because it's not the job they thought it was. The number one reason they give where they're leaving is because of the 50% of the time they spend doing admin. And that's awful to think that these people have gone through their training, it's been their desired world. They, in the main, enjoy the part that we would think of teaching, standing in front of a class of students, but the rest of it drives them away. And that is a prime example for me of where if we can take away or at least help enormously with all of that drudgery, we can make teaching an amazing job and give people back the time to be amazing at it and focus on what they need to do. So, and it's so true of every job in different ways. Teaching is a bit more emotive, I think, as a subject than, but it's true of you know, a data entry clerk or whatever, they probably don't want to do half of what they need to do.
SPEAKER_04Um, I think it applies to healthcare as well. Like we've been seeing how you typically think, oh, doctors are so well accomplished and they probably won't feel threatened by technology, but there is an adoption issue in healthcare as well, right? Because, but then they hide behind the governance and compliance and all of that, which is probably right, but then you know, you can use it for your admin work. You know, when there's no patient data, you can still use it. And there's been results and experiments that have been run where with say a pathologist reviewed 100 samples and they were 97% accurate with diagnosis, yeah. AI did it independently and it got 98 or something like that. But when you put them both together, they got 99.8% right.
SPEAKER_02Right.
SPEAKER_04So it's that co-intelligence that's really going to push the envelope. And the sooner people adopt AI in their workflows to then understand, hey, this is where my intelligence is required, and this is what I can automate. That's uh that's a behavior change for them as well.
SPEAKER_03Yeah, indeed. And I love so and measuring impact. Sorry, Neil, I'm conscious I'm I'm uh there's last question for me, and then I'll let Neil get a word in edge questions. Yeah, probably not. But so so one of the things I often try and get people to think about is not measuring just on efficiency, but measuring on where it matters. Yes. Um, and in again, education, application processes, so students waiting two weeks to hear about their application, it needn't be because there's just compliance checks, that's stuff AI can smash, and you know, quick oversight. It doesn't take a lot of human because it is literally is that data correct? Does it fall inside the rules? And then a few cases of outside. You could give them an answer within minutes. That's a wonderful thing for a student, it can help them get on with what they need to be doing in their life. And I'm always reminded of the hospital example in the UK here where they were using AI to look at um, I think CT scans. So um, and that process is is you get the results two weeks later if you have a human do it, which is probably not two weeks of your life that you're really happy about. And whereas but the AI was giving the early the pretty much the the the you're fine, you can go home, not worry about it, before they'd got their shirt back on. So you're just you're just getting dressed and it's going, no, you good off you go, don't have to worry about it. That kind of world, I think that's the impact I really seek out is where can it actually make a real difference to people's lives rather than that efficiency, which is nice. I'd like to take away the drudgery from people's jobs, saving people an hour here and there is great, but ultimately the real prize is in my mind is how do you do those transformational things that really make a difference?
SPEAKER_04Listen, if people not everybody sees saved time as a good thing, right? So there's there's always people who would come back and say, Well, why would it take a doctor two weeks and this one did it in minutes? I don't trust its output. You get those semi knowledgeable people who come and say, Have you heard about hallucination? What if this hallucinated? Have you heard that AI outputs can be completely wrong and it's biased? What if it doesn't know how to read my data? And I think building those solutions with governance and human in the loop design can help kind of squash those concerns that people have because those are valid concerns. Like we all know that AI will get analytical stuff right. If it's looking at that scan and all it has to do is based on data points, it'll get it right. But it'll miss the nuances about cultural bias or um ethnicity. All of those points that that's the nuance that the human can add.
SPEAKER_03Yeah. And I've heard in that same exact story, what it's not the same example, I don't think, but where the AI is completely correct about saying you're all clear, but it's missed the fact that actually there was something else wrong. But it wasn't trained it wasn't trained to look at the th that it didn't care about that thing, you know. There's the obviously the human will potentially notice other things as well. So yeah, that is there's a little way to go, I think, isn't there, in trust and and accuracy in the models themselves.
SPEAKER_04But absolutely, absolutely. There are some ways to go, and we can only get there if we use it and then give it feedback. I've seen a lot of people that treat AI as a one-way thing where they just consume it. If they don't like the output, they'll drop it. And I'm like, no, give it feedback. It's looking to learn, it's looking to get better, and the better it gets, the better your life gets. So it's so important to have that relationship with this technology. Um, one caveat because I think you said some of your users and customers are also in education. Something that I always tell kids uh of all ages that are using AI, like my kid is in sixth grade and they have an AI class, they have uh a subject. Yeah. And so what I tell them is don't humanize AI. We already call, you know, we we have names like Alexa, City, and you're you're starting to humanize it. And what that could land up doing is blind trust. So it's good to be, you know, it's good to question AI as technology. And I always say, don't humanize it. It because don't think it's going to give you life advice, and uh you can just rely on that and not include the human that's your teacher or counselor or someone who has good intentions. Don't uh they're they're not a replacement for any of those things, especially teachers as you got up.
SPEAKER_03Definitely good advice, isn't it? Definitely good advice.
SPEAKER_00Yeah, it's something we've pushed uh pretty heavily here, and um we will continue to do so because there's been lots of um uh things in the press, as you know, about uh people using AI for uh uh uh advice and ending and ending badly. Um and speaking of just uh just speaking of ending, uh uh we're coming to the end of our time together, unfortunately, which is really and I've really enjoyed listening to the answers, even though Kieran wouldn't let me get a question inside with uh that's okay, as I'm used to it by now. It's only it's only 25 years. Is it 20? I don't know how many years it is now. So I'm I'm kind of used to it. But but so we should give you the opportunity. I don't know if you've got any questions for us. We've it's been like it felt uh just sitting here watching it, it's felt like a bit of an interrogation. So apologies for my uh my fellow uh pantomime horse interrogation. Um that's that's sweet of you.
SPEAKER_04Um I my question was are you the head or the back of the pantomime horse?
SPEAKER_03I think Neil is very firmly in the front of this pantomime horse. I'm very much uh at the back, but driving.
SPEAKER_00It feels like we take turns, is the honest truth.
SPEAKER_03That is true.
SPEAKER_04When you said uh speaking of ending it badly, this is time to end this, and I was like, oh my gosh, Neil now. But um I think I think that I'm really curious about what the state of AI is in the UK. Like in the US itself, I feel if you're on a different coast, like the West Coast, San Francisco, California, even every hoarding you see when you're driving from your airport to the hotel is going to be an AI ad. Whereas on the East Coast, it's not there yet. But we're a lot of good companies here. We have lovable, we have uh, oh my gosh, I'm blanking and they'll hate me for this. Sorry, Boston, but there's there's a really good AI push in the Boston area as well. But there is a difference. So, what is the state of AI in the UK and are people adopting it? Are they, you know, still on the fence?
SPEAKER_00I think it's a really interesting question, actually. We should give a little shout out to uh Nicole Alos, who uh actually put us together. So thank you, uh Nicole. It was very kind of you to uh to recommend to. Um and she asked a similar question, but she had a slightly different one, which was how do people in the UK um perceive the tech companies? And I and I don't know whether you know some people I think it's probably in all countries, you know, some for some people AI is bad and or tech companies are bad and blah blah blah. But I think I I think it depends on who you took. There's a lot obviously we we move in circles where there's lots of talk about AI because we're in that business. Uh and I think uh I I was reflecting on some of your earlier answers about adoption in in various uh sectors and various industries, and I think lots of people know about if you just take education. Um we all know kids are using AI to help do their homework, and and that, but actually, how are teachers helping them get better? And I think there's a lot more of that stuff than there probably was two or three years ago. There's a lot in the press about AI all the time. There was a uh a Claude Leak um uh just in in the press uh uh yesterday, um which was reported here. Um so there's lots of kind of data security concerns, there's lots of concerns about things like autonomous weapons. I don't know whether, you know, I don't know how much in the US you get to see things like uh Ukrainian war or or whatever, we get a lot of that stuff in in Europe, of course. Um, and then uh kind of automated systems, drones, etc. Um, I think in the UK we have a uh some big challenges. You like why don't the UK have our well there's no sovereign LLM? You know, there's talking there's talk about sovereign data centers in the UK. Um I think for one of the challenges for Kieran and I, and Donald, who's the who's the CTO for us, you know, we we get lots of people who are using the latest models like Claude or um GPT 5.6 and uh sorry called Ferble and ChatGPT 5.6. But actually what we produce is is uh confidential and private, and we actually build the solutions in um customers Azure tenants, but you can't get the latest models in the UK. Um we we are running months behind. Um so it's really interesting that we we don't get access to to those things within the the Microsoft Azure Foundry. So um uh so we we're trying to we're trying to keep up, I guess. Um uh there's lots of talk in the UK right now about things like you know, should you be using the Chinese models because they're open and fast and and free and lots of concern about data. I don't know whether you've got those same concerns in in the US.
SPEAKER_02Yeah.
SPEAKER_00Uh so I think it depends on on uh I think it depends on who you talk to, is I think the understanding you asked a re you asked a really small question which has got a very big answer.
SPEAKER_0376% of teachers use AI. That's based on a February 2026 survey of many thousands of teachers, so that's pretty good. I like that. And um I think 83% of businesses in the UK are using AI in some way. Most of that just means they've used copilot. So it's kind of and and in my mind that's kind of makes me cross because people I always say deploying copilot as your AI strategy is like giving everyone Excel as your data analytics strategy. It's like great, it can help enormously, but it's not going to be the answer on its own, that's for sure. So it's so yeah, so yeah, so so it's yeah, it is getting quite a lot of use. I get I feel over three years of doing this job, um, the argument is no longer about you need to adopt AI, it is how should we adopt AI? Not it used to be, I'm not so sure. Let's see. Yeah, that's that's gone. People are aware they've got to do it, it's just they're not sure now what to do.
SPEAKER_04Right, right. That makes a lot of sense. Yeah, adopting it responsibly has been the theme here as well. Um, I want to quickly touch upon two small things you mentioned, Neil. One was lack of access to the latest models. So you guys get the fable access, access to fable footage.
SPEAKER_03So this is inside Azure Foundry, what's as Neil's talking about? So Microsoft released to the different regions different models. Sweden get a lot of favor in the EU. They can put they can, or in Europe, they can um get they get all the latest models straight away. We get them a long time later. We just got access to 5.5, chat GPT 5.5 within Azure Foundry. So which is yeah, it's slightly annoying. We did we got fable until your uh Washington decided that no one was allowed it. Well, they they said non no non-US citizen, didn't they? And so it was pulled for a while. But we did so we do get those models, I think, pretty much as they're deployed on the public domains, but those aren't that useful in enterprise because they're secure.
SPEAKER_04Right, right.
SPEAKER_03For what we do, at least.
SPEAKER_04Um I I hope they push the other models out soon because I did check out both Kimi AI and Deep Seek. I think Kimi's K3 is doing pretty well.
SPEAKER_03Um you can run it locally and get fable style answers. I mean, that is very impressive.
SPEAKER_04Yeah, yeah. So thank you. Thank you for that answer. I and I think like it's you guys are if it's 83% and 70 something percent you said in schools, that's fascinating. That looks like we're all moving towards the same question, which is how do we adopt it responsibly and sustain it? And it's not it's not just a phase, it's not like the it's not hype, it's moved beyond that. So that's that's always a good thing.
SPEAKER_00Yeah, yeah. And it's been really good to talk to you about adoption on uh the other side of the pond. And uh to both compare and contrast, there's clearly a lot of similarities in what we're doing, and um, it's been a real pleasure to talk to you, Sue. So um uh thank you so much for taking the time. I hope you have a great day over there in Boston. Uh Kieran, thanks as ever, but a special thank you to you, Sue, and another special thank you to Nicole for putting us in touch. And um, I hope we can talk again soon.
SPEAKER_04Absolutely. Thank you for having me. I loved having this discussion and love to see you.
SPEAKER_03Keep in touch. I'd love to know, I'd love to hear how you how you uh keep getting on and keep learning. So uh we'd love to have you back at some point if you're willing.
SPEAKER_00That'd be great. I think we should do it at Christmas when it's pantomime season.
SPEAKER_03We should oh perfect. We can actually get a pantomime horse outfit and everything.
SPEAKER_04I would love that. I would love that. I will hold you to this promise.
SPEAKER_00We will speak to you in December. Thank you very much, Sue.
SPEAKER_03Joseph, thanks, Sue. Bye bye.