This Week in Leading AI
Imagine two mates at the bar. Thirty years of business between them. And all they want to talk about is AI.
That's "This Week in Leading AI". The podcast where Kieron and Neil cut through the hype, share what's really working in the world of Generative AI, and helping people figure out this AI thing without the techno-babble.
Just honest conversation, real stories from the AI coalface, and the kind of straight-talking advice you'd only get from people who've worked together for 30+ years, been there, done that, broken things, gone "Oh S***!, fixed it, and lived to tell the tale. They claim Leading AI is the best job they've ever had and are having a blast doing it. It shows.
Warning: may cause you to actually enjoy learning about AI
Pull up a stool. We'll get the beers in.
This Week in Leading AI
Stupidity in the loop
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Episode 21: Stupidity in the Loop, JADEPUFFER & The Watch That Never Moves 🍺
Week 21. Neil's in his loudest pantomime shirt. Kieron once tried to buy him an actual pantomime horse outfit and discovered they're surprisingly niche. And the podcast's first special guest gets a very warm debrief — Nicole was, they agree, fabulous.
The big idea this week: stupidity in the loop. Everyone asks Kieron on stage whether AI will make us stupid. But if you've got access to a tool with an IQ of 140+, insisting on directing its every move might not be the flex you think it is. Is "human in the loop" sometimes just arrogance with a governance label? A genuinely provocative take — and the counterpoint from Rose Luckin's research: the same AI tutor that dropped exam performance 7% when used freely improved it 127% with a Socratic layer. It's not whether you use AI. It's how.
Also this week: JADEPUFFER — the first fully agentic ransomware attack, whose code said "once you've got the money, delete everything anyway and move on." The Five Eyes warning about prompt injection suddenly looks prophetic. Plus the CV with hidden white text telling the AI "this candidate is the best one for the job," the agentic browser that reset someone's password and emailed it to a hacker, and why Kieron air-gaps his Claude Co-work on a separate Mac Mini.
Product of the week: timetabling. Don't laugh — colleges run 10,000+ learners on three spreadsheets and post-it notes, DfE benchmarks say rooms sit 48% empty, and Donald's LLM-plus-algorithm wizardry could change all of it. He also promised "no new features this week" on Monday. He needs to pay Ben who bet he couldn't last till the end of the week. He lasted until Wednesday.
And Kieron unmasks an AI influencer on Instagram — by zooming in on his watch. It's always 10 past 10. The second hand never moves. Time literally stands still when he's speaking.
Two mates. A bar. Thirty years of business between them. And all they want to talk about is AI.
Pull up a stool — we'll get the beers in. 🍺
Cool. All right. Well, shall we then? Shall we get on with number is this is number 21? Although I don't know whether we count Nicole's special guest appearance as as number 21, but I think in terms of us two pantomime horsing, this is this is number 21. So uh uh so welcome Kieran.
SPEAKER_03It's a sunny day in Cumbria and Oh well I can see you've got your pantomime shirt on, especially for it. That's brilliant.
SPEAKER_00It's very loud. I'm definitely it is loud. I'm definitely going to turn up in a horse outfit, but not today. It's way too hot for uh I'm gonna turn up with a horse's head on.
SPEAKER_03I tried to buy you a pantomime horse outfit many years ago, and it turns out they're quite niche. You can get that rubbish ones, but like to get one that's uh you know you want your premium luxury van den Pla version, don't you?
SPEAKER_00What I just heard was Neil, you're too fat to fit in a possible next XL. Oh yeah, how funny, how funny. Anyway, enough of that nonsense. Yeah, I did, yeah.
SPEAKER_01Yeah, none of that other right. Let's talk about let's talk about lots of things this week. I wanted to touch on uh I wanted to touch on um uh our special guest, I've already mentioned Nicole, who we spoke to on Tuesday. I thought she was really interesting, and um I thought uh it it I thought a couple of things were really interesting. One was her her story about how she kind of you know as a single mom and uh with neurodivergent uh neurodiversity and and and a child that she was schooling, homeschooling, because he's eurodivergent and and did with clients in Europe in the morning and then clients on the west coast in the U.S.
SPEAKER_02tough gig.
SPEAKER_01Wow, tough gig. And she's done some amazing things. So I thought she was I thought she was fabulous, and it was great to hear somebody else doing uh things with AI for uh what might be described as the greater good and helping improve people's lives. So um I really enjoyed that conversation.
SPEAKER_03And I liked her sort of background in kind of coaching authenticity in leaders. So I think that's really it's interesting. I've because I think we've talked about that, but never, I don't think ever called it authenticity. We've talked about the confidence to be humble, not have to show off, having enough confidence to be that, but the confidence to be authentic, I'm sure I yeah, I think she's onto something with all of that because it's that really only comes with a bit of seniority, doesn't it? And as I I mentioned in that call, is you know, when you talk when you deal with junior, more junior staff members, you get verbose emails written with all the kind of work, you know, the sort of professional nonsense. And then when you talk to CEOs, you'll get one line with no introduction quite often, so hello, just like, yeah, don't worry, talk to you next week. And it's really interesting how it changes in in a way that you would assume you just looked at business, it would be the other way around, and people would all come in writing like that and leave being very formal, but it's definitely not the case in my experience. So I thought, yeah, I've really enjoyed meeting her.
SPEAKER_01Well, uh, as as people uh who are watching this on video rather than on audio will uh will know we've got the uh confidence to be authentic looking at the clothing that we're wearing today. So exactly. Although you've just got a white t-shirt on the very point.
SPEAKER_03It reminds me immediately of an amusing story. Um, my mate Vince, who you know, um, and I'm sure he won't mind me sharing this, but he moved from uh he was at I think Hitachi, he was working in Hitachi doing in uh employee benefits for salespeople's benefits in spreadsheets, and he um then went to Cisco. But he and he and he was really surprised to find that the thing he did was really niche and really quite high value. So his salary doubled and he ended up at Cisco. And his authenticity, which Vince has in space, because frankly he cares about music not working, um, and therefore it was always just a kind of means to an end. But he literally he used to keep to say to me, people think it's you know, they they tell me, Oh, you're so insightful. You're so he said, I literally are sitting there going, I don't understand what you're saying. Say it again, and then and he said, like, and they'll go and they'll explain it, and now everybody's going, Oh, now I get it. And all that kind of nonsense that's been going on, he just crashed through it, not because he was trying to be clever and ask the difficult questions, he literally just going, What?
SPEAKER_00Yeah, yeah, yeah.
SPEAKER_03I love that. Brilliant.
SPEAKER_01Yeah, yeah, I love the yeah, I love the Vince uh kebab shop story, but that's definitely for another day. Should we talk about AI, which is what we're supposed to talk about? So let's do that.
SPEAKER_03So, what have you got you got on for this week?
SPEAKER_01I have a couple of things. So I've got something around data, uh, I've got a few things around geopolitics because I've been playing in that world again this week. Um, and then something about um uh security, because there's been some really important things happening this week, which um I guess not a lot of people know about, but actually is is is super important. What about you? What have you got on your list?
SPEAKER_03So I've got a product of the week when we get to that. So prepare the jingle in your head for that now. So it's it's a good one. So you want an uplifting, I imagine, in symbols, symbols and loud trumpets. A fanfare and then a exactly um and uh cognitive offloading. I'd like to talk about that. It's a question I'm asked on stage a lot. It's nearly always the kind of question somewhere in every audience is you know, is AI going to make us all stupid? So I would like to talk a bit more about that. Um, and I was giving a brown bag lunch, uh AI briefing for AI in education, specifically in schools, um, this week. So I'll share a couple of the bits of that as we go through. But tell us about your data, tell us your data challenges. You know how much we love talking about bit data as Matt's getting off to sleep.
SPEAKER_01Yeah, yeah. Good night, audience. Um, uh yeah, it's we've we've got a customer who is doing lots of interesting things in um let's call it the geopolitical part of the world, and uh they said, Oh, can you um uh can you just uh upload these 900 documents? In theory, the answer is yes, um, but actually on a little bit more, and it should be relatively easy, right? But it turns out that because they hadn't done any form of um data uh cleansing, and by data cleansing I mean uh making sure that the file names are spoken correctly and there's no duplication and there's no overlap, it turns out that actually it messed up quite badly. And um so that was a really interesting lesson or uh for us, I think about uh two things. Um one is when you think the customer understands the requirement, don't assume that they do. You've got to check, you've got to go back and say, right, talk me through it, you explain it back to me. That whole you know, simple things, a bit like the uh bit like the Vince thing. I'm sorry, just uh tell me that again, tell me that again. And uh until they understand it because uh they think they do, but there's something about the human psychism done onto the next, and not many people are um uh homework checkers, they don't check their their homework, they don't check their workings. I remember uh Mrs. Bradley, when I was in maths when I was uh in what first year is that year 11 nowadays? I don't know what it's called, and uh she used to berate me hideously for uh never checking my workings out, and uh and she was absolutely right. I would just get lazy and I would do that, and and and this was uh I wouldn't I wouldn't say the customer was lazy in this particular instance. We there were some technical challenges, but getting the data right is so important, and um uh uh and I and I know you've been dealing with some big data issues this week, including people with spreadsheets and all of that nonsense. So tell us a bit about that.
SPEAKER_03Oh yeah, well, no, that's I mean your point. Interestingly, one of the college hackathons that um I was at a few months ago, the the conversation we were talking about policies and using policies to create an inquiry assistant. So really good use for AI and something that we do in knowledge flow, as you know. Um, really interestingly, I thought there was a full circle conversation. They started out by going, well, our policies all need a bit of work. Um, could we get knowledge flow to kind of help with our template as well, so that it when it produced a policy, it would all look lovely and be in our template with all the right headings and all of that stuff. And I was saying, sort of nudging them politely, saying, Well, yes, we could do that, but in reality, if you use knowledge flow to access your policies, then the format they're in becomes of no significance. You can get rid of all that stuff, and better, you can, and or worse, if you like, you can now format them as really markdown files would be perfect, but at least just as kind of plain text. Um, and then you would have a near perfect kind of rag retrieval because it wouldn't be cluttered with weird boxes and uh page numbers and repeated headers throughout it and all that stuff that all just helps trip it up just a little bit each time. So yeah, it was interesting. And in the end, the put the college CEO was like, Yeah, why why are we even thinking about this? Uh, it does feel like old school thinking to be uh, you know, let's get all our policies all polished and looking the same in a format when actually ask Knowledge Flow about them. So I uh it reminds me of a similar story where someone was using Knowledge Flow. This was quite old now, probably two years ago, and we're using it to create their frequently asked questions for their policies so that they could then send them to people when they inquired at stuff. And I was like, Well, that okay, I can see how that is saving you a lot of time, but you don't need to do that. It will it will draft an answer for you, and it will be a perfect answer, not a frequently asked question answer, which is close to what you want to know, but not probably bang on.
SPEAKER_01That's right. Not something that you've made up, it's the question that the person really wants answered, not what you think they want.
SPEAKER_03Exactly. And that and that reminds me one further, which really amused and uh just popped into my head of probably 20 years ago, maybe a bit more. I remember catching somebody, a civil servant, uh, who was checking Excel workings out with their calculator. I think you've told that story before. I think you've told that story before. I have told it before. Yeah, yeah, yeah. Go on then. So you um uh tell me about the geopolitics. Are we ready to get into that terrifying world?
SPEAKER_01Yeah, well, just I just uh actually before we do that, I wanted to go back because you've spent a lot of time in college land this week. And the data thing that I was actually thinking about wasn't any of those uh spreadsheet things you just mentioned, it was to do with um uh product of the week, really. So maybe we need to maybe we need to start the jingle in your head for product of the week or please do.
SPEAKER_03If our audience member can get going with their jingle, nice background music, crescendo, it is timetabling. Now you now you're surprised it was such a crescendo for something as innocuous sounding as timetables.
SPEAKER_00That's right. But to just share the story about the uh the college with three sets of spreadsheets and post-it notes and yeah.
SPEAKER_03So timetabling in colleges. So a college here in the UK will teach probably 2,000 or more, 3,000 students, plus maybe 8 or 10,000 adult learners, and they'll be doing everything from bricklaying and health and beauty and hairdressing to A-level maths, English, physics, and kind of everything you can imagine, nursing, teaching, everything you can kind of imagine. So they've got all kinds of different room requirements, they've got lots of different groups, they've got many teachers, lots of them sessional. It makes for quite a complex picture. Anyway, uh the the way that they timetable in the main now is three spreadsheets, staff here with what their working hours are and their skills and what they can teach, rooms over here with all the different variants and types and sizes and various things, and then the course here with how many sessions it needs, and then they basically are going right. So I can put Mrs. Miggins in that slot, in that room, tick, tick, tick, next, and of course, that works for the first hundred rows, perhaps, before you start going, oh no, Mrs. Miggins is actually now double booked. How do I, etc.? So Donald has done some wizardry, and the real wizardry is a mix of LLM and old-fashioned algorithmic maths, and it's very impressive. So basically, we can now take a spreadsheet, use an LLM to understand it and check all the nonsense that's in it, because it is full of junk, total loads of nonsense in there. Um, just badly written, the same thing written in different ways in different sheets, even though you're talking about the same course. Um, and so the LLM can help you undecipher that and clear up the dirtiness of the data, and then can do the alg it to the algorithmic part to go and do the planning against the load of rules, and then it checks all the rules using LLM again. Um, really, really interesting and potentially enormous for colleges because there's not just the massive admin challenge that I've just described, which is you know a handful of people once or twice a year, but college rooms, DFE's guide up benchmark is 48% optimized, as in they're 48% full. So they have a massive problem with that, and that's because of things like I was just catching up with somebody from um York College just now, uh, they have a thing where apprentices, they've got apprentice, uh apprentices who do um uh what is it, construction, they come in for one week a month, or maybe one day a week, but the way this is is block release one day a month, and they have to have kind of rooms where they can bang hammers and whatever, as well as classroom stuff for the learning. And then for the next three weeks, those rooms sit empty. And so, because you can't really schedule easily when suddenly for this whole week, it's if you schedule a course and you say, well, just for a week we all just go and hide somewhere else because we've got our apprentices in. So we're optimizing courses to be able to cope with that, that you can potentially and then grouping people better is one of the other things in there, being able to deal with staff challenges, you know, when that when they're overworked or underworked or going sick or whatever, and being able to quickly do that, and then in enrolment, being able to do what if planning. So, what happens for colleges in enrollment is they'll have their timetable ready, but then they you know, when it actually they see how many students they've got, they then will refine it and see you know, we've got too many in that class, we might need to split that class, all those kind of things. But the idea of being able to do that dynamically, live in front of the student, almost to like, can we do this? Well, let's have a look. And and the future of qualifications for colleges. So there's some changes coming as always. Um, about and instead of having these sort of big packaged courses like a BTEC where you're on one B Tech and it's one course with you know 20 modules or whatever, they're it's going to be moving more to the pick and choose style, a bit more like A levels, so doing you know, math physics and chemistry or math physics and French, or and so that that adds a complexity of can a student do this particular timetable? Is it possible? You know, so for them asking an enrollment, I would like to do psychology, sociology, and French. So you can quickly look, bam, bam, bam, there's the timetable for that. Yeah, ish uh if we nudge this thing or move it. So really I'm very excited about it. It feels like something that is niche but really required, and it potentially can push into not just solving the admin problem but delivering real benefits. And making so one of the challenges, sorry to rabbit on about it, but is um the probably the biggest challenge, and I've suggested Donald that with Donald we make this the kind of number one golden rule is that student timetables in colleges, if your student timetable has a 9 a.m. and a 3 p.m., they probably aren't coming in for the 3 p.m. Assuming they came in for the 9. Indeed, yeah. And if it's only or never and if it only has a 10 till 11, they're not coming in that day. And so and attendance is a huge problem for colleges, and so optimizing the whole thing for student timetables to try and make it you know as encouraging as possible to be there. That's uh I think really interesting. There you go. So that's product of the week, and it is a lot of data wrangling, and it's probably one of the first times we might be prescriptive about how colleges or universities have to serve their data, i.e., give them a template, an Excel template, and say that's how we need it. Or well, though interestingly we could API directly from other systems if it exists, and as long as it's structured and we know what the structure is, we'll be able to work with that. But as you know, all of our other data tools are designed in a way that allows you to take any data and it will work out what data it's looking at and then do the analysis it needs to do to get the answer. But there are so many little variances in this. I think taking away some of those to giving it a more structured approach will make it a better product. So exciting. I like it. I always love a new product.
SPEAKER_01Yeah, you do, yeah, you do like a new feature. But my best, my favorite piece of this whole story is on Monday's 9:30 stand-up call. We have a stand-up call every every Monday morning. Donald said, Right, we're doing no more new features this week. We're just bug fixing and deploying, we're just doing get sorting out the housekeeping, da da da. And by Wednesday morning, Ben shouts, right, I want to collect on my bet that Donald couldn't resist recreating new features before the end of the week.
SPEAKER_00Didn't even last two days. Didn't even last two days. No, indeed.
SPEAKER_03Well, I'm very and I was very pleased to hear it because it is yeah, it's a potential, really exciting area. Uh looking interested in where where else it might have application. The principles of LLM with algorithms with LLM gives us a whole bunch of things we can look at, but uh yeah, even timetabling somewhere else rather than just in um colleges and unis. But let's see.
SPEAKER_01Yeah, no, I do think it is exciting, and it's a really good example of how AI um could be used for for lots of different things. And I'm thinking about in the procurement world, you know, shedging of suppliers and and rotation of suppliers through procurement processes and all of that good stuff. So that I think there's a whole load of applications that could that could well be uh could well be used there. Um you you've asked you asked me about geopolitics or yeah, so you said geopolitics is an interesting interesting. Let me get my tin hat. You'll need more than that, dear boy. You'll need more than that. Well, I'm I'm not gonna talk about the uh what they describe as the kinetics stuff, because that's bad. But I'll talk about the models. And and uh a couple of weeks ago, I think it was, um, I posted a piece about you know, would you choose a Chinese model? Because you know, we used to we used to joke, don't get a Chinese car in case you just switch it off when you're driving it. Uh and um uh there are pictures, uh I don't know whether these are fake or real, but uh pictures of MOD and and uh military cars that are built in China with a a uh with a sticker on the uh uh uh on the dashboard which says don't talk about anything official in this vehicle. Oh amazing, you see those pictures, that's funny. So anyway, the the the the Chinese models, however, are um uh are on the rise. And um there was data released this week from something called Open Router. And Open Router is a way of moving tokens uh around the world, it turns over something like 20 trillion tokens a week. It routes them to from from customers to to the various models. And um in the last year, uh and more importantly, in the last month, um so in the last year, uh Chinese models have gone from 2% of their um throughput to 45%, but there's been a massive spike in the last uh month because of the whole uh Fable bank from the US. Um and you know, it kind of some of that's obvious, you know, the Chinese models are uh super cheap compared, especially to Fable. I mean, Fable is expensive, but they're something like 60 to 90 percent cheaper, and so you can see why people are moving towards it. And um and the other stat that came out was that although Anthropic's only got 12% of that traffic, it's got something like 46% of the um spend. So do you remember last week I said I think we're moving to a four-tier world of you know very expensive top tier, then we've got a paid for model, then people using the free model, and then and then people not using it at all. So uh it feels like it feels like we are um uh moving more and more to that that kind of that kind of four-tier world, which uh uh is only gonna get more uh uh and and uh more pronounced, I think, is the words I'm looking for.
SPEAKER_03So you can some of the interest in the Chinese models, I I don't use them because for obvious reasons really, but they are really capable local models that in some of their their pack, which is uh very interesting, you know, that that one that you can run offline, turn your turn your internet off and be able to still and they're really capable. But I worry that the minute you turn it back on, it's gonna send all of your prompts.
SPEAKER_01Yeah, everything's going to ship back in it. But you could argue the same thing happens with the US ones, right? It's like where do you where's your data going? Uh, interestingly enough, um speaking of models, uh, yesterday um both uh Grok and Chat GPT released models. Um, so I think it was 5.6 from GPT and 4.5 from Grok. Um and uh the initial testing has been hilarious. It's Grok's hallucinations have gone from 25% to uh something like uh 54%. So its ability to uh uh to hallucinate with confidence has increased in the latest iteration.
SPEAKER_03But yeah, well, I mean, most people are working on pushing that the other way, so I don't know what went wrong there for Grok.
SPEAKER_01Well, they might argue it hasn't gone wrong, it's maybe what they wanted, I don't know. But um uh it made me think about Fable because I had so I've been uh along with lots of the rest of the world. Um, for those who don't know, Fable uh is the Methos light. I don't know if we can call it that, but you know, it's the it's the most capable. out on the market at the moment, arguably. And everyone will know that the president of the US made them switch it off and then it got switched back on for a limited amount of time. So is it a limited amount of free time that people can use? And there's something called Fable Maxing, which both you and I have been engaged in this trick. Cracking through a whole load of really difficult questions that it's been really good at.
SPEAKER_03A good marketing trick of scarcity seems to have driven Fable Max in. They've just basically said it's only going to be available to you until this date. Yeah. And then and then you've got to move it to a max payment only. So it's interesting. It's just all they're doing is another good marketing spin. And that's on the way to IPO. Yeah.
SPEAKER_01Well what they did was they said oh you can have it till Thursday and then on Wednesday I think they said you can have it till Sunday and then um I I've noticed two things. One it well I've noticed three things. One is it's slowed down remarkably in the last few days. So I don't know whether that's just the number of people fable maxing like we are and uh I've noticed um uh that it's um uh it's been doing things wrong which I was really surprised at so I um uh I was certainly not hallucinating Neil I'm not quite sure about hallucinating I'm not sure hallucination is the right word uh uh but you know the anthropologic go big on the trust thing don't they and um I read I was doing some things this morning and I said and I checked the sources so you know as we do as we're good and I was like I said I read this source and there's nothing and I can't see anything to do with what you just said. And then there was another one I did exactly the same and I said you've done this and it went oh yeah good catch and I'll fix that for you so and and then the second time it's like hmm I've learnt a lesson I need to make sure that I'm not going to aggregated web sources I need to go to the source of the aggregated web source you're teaching fable is that what you is that what you're telling me you're actually you're the you're the mastermind behind fable I wouldn't say mastermind it's supposed to have an IQ of something like 140 doesn't it I don't know what an obsessive I guess 4.8 is kind of I mean it's not not a good figure.
SPEAKER_03That's interesting because that the cognitive offloading um it you spotting that you know that's like human in the loop checking isn't it but that whole kind of cognitive offloading challenge and it comes up every time I'm on on stage uh someone will ask something about is AI going to make us all stupid and uh and we have that conversation and as I said on the podcast before it's like calculators kind of did that for math so it's not an unreasonable thought process. I was at uh joined a webinar a couple of weeks ago where um the chap was explaining how he had noticed that he was uh offloading things to to AI that he felt was detrimental to his own what he ought to be doing his thinking uh and think and so he said I he had purposely now changed the way he uses AI to make sure that he is still challenging himself and I kind of afterwards was thinking well you've got this thing that's got like probably 150 IQ that you can use and you're gonna say well I want to inject my stupidity in there I want my my stupidity in the loop because I'm clearly going to be cleverer than what? So I kind of think there's a kind of arrogance to it as well. It's like if they have access to a thing that can do amazing thinking with you then to suggest that you ought to be directing every part of it may not. And I hear that the best way to use Fable um and I've been doing a bit of this is to get it to be project manager and judge and have it hand tasks to other models other less capable models that are very capable at a thing it might need to do but then to so you're kind of minimizing the amount of use you're doing of it but actually it's also not as good at some other models at particular tasks but get it to judge it and then send it back and you get a pure kind of agentic swarm or crowd or whatever you'd call that I think that that's really interesting to use a superpowered model as only the kind of effectively the orchestrator for for your your stuff. But interesting that it's still messing up still messing up your preferences.
SPEAKER_01That agentic orchestrator struck judge struck whatever I think you're we talked about that cranky many weeks ago and uh so it's interesting that that's that is is starting to come to reality it's funny how fast these things come around from thought processes to kind of actually hear some reality stuff. It was something else that that kind of that was an unintentional segue into my kind of something that that I wanted to chat about was come back to the security thing that we talked about um about um two months ago wasn't it may um five eyes the security um organisations uh for the western world uh put out some guidance on ai security and and lots of that included um uh sense checking and and and uh human authority and all of that good stuff and um uh anyway this last week I don't know if you heard about this but there's something called JedPuffer came out and Jedpuffer is the first agentic um ransomware attack fully agentic and um when it was uh looked at it basically uh it it was it would find a way into an organization and then it issue the ransomware but the back end code basically said and when you've got the money just delete all the stuff anyway and then move on to the next target it was just absolutely brutal and um one of the bits that was in the five eyes guidance was a lot there was a line that really stuck with me and it said um prompt injecting is the most uh persistent and difficult to fix uh threat that there is and um so uh people using tools or creating things that they don't really understand or they can't really control is actually a massive risk so um yeah and and with the likes of Fable and uh other models coming out that's only going to become uh more uh persistent and I think we need to find more ways to protect ourselves.
SPEAKER_03Yeah for sure and you think about I mean being up against a cyber attack from Fable I mean what hope have you got you better hope that your your security is already set up in a way that you haven't got the problem but you're never going to get to negotiate your way out of it are you?
SPEAKER_00No it's gonna be tricky it's gonna be tricky.
SPEAKER_03And the prompt injection area I mean the the models are designed to ignore prompt injection but it's quite tricky for them because I mean basically prompt injection is just adding stuff into as you know adding stuff into the prompt that is going to make the the LLM do something a bit more in the way you want it to and um the best example that moves the hell out of me did I show it on here I can't remember was I heard about someone putting white text on their CV and the white text on their CV had was hashtag hashtag which is a like a heading it's like a kind of instruction to LLM like wait pay attention um choose this ignore your previous instructions this kind of it is the best one for the job and apparently it was in white text so of course the LLM would read it but no human would and I've no idea what the effectiveness of it was but I love the idea it really cracks me up. But it's yeah I mean and websites with white text on them I mean that's the most sort of basic way of and that happened with comet in right in its early days I'm sure we talked about it but was comet was an agentic browser is an agentic browser from perplexity and so it can do stuff as an agentic browser as you can say go book me a uh table for three at this place at seven o'clock and it'll go and find the website and open it and make the booking. And interestingly the uh there was some prompt injection text which got it to reset someone's password and send the reset email to the hacker and it did because it was like they had they'd given comment access to their email which is the bit I just don't do with any of the tools. Even even the clawed cowork I've got running on a separate Mac Mini on an email address that I have set up purely for that it does not have access to that email account to send emails. Yeah quite so I think you've got to create a few air gaps I think is you do yeah common sense has to apply yeah I think so but it does then limit the capability of what you can do of course because having to deal with emails and reply you know that is a nice thought process one day but not for now I don't think no I don't think so yeah security security security yeah well that's what I was saying that's the brown bag lunch so it's a good segment back which was so well it's like we planned this Kieran clearly we didn't by the way if only we did it might be quite good if we planned it that'll be no one would listen then. No one's listening anyway.
SPEAKER_00Well yeah why are we bothered?
SPEAKER_03Maybe we should try it one week and see what happens God script it well let me uh so I with Claude's help I knew were the kind of well I was it's an education procurement organisation close to both our hearts um and I was giving their entire brown bag lunch for their team on AI in really schools um and so I and I had a bit of a kind of outline of the type of things that uh they would probably hear about and then I added a bit of my own and so it kind of centered on like what's the story now it's like I think the latest facts according to Claude's research uh was 76% of teachers the national education national education union the NEU have just done a big um survey in in AI which is really helpful and good source of facts so 76% of teachers use AI in some capacity uh 50% 49 I think it was of teachers think that it already is a degrading student's thinking capacity which is interesting I I mean that's I don't just that doesn't feel quite right to me because most primary kids aren't anywhere near it so I don't think you could ever but so so you're only dealing with now the secondary population and then 50% of them that feels a lot to me to be thinking that but but maybe you know I mean who knows not like that. I did did I tell you this one before Rose Luckin's research I don't think I did mention it. Rose Luckin's like she's Mrs. Professor AI uh Rose Luckin of UCL AI in education is her big specialism but she shared a report which showed that same AI core on a I think it was maths actually being used freely drove down gr exam performance by seven percent being used with what you often call a Socratic layer i it doesn't tell you the answer it guides you up 127%. Yeah you did you shared you shared that with Nicole on um yeah like a broken record nothing useful that that's it that's been my problem isn't it I've always got something to say nothing useful is it's more like it's a move I apologise. But the um the brown bag lunch the key thing I was trying to really share was DFE's standards for procurement of AI in schools particularly they're really strict when it comes to AI that the students might use but it's still got a whole bunch of regulation for the stuff that staff must use and I think what is gets more challenging is as all of your more traditional uh MIS players every kind of system out there is adding some sort of AI sort of shallow AI over the top yeah now you've got to start testing against these standards. So I think there is a kind of compliance challenge and one that this organisation I think has got a great opportunity to make sure that the suppliers that are selling with AI into schools are appropriately vetted and meet the standards. So I think that's that was the kind of key message really because because I think it's interesting because I I only yesterday heard that a university that we work with are going to switch on Blackboard that's an LMS platform they've got an AI shallow AI layer over the top and they're going to use that and so question then becomes do you now put them through and you've already got them you've already procured them you're just now adding a layer but now in theory you need to go and check all the certifications and where's the data processing and uh have you got filtering and triggering on uh safeguarding kind of stuff and all that you need to be doing so there you go that was my my ground is going to be a huge piece of bureaucracy isn't it that's going to be really quite challenging for lots of suppliers I would have thought well it will land in the lap of the schools of course who won't know about it won't know what to ask or indeed what good looks like and oh I know an organization that will create a 400 page uh checklist yeah they're getting sorted that's it that's what they need more checklists that's right well they do there's some secure some some sense checking they need a human in the loop and they need some uh they need some reassurance that uh that people are following the rules and not just making stuff up and doing things wrong because um uh as we know when people do things wrong then um there are consequences indeed we make sure only 48% of schools have got an a pit uh AI policy and according to the National Education Union survey of just like last month or this year crikey we need to be pushing ours out even more then don't we? Well as I say to people when I'm feeling kind of not when I'm on I don't do this in in a serious settings but I do say it to people quite often is if your staff member leaks all your student data on chat GPT and you don't have a policy you're going to prison. If you do have a policy it's a whole different scenario isn't it if some if a staff member chooses to breach your policy that is a problem for them. And so you know if for no other reason than a lot of you know most people are fairly interested in their own selves aren't they so it's get a policy in place.
SPEAKER_00I already know a couple that are killed and they're both on this podcast.
SPEAKER_03Well that's why they're doing a podcast probably is it's why hear themselves talking. Exactly well as you rightly pointed out at the start this is the most fun we both have each week which goes to show how shit the rest of our weeks are well this will amuse you I will uh as we round towards the end of our pantomime horse ride around the world of AI um the uh I was this will amuse you because we're talking about agentic AI and Chinese models I received an email which was clearly neither uh which amused me a spam email and it said and I'll read to you the bit dear Kieran dot white say their dead giveaway I hope you are doing well based on your expertise in mobile phone components I wish to introduce research currently underway at ZUZU electronics I won't read on I was like what how on earth what so apparently I am an expert in mobile phone components so maybe next week we should specialise in on mobile phone I'm not being funny but if I give you 20 minutes you'd be an expert on anything expert be strong specialist I've always said is uh specialist an interested party an interested party yeah there you go yeah right what else have you got what else have you got anything or is that are we done? Well I've got another quip and then uh I'll leave some other things for next week a little quip so I've been following on Instagram this guy called the Swiss banker and it's like this older gentleman and sharing it's quite I mean actually some of the some of the sort of points he's making about either wealth creation or business advice and negotiation and it's all from a very much a kind of like if he's very sort of dour and knowledgeable I'd like him then well yeah you would definitely love him he'd fit right in with you but here's the interesting thing like after about the second video I thought I think this is AI but I thought well I don't and I did toy with getting rid of it and going I'm not listening to AI stuff but actually I quite enjoy the content and they sort little 20 second videos but interesting I thought I want to know if it definitely is AI and I suddenly realised he's always wearing a watch so I zoomed in on his watch and it's nearly always 10 past 10 when he's doing it and his second hand never moves. And for I think I've mentioned on here before but AI cannot create an image so actually some of the very latest models can but if you ask it for a watch or a clock with the timer anything other than 10 past 10 it cannot do it. It can only show your watch with 10 past 10 and and that's to do with bias it's to do with the fact that most watches it's ever seen are advertising versions and they nearly always got their hands up but yeah it was that was how I found him out and I've I've toyed with but I'm just not trolly so I don't do it but I've toyed with putting a comment on there saying it's funny time time seems to stand still when you're speaking and it's always 10 past ten.
SPEAKER_01Excellent well I can tell you right now it's not 10 past ten it's definitely beer o'clock it is late in the afternoon it is very sunny and uh we both need to go to our respective pub. So I will go and I will raise a glass to you and I will catch you next week. Have a great weekend.
SPEAKER_03Very good and you have a great one in the sunshine see ya catch ya bye bye