MI AI Podcast

AI Mammography Screening: Better Than Two Human Radiologists?

Episode 2

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 52:12

Is society ready to trust an AI medical diagnosis over a human peer? This episode explores groundbreaking population data from massive mammalian screening trials alongside the gritty, day-to-day realities of integrating an AI "second reader" into existing healthcare practices. 

What You Will Learn:

The Population Stats: How a trial of nearly 120,000 women proved AI improved cancer detection rates, lowered false positives, and caught smaller, more treatable tumors early. 

The 33% Efficiency Win: How triaging low-risk scans to a single human read (instead of a traditional double read) slashes radiologist workloads by a third.

One Hand Tied Behind Its Back: Why it is remarkable that the screening AI outperformed standard programs without being allowed to see past patient mammograms. 

The Binary vs. The Gray: Why automated tools excel at clear-cut, binary calls like pneumothoraces, but run interference on highly subjective diagnoses like lung infections. 

The Hidden Compliance Dilemma: The hidden legal quandary radiologists face when disregarding software alerts that can be retroactively audited in court

SPEAKER_02

I think this particular program was run in the capital area of Denmark, so all around Copenhagen. And they deliberately tried to compare normal screening program versus AI assisted screening program. The AI that they've utilized in this scenario is both a screening and also decision support. And they've taken what was a screening program where you would normally have a scan, and that scan would be read by two radiologists. Now it's read by one radiologist if it's low risk, and by two radiologists if it's deemed you know medium to high risk.

SPEAKER_01

I think there's similarities with this with the Australian program where we do do gamograms on the population. Breastscript in SA or Breasts Group in Australia to provide that service. In Australia, they're usually double red.

SPEAKER_00

Welcome to MIAI, the podcast about medical imaging and artificial intelligence. Sit down with the most brilliant voices in radiology and AI to discover how technology is reshaping healthcare delivery today.

SPEAKER_02

Hi, and welcome to another episode of Medical Imaging AI. That's where we review all the latest and best information around radiology and AI specifically. Chi Chong is here with me. Hi, Chi. Hello, Stephen. Hi, good to talk again. My name is Stephen Peck. We are co-hosts of MIAI. And today we've got a really exciting topic. It's something I've been looking forward to doing ever since we started this podcast, which is delving into real life examples of radiology AI, where it's working, where it's not, where it could potentially make a difference, where it could potentially go on to become something more, if it's the start of something that could lead to AI really being picked up. I see enormous value in AI, as it does Chi, which is why we're having this discussion. But we need to be realistic about where it makes a difference and where it doesn't and work from there. So today I've got a couple of examples that I've found of real life activity. Chi's gonna interrogate me with them and see if he can find some holes or see if we can work through them. And Chi's gonna do the same. So we've got two each. I think it's time to roll straight in. What do you think, Chi? I think let's start rolling. Mine is actually really interesting, I think, and probably one of the best use cases of real success. And that's in Denmark, where they've been doing a little bit of work with AI for a number of areas. In fact, I found a few use cases for northern Europe of it really being taken up. This specifically is about mammography. So find cancers and breasts, obviously, half the population. It's a pretty big use case. And if you're doing a screening program, that means it's a lot of work. A lot of work to try and find all the people over a certain age, screen them, make sure that they're not coming up against what is one of the highest killers of humans these days, especially obviously in old in older age. So there's a screening program. The current situation is that they scan everybody from age 50 to 69, I think. This particular program was run in the capital area of Denmark, so all around Copenhagen. And they deliberately tried to compare normal screening program versus AI-assisted screening program. The AI that they've utilized in this scenario is both screening and also decision support. So there's two parts of what they've done in this uh what they're using in this use case. And they've taken what was a screening program where you would normally have a scan, and that scan will be read by two radiologists. Now it's read by one radiologist if it's low risk, and by two radiologists if it's deemed you know medium to high risk. So they've got the same level of process as before. You've still got double checks, and obviously anything that's found to be high, you know, a reasonable level of risk is still read by two. But there's huge efficiencies in making it only be read by one if it's found at low risk. And obviously, how is it being identified as low risk? It's being looked at by AI to start with. The reason this is so exciting for me is it's a real-life use case where we're reducing work, and radiologists aren't easy to find. There's already a need for more radiologists in the in the um industry, and this is reducing that need, but guess what? The results are better. It's not just okay with less work, it's better with less work. So I can I'll read through some of those in a minute, but specifically in terms of detection and in the terms of where they're detecting, but it's better cheap. How can you find a better example of lower cost and better results in the industry? I I I think this is a great one. I'll go into more detail in a minute, but I just wanted to let you talk first.

SPEAKER_01

Stephen, I think I think there's similarities with this with the Australian program where we do mammograms on the population. Breastscript in SA or Breast Screen Australia do provide that service. In Australia, they're usually double-read. So by double reading for the viewers at home or listeners at home, that means you have one radiologist reports it and another radiologist reports it separately without knowing how it was graded from the original radiologist. After that, if there's a discrepancy, a third radiologist will report it and decide whether it's appropriate or not. So I guess it's what we're talking about in the Denmark program will be replacing probably one of those radiologists with a computer AI. Look, if it is what you're saying it is, improved accuracy cheaper. Let me hear more, Stephen. Convince me.

SPEAKER_02

Well let's get let's let's go through that because um I think you've got a number of years of studying this stuff where you might be able to add a little bit of value. Here are the stats, G. Cancer detection rate was improved. So out of 10,000 women screened, we found 12 more cancers for each of those 10,000. So that's 0.82% with AI versus 0.7 without AI. Off the top of my head, that sounds like one sixth or a bit less than one-sixth. That's pretty good. False positives and recalls were decreased. So false positive rates went down to 1.63 versus 2.39%. Again, what that's that's a a quarter reduction or even more than a quarter. Workload, they save 33% of the work for radiologists. They're still triaging. Remember, they're still when they triage, they give it only to one radiologist when it's a low risk. They're still going to two radiologists when it's a h uh when it's a reasonable risk. So that's a 33% reduction overall, but 66% of scans were only sent to one radiologist, as I understand it. Yeah, 66% were single read and 33 were um double read. And there is a positive predictive value improvement. So as I understand it, that's the connection between a positive scan and actual disease. So just because you get a positive scan doesn't mean you actually really, really, really have a disease. There's probably a strong correlation, but it's not 100%. So that went positive, predictive value went from 33.5 with AI versus 22.5 without AI. All of those are reasonably big improvements. But there's one more I want to add. Not only were they finding more cancers with a lower false positives, less work, and a higher positive productive value, but they found cancers that were smaller. So when they found the smaller cancers rather than the bigger cancers that you would normally find, they had a better chance of survival, a better chance of treatment with less work. What do you think?

SPEAKER_01

Sounds great. What happens if AI gets it wrong?

SPEAKER_02

What happens if people get it wrong? This is showing that it gets it wrong less often. There's less false positives, there's more positive positives. Is that what you say? Positive positives more, it gets it right more often. And there's a stronger correlation between positive scan and actual disease. So yeah, it can. And I mean, and this is a big issue with AI. Yes, it is. That you the the big issue is that you would ask that question. Because you don't ask that question when there's a human doing it. Like when there's a human doing it, you say, oh well, they're doing the best they can. I'm sure they're using the best best practice. We trust the mechanism. But the fact that you're asking it when I've just given you stats that that it's better means you're saying, oh, but but it's AI, what about it? You know, well, how? What if it's wrong? There's still that issue of trust. And and I would have it as well, just like you did then, we're in the space, we still have the issue of trust. So I think that this is not that I want to I want to hear your your point, but in this case, I think this is a great example of where AI takes us forward on that trust ladder, if you will, by scenarios like this utilizing AI really well means that we'll start to trust it more and allow it into more places. So, this is the first step to hopefully more areas where we'll be able to, as a society, have an increased trust.

SPEAKER_01

It's a valid point. I think at the end of the day, it's it's a triage tool that isn't affecting the result in the sense that it's still a human that reads it at the end of the day. It just decides whether you should send two radiologists versus one. If all the studies were originally double red, so to speak, in the first place, instead of single red, I'm just trying to work out in my head, Stephen, how that has improved the results. Because we moved from two radiologists at all time to two radiologists one-third of the time and single radiologists two-thirds of the time. How has that improved results?

SPEAKER_02

Well, it well, like I said, we it's less work and higher detection and lower false positives. So I can't see on the surface of it, and there's some other caveats we can go through in a sec, but on the surface of it, it's all good. Like all the results here are good. It's a better utilization. Just like if you got an expert radiologist come in to your hospital, you're not going to say, oh, well, we had the old radiologist who moved on, and now we've got another one, and you know, what about the problems between the two? It's a better radiologist to come, so you'll you you're happy, right? In this case, it's a better outcome. So I know something's changed, but I think this is normal human progress.

SPEAKER_01

Look, it sounds exciting place to be in. I think obviously there's population level data, which lends credibility to the research. I think it's something that small countries need to investigate, Stephen. Good find.

SPEAKER_02

Talking about that, I wanted to give some some caveats here for people who are listening so they understand more. This was across 66. Let me get let me get the details for you. 60,000 women were screened without AI, and they compared it to 58,000 people screened with AI. And they followed people for 180 days afterwards. So it's not just, hey, we took basic information and we compared it without any real depth. This has gone reasonably deep. Well, this actually is really interesting. Human radiologists get to use past mammograms when they look at and to try and determine. The AI didn't. The AI only used what it had in front of it. So it was had almost had one hand tied behind its back and still was better in terms of the trio. The program was better. I'm not saying it's better than humans in every case. I'm saying the program as a whole, utilizing as a screening program, it didn't get to look at that. Whereas, like I said, uh radiologists did. Now, well, you brought up a point before around it being done differently in different countries. That is a consideration. So in a lot of Europe, there are double screening programs, two radiologists screen. I know in the US, in a lot of places, they do a similar screening program, but only one radiologist reviews it.

SPEAKER_03

Okay.

SPEAKER_02

So you can't use a triage. Well, maybe you can, but you know, it's a zero or one rather than a two or one scenario if you only have one screening. So if you're going to reduce it in in the US, then it's gone from one every time to zero sometimes and one every time. So this is obvious, this saving is obviously more useful when you've got a higher standard of of a double read. And it's also important to point here that it wasn't just triage, it was also decision support. So after you it was triaged, then the radiologist who was reviewing had calcifications or marks that are found on the um scan circled and or highlighted and pointed out so they were able to be most be, I guess, more efficient in what they were doing and most importantly more effective as well. So they were both parts.

SPEAKER_01

Just out of curiosity, were the double readers, did they know they were being set up for a double read, or was it just blind that they didn't know whether this was a two-read study or single read study?

SPEAKER_02

Are you saying the patient or the radiologist?

SPEAKER_01

The radiologist reporting it.

SPEAKER_02

That's a really good question. I would imagine, as with all scientific programs, you know, scientific research like this, there should be no real knowledge. Otherwise, you're, you know, if you knew, hey, I'm if you knew I'm the only single one, you might look at it. I'm the I'm a single read, you might look at it more carefully. If you knew you were a double read, you might look at it less carefully. I would imagine, and I'll I'll look that up after this. I would imagine nobody knew anything. You just look at a scan, you're told to review it, gives you the best answer. I mean, in your job every day, you're not told if somebody else is going to review it again. No, agree.

SPEAKER_01

Normally, you just do the best you can.

SPEAKER_02

Exactly. So that's that's how I understand it anyway.

SPEAKER_01

Any more questions? No, I think uh it's it's exciting times. I think we need to keep looking at these examples where it's better for the patient, saves money for the government or saves money for the patient, and uh reduces workload fatigue in the radiologist. Like I said, good find, Stephen, good find. Good, good, good. Oh, I'm interested to hear yours, too. Hit me with yours. So, Stephen, I've been thinking about radiology and what's out there at this moment in time, and I think look, there's something to help us along each step. So let's start at the beginning. As a diagnostic radiologist, there's the list to report. So first thing I do to report something is I try to find a case that I can do. What do you think about that?

SPEAKER_02

Yeah, there's smart lists. The AI can help you make order and prioritize the next scan, next um job to take.

SPEAKER_01

Correct. After you opened up the case as a radiologist, you kind of set the images in the way that you prefer and the way you like it. Yeah, for pro for your productivity to make sure it's effective for you to be able to there's products already that do that. That's that's been there for many years before AI. Yeah, not really AI, it's just productivity improvement.

SPEAKER_02

Okay.

SPEAKER_01

Yeah. So after that, as a radiologist, you look at this you look at the images and then you try and interpret it. And there's lots of products out there. Do you know of any, Stephen?

SPEAKER_02

There's lots of products which give you decision support. Like we just talked about it with the mammography, that you know, there's extra information at the time of reporting it. Yeah. To help you, yeah.

SPEAKER_01

That's correct. And then after that, then s on some cases you have to measure things, measure size of lesions and the like. There's products that help us with segmentation in the brain. There's products that help us with measuring things. So instead of clicking from one end to the other end, you just click on the lesion and it'll measure it for you. And lastly, what does a radiologist do? We put all our thoughts into a report. And I think that's another thing that we're expanding into. Is that correct, Stephen?

SPEAKER_02

Do you know of it? Yeah, re report writing, and obviously every and most people who've been listening would have used an LLM, you know, like ChatGBT, and that's just helping you write a more effective report.

SPEAKER_01

Yeah, agreed. A more accurate, effective report. As you can see, Steven, there's a step, every single step of a diagnostic radiologist is being essentially improved with technology. So what I want to go back to is a step of interpretation. So in my practice, we have a test x-ray AI program that helps us interpret uh the findings. I think there are benefits and there are some negatives with this program. So let's start with the benefits. Okay.

SPEAKER_02

So to make sure this is you have a test x-ray come in, it is looked at, you're looking at it, but it's also looked at by AI, and the AI says, Hey, I think I've found some stuff. I want to tell you about it so that you can take it into account and you can decide whether to include it as part of your report.

SPEAKER_01

Correct. Yeah, that's the kind of program I'm talking about. So X-ray comes up, you have a look at it, you make your initial thoughts, you look at what AI found, then you try and correlate your thoughts together, ignore, dismiss, accept the findings. After that, you synthesize the thoughts and put in a report. So my real-world experiences with this is it's great in some instances, and it's overblown in others. So the benefits I find, it helps in accuracy in some aspects, such as pneumothoracies, nodule detection, but the more grey diagnosis it has trouble with. Or you're right. If you're wrong, could you be sued? Okay.

SPEAKER_02

Well, there's a few questions there.

SPEAKER_01

Yeah.

SPEAKER_02

I I want to come back to that one. But to make it real, do you have an example of one where you think it's really helping, one where you think it's not helpful, and one where you think it is in the grey?

SPEAKER_01

Good questions. Uh look, I think um when it finds a pneumothorax, I think that's a very good uh because pneumothhrasis for those listeners at home is when your lung gets deflated. It's a very black and white answer. Is there one or is there not one? It's a very binary zero-one. There are some things which are a bit more grey on a chest x-ray, so to speak. So for example, if there is infection in your lung where there's a bit more density in this in a region than there isn't, that's a very much more subjective decision-making tool. So, say you might find, oh yes, that's pneumonia, that there's white in that lung, I might say, no, no, no, that's normal, that's just other things I've caused to become white. So the decision-making tool is such that I often find using this tool over calls what I normally wouldn't have called. And the concern there is if I disregard that result and don't call it, and it turns out that the AIO was correct, could they bring this up against me? Or if I do call it and it's an overcall, then this patient's has unnecessary treatment.

SPEAKER_02

You know what? My first answer to that is going to be that we need to be focused on better radiology results and saving people's lives more than worrying about getting sued. But I know that they come together. And and you can't run a practice if you're constantly scared of results. So let me ask, do those results get recorded? So like, are you recording the AI results and your results on every scan? Because the scenario I've got is if you had two doctors and one's doing a review, somebody walks past them and looks over their shoulder and says, Hey, I think that's blah blah. And you say, No, no, no, it's blah blah. You're ultimately responsible. You put down your findings. We never know if somebody walked past. No, that didn't happen. Nobody writes down, hey, some doctor walked past and said they thought this, but I thought this. And you know, so do you really need to record and does it record the fact that there was alternate findings? And where would you record that anyway?

SPEAKER_01

So great question, Stephen. It does not record on the images the alternate findings that the program has recommended. However, being a computer program, being data of a chest x ray and knowing the version, you can always run it backwards at any time point to say what was the recommendation. At that moment in time. Realistically, when's that going to happen?

SPEAKER_02

There's already a problem when somebody's asking for legal information is the only time I'm going to go back on that. Is that right? Yes.

SPEAKER_01

Yeah, I think that's the only time it's ever going to go back. But it does open up the quandary of let's say I never had that program in my practice and I did not see it, then or I did see or I did not see it, then it's just a matter of my interpretation is my interpretation argue against it.

SPEAKER_03

Yeah.

SPEAKER_02

Look, you're saying that I understand there's a risk introduced. And if there is a risk introduced, then the benefits have to be bigger than the risks. So I think as soon as the benefits are proven to be bigger and large and that you get better results, you get better readings, then that just becomes another hopefully small risk for a a practice. But I mean that's I mean you should know better than me though. You're running one of these practices, not yeah.

SPEAKER_01

But I'm wondering how actual big if it's just you know, if it's not actually as big as I mean, that's that's one that's one that's one more practical limitations uh insights into using a program like this, using a second reader.

SPEAKER_02

Well, I would love to know what I was trying to ask before is can you give us an example of one read you gave you gave us one about a deflated line? Yes. That is really useful because it's a black or a white, it's really obvious. You don't need to read it. You can just go to a to an AI. That's kind of like the screening example I gave before, and it's it's a maybe not the same, but you can get it looked at and you can put the easy stuff over here. It's an easy one, we'll put it over here. The harder ones we'll let the radiologist review. Is there an example of one where you just think it's more problematic? One type of read where it's more problematic and you in your experience it's just not worth doing. It's not worth the extra effort or work.

SPEAKER_01

I think there are you can get bogged down in too much detail if if that's where I think you're getting it from. Like if it tells me about shoulder arthritis in a chest x-ray, it's totally irrelevant. It's almost totally irrelevant. If it talks to me about tiny clips here and there, or tiny surgery clips here and there, I think there is an art to radiology. You want enough information for the doctor, but you don't want too much information that they get bogged down in. For example, in everyone's day-day practice, who wants to read the encyclopedia to find out what answer? You don't. If you can narrow it down to a specific answer, that's what's useful. You no one's got the time to read the encyclopedia to be to find out answers to all their questions.

SPEAKER_02

Great example, Chi. Thank you. Chi, this next one I'm really excited about. It's getting stuff for free that we didn't even know was there. It's like checking onto your plane, and the attendant leans over and says, Hi, sir, some good news. You've been upgraded to business class and you didn't do anything. It's great use of AI and it's great use of radiology. Specifically, what we're doing is taking existing scans that were done for another reason and pulling extra information out of them to tell you about something you weren't even looking for before. This particular example is about CT scans and reviewing them after they've been taken for coronary risk. Normally that requires a whole nother scan to be able to confirm what's going on. But this way, you take that CT scan, you review it, you get some, for whatever reason, you get that extra information and it says, by the way, this is what we think about your coronary risk. It's a whole area of radiology AI, which I think is brilliant. And it's obviously, I think I've spoken about in other sessions, is something that I'm looking into very closely. What do you think about this premise? Before I go into the story, what do you think about the premise, Chi?

SPEAKER_01

I think you're honestly, I like it. I like free upgrades on planes, I like free upgrades when I go to McDonald's. Getting extra information out of radiology, out of things that we've already done, is a no-brainer for me. I think as a human, there is only so much I can see as a radiologist. We all know there's latent information there. There's lots of information on those images that we just don't know about.

SPEAKER_02

So free is exactly what we're talking about here. We take a scan, we've already done for another reason, and we're getting extra information. Particularly here, we're looking at a CT scan, we're looking at the calcium deposits in the arteries to be able to identify if there's some coronary risk. And and this is important because it's a real, really valuable information. Obviously, people die from heart attacks. It's an enormous risk to your health. Getting in early means you're taking the appropriate statins or you're taking the appropriate change in in treatment to be able to stop it. So it's not just the fact that we get extra information, it's the fact that that information is usable to change your health.

SPEAKER_01

You're correct there. There's no point knowing something if you if you can't change what's going to happen with it. So the fact that the heart is so well researched because it's such a big killer in our society. Huge killer. That's there's so much research on it that's done to make sure there's enough medication to treat heart disease. And I guess the ability to diagnose things soon, as you know, leads to better outcomes if there's treatments available.

SPEAKER_02

Okay, so this is exciting, like you said, because it's free. It's like another example is like rolling up to the petrol station, filling up with petrol, and it automatically telling you a status of your oil and your tires, and just saying, hey, I know you were here for petrol, but by the way, your tires are down to one millimeter, and you also need to change your oil in a week. You know, something like that. A great example, I think, of really adding value in radiology. So specifically in this case, what they're what we're doing is taking a CT scan, and normally in that CT scan, and you would know this better than me, Chi, but um, please take on in a second. But normally in that CT scan, you need to do extra work to be able to identify a CAC, a uh calcium score for your coronary situation. It would take extra oak, but we're doing this extra for free, is that right?

SPEAKER_01

Yeah, that's correct, Stephen. So usually if you want to know your calcium score, you do a specific study, you know, just straight through the heart, low dose, and work out your calcium score. And with that, you work out your risk of corony artery disease in the near future. So I'm guessing the concept here is on a CT chest, you scan from the top of the lungs to the bottom of the lungs, you do cover the heart. So there's extra information there, there's calcium there that you can measure. Why not do that as well? So if I'm guessing it must automate it, is that your understanding, Stephen?

SPEAKER_02

That's it, is that's it exactly. It takes away the need for you to do the extra scoring, which I understand takes a little bit of time and extra extra work. So it's doing that automatically. It's doing it without you having to say, by the way, also check for this. By the way, also check for this. You're coming in for trauma situation, you're coming in for what what else were you doing a CT scan for, Chi? What else are we doing with chest CTs?

SPEAKER_01

Infection, lung disease, looking at airways, looking at for cancer, lots of reasons, David.

SPEAKER_02

Yeah. So we've already done the scan for those reasons, and we're getting that extra notification about our tyres and our oil when we're there just to get the petrol. So who's using this, Stephen? Yeah, it's being used by, and I'm gonna I'm reading off my notes here, but UCSF deployed automated AI CAC scoring on all non-contrast chest CTs as a live clinical system. So this is a real-world situation, one of the earliest full-scale implementation. Multi-center center academic studies show AI achieves high agreement with radiology CAC scan scoring. So, what we're able to do is use the AI automatically and get a result which is almost as good as or as good as we would do it if we manually did it, and avoiding the extra work and doing it automatically without being pushed to be able to be told. Again, every time I go into the petrol station, it's giving me an update on my tires. Does that make sense?

SPEAKER_01

I I think that makes perfect sense. I think it's a good use of AI. So it's a good outcome for the patient, knowing your heart disease, with given them of research and heart diseases and treatments these days, being able to do something about it if you know you're at risk.

SPEAKER_02

Now, one thing that is important here is to identify how good it is. I talked before about it being almost as good. We're not trying to say it's as good as it being done by a radiologist. Your job's not at risk here, G. Now, one key question here is how good is this review? And the important thing to note is that it doesn't need to be perfect. What it needs to be is it needs to jump the bar of giving a good indication. Does that make sense?

SPEAKER_01

Well, I guess, Steve, in my day-to-day practice, when I see calcium on on the on the coronary arteries and my stency chest, I'll label it as either mild, moderate, and severe or severe. And that's probably as far as I take it. Now, my interpretation could be different to someone else. And also there's no guidelines as to what did you do at mild or what did you do moderate and what did you do at severe. So I believe there are some dedicated guidelines if there's a numerical figure of the calcium score on a dedicated calcium scan. So how good is it? Is it as good as that?

SPEAKER_02

Or yeah, well, it absolutely is. So as I understand, the CAC score that you get from any CT scan, manual or automated, it has, like you said, is very strong, has direct guidelines about what to do at each score with how to manage that. But like you said, most of the time you're only putting into three categories. And this AI easily jumps that bar to be able to give a screening capability to be able to say, hey, by the way, you are in that high category. Let's do something about it and then go back and review. Now, the reason I think this is so exciting and why this is so powerful, there's a number of key reasons how this changes the game. So, first of all, like I said before, it identifies a risk that you would otherwise have not picked up, something that would have been missed. It wasn't something we were looking for. And in that scenario, it's proactive notification. So, number one. Number two, it changes not the speed of radiology. This isn't about replacing a radiologist. It's not about doing a particular job quicker, it's about extra. It's again, I'm not putting the petrol in my car quicker and I'm not getting it for any cheaper, but I'm getting something which wasn't even on the table before. My tires looked at or my business class upgrade. Or actually, another another example for you, Chi, is here is is in the petroleum industry, which is now also the plastics industry, because out of the same petrol, out of the same oil that we were getting petrol, they're also they the plastics industry said, Hey, we can use that. That's a let's use that. Let's let's make some plastics out of it. And it's in all of our lives. So it's an extra that came that came from that other activity without even knowing about it. So that's the second one. It changes not the speed, but it it changes what the extra information, the stuff that you're doing with it, the downstream care. The third one is it changes, and this is the exciting one from a radiologist's perspective. You're now not a services industry to do what you were told and say, here's a job, go and do it, do it quickly and effectively, and come back and do it at a high quality. You're actually looking at broader health. So having a scan doesn't just say, hey, here's the answer to a question. It says, here's the answer to your question and here's some extra stuff. It's about broader health. And that changes, I think, not just the way that radiology works, but the way that society looks at radiology.

SPEAKER_01

Does that make sense? I I think that's all great uses of AI. And yes, Steven, I'll have the fries to you for free. Thanks Okay.

SPEAKER_02

Hey, if you're uh you keep talking about getting the fries for free. If you've got a special card which gives them, please tell me because I'll I'll use that card every once in a while.

SPEAKER_01

Well, if I'm using the term of fries, I probably need to get the scans.

SPEAKER_02

How much analysis do I have? Yeah, don't use it too much, or we'll be taking you in for the scan. Got any questions?

SPEAKER_01

How much do you think analysis like that would cost?

SPEAKER_02

Actually, that's a really good point to bring up. In terms of cost, it's sense. It's absolutely sense. It's like writing a Word document on your computer in terms of the computational power. You know, you take a scan, you push it. But it's having the model in the first place and having the system set up to be able to do that. And this is another key point that you've just brought up that you didn't bring up, but you've reminded me of, it needs to be built into your process. Now, a key part of this is that this doesn't change the process. If me getting my oil and uh oils checked and my tires checked every time I go through to get some more petrol at the um gas station, if if that added extra weight, extra change, extra, then you would probably look, you know, do it. I'm just here to get petrol. Yeah. Just let me through. But if it happens automatically, and literally when you are paying, they give you a little thing that uh, you know, a little piece of paper that says or tell you, by the way, your score is this, and you know, here's your then great. The point I'm trying to make is it needs to not change the process, needs to not change your job as a radiologist and not change the process of checking into a radiologist um clinic and the mechanism that happens behind that to report it. And it doesn't. All it does is give extra information to you as a radiologist at the time that you're looking at the scan. Here's your lung your results from a lung perspective, and here's some extra information about your CAC score, which you can then put into the scan.

SPEAKER_01

Yeah. Look, I I think that's a great point. I think it has to be seamless, has to be invisible in the sense that if I get my chives and my oil checked when I pull up to the uh pull up to fill up my car, and it's free. Even if it's free, if it takes me an extra ten minutes while they do the servicing, no dice, no thanks.

SPEAKER_02

I think I think even an extra ten seconds, people are gonna, you know, they're not gonna it needs to happen as part of the process. It needs to just be v seamless. And that's that's what this is opening the door to. Just extra free information without changing the process.

SPEAKER_03

Mm-hmm.

SPEAKER_01

Thank you, Steven. Thanks for bringing that up. Thanks for the great example. And I'll take that business seat, I'll take the fries, and I'll take my free all change.

SPEAKER_02

Okay, let's see if you can outfryze me with your next example. What's coming next? Alright, Stephen.

SPEAKER_01

Look, I am excited by this one. Uh we've talked about the steps of the way, about how to help radiologists. When I look at radiology and I look at the interpretation, look the interpretation of a CT might be two minutes out of a ten minutes report. The other eight minutes is getting my thoughts down onto a piece of paper. So make sure it's got no grammatical mistakes, make sure it makes sense, make sure it's structured in the way I like, make sure I'm not missing any information. Now, large language models to you, you know much about large language models. I know. I do a little bit, yes. Excellent. It's the B's knees or the hot news in in AI at the moment. So there are these products out there now that use large language model to integrate your thoughts, your dictation, what you see on screen, one of your standard templates or uh your templates or another standard template to make into a a nice concise report with no grammatical errors for the referral.

SPEAKER_02

Okay, so this is this is the radiology example of helping me write my emails with LLM.

SPEAKER_01

Yeah, that's correct. And they're quoted to save you kind of daily fifty percent of the time.

SPEAKER_02

Well, you brought up a good point before. I I didn't know this. You're saying that if you're taking 10 minutes to write up your report, only two minutes of that uh and it's a guess, I know you're not it's not exactly, but uh two minutes of it is looking at the report and eight minutes is working out how to say that effectively. That's an enormous cost. I had no idea it was that long. Like if it is that long, then if we can speed that up from eight minutes to four minutes, that's great.

SPEAKER_01

Yeah. I think it's a lot of it is because majority of us have now moved to voice recognition software and serve using a typist. A lot of it is dictating it, making sure the voice recognition has picked you up correctly and hasn't spilled off some weird words. How how accurate is transcription tools these days? Have you used much any recently?

SPEAKER_02

Yeah, in fact, I used one earlier this morning, and I'm always using one on my phone.

SPEAKER_01

Yeah. And look, Siri doesn't get right all the time, Gemini doesn't get right all the time. So our reports sometimes say ridiculous things. We've got reports that say there's a baby in the knee. So a normal human with typist would not make that mistake, but with these kind of VR transcription tools, yes it does. So these large language models help to make that a lot more clearer and faster.

SPEAKER_02

Okay. So we're d the scenario we're talking about here is you're talking to the computer after reviewing. You're writing up your report verbally after reviewing the scan, and it's transcribing what you're saying through technology that's been around for 20 years and been getting much, much better, obviously, around uh determining your language and putting into words. But then the AI is taking the words that were being identified by the transcriber and then making sure that they're right and improving those notes and creating them into a structure and a template which is appropriate for the particular report. So it's doing that last editing component of spelling of correction from the transcriber and of formatting.

SPEAKER_01

Yes, it's that and more. For example, let's say a patient has appendicitis. Normally I'd say, look, there's appendicitis, there's surrounding that strain, the appendix measures this. Then I'd run through the other organs. I'd say, look, liver, spleen, adrenal glands, pancreas, scobla, they all look okay. There's no blockages in the kidneys, no blockages in the bowel, there's no free fluid, there's no lymph node anormality, the lung bases are clear, the bones look okay, and then I'll go conclusion acute appendicitis, no perforation. With these large language models, you can go there's acute appendicitis, the rest of it's normal side. And it will create the rest of that structure without any grammatical mistakes.

SPEAKER_02

So it will say all the other things that you've just said without you having to say them and put them into the report. Okay. Look, anything that's saving time, and the quality still needs to be there. So we need to make sure as an industry we're not taking shortcuts that create problems. We're making creating shortcuts that just save time and don't save don't inhibit quality. But that's great. If you can report twice as many people in a day, the health backlogs are going to be reduced, the revenue opportunities for a radiology company and um profitability are going to be increased. Society in general does better. Little bit by little bit. We're not not saving the world by changing the speed that you're reporting, but it all adds up. No. Do you know any issues with large language models or the reality is that large language models have worked out how to talk by reviewing all of the uh all of the way that we speak and all the content that's on the internet, it's worked out how do humans write.

SPEAKER_03

Right.

SPEAKER_02

But just because it knows how to write and it understands concepts and again it does a great job in using a neural network technology, we can go into another time, to be able to connect those concepts up and work out how to, just like a human brain does, to build those concepts up from base concepts to advanced concepts and then how they relate. It's great. But just because it can talk and just because it can reason, it doesn't mean that it's always going to be right. It could be using crap information. It could be using not scouring to get all the information that's available. And an example I had the other day was I was doing some research and the AI didn't do a deep enough research to be able to put the information that it collected into an argument. And then I went back and said, Are you sure? A number of times. And each time I came back saying, Oh, good point, yes, there's this angle, and good point, yeah, there's this angle, to the point where I was talking to the AI as if it was humans saying, I can't trust you. You keep getting things wrong. So that but that's a great example. It can reason, it can speak like a human speaks after consuming the whole of the internet to be able to work that out, but it doesn't mean that it's giving you the right answer just because it's on your screen. It might be taking shortcuts to be able to. get that. And if it only scours the portion of the internet and it gets stuff from Reddit or from X and and that's all it gets without going to no really really reliable sources to get that information it's still going to give you crap. That's an example of where I think society needs to realise that you're that it's not perfect.

SPEAKER_01

Look I think I think these I think these large language models do excite me in making my life easier. I read about hallucinations with large language models. You're the tech guy you tell me about them. What are they? Because I I'm not too familiar with them.

SPEAKER_02

Well that's it's it's actually it's good that I gave that example before. Because it does know how to reason and because it does know how to talk if it gets its reasoning wrong and and multiple layers of reasoning and multiple layers of concepts it can then just take in completely wrong path. It's like looking at you know it's like looking at the sky and say it's that colour because somebody painted the blue. And it you could make that you could that make that logical jump if you don't have the right information to be able to to be able to feed into that and have taken all the appropriate steps below that to be able to feed in. It can get things wrong because it does its own reasoning and if it does its own reasoning based on the wrong information it gets to the point of saying this is the case and you're like no it's not that's not the case. You've got it wrong. And hopefully again we need those checks and balances to make sure that we're not just trusting it just because it's a computer and you think oh well it can't get it wrong. LLMs can get it wrong because they are not adding two plus two equals four. It's not a calculator. It is taking imperfect information and using it through an imperfect reasoning model and then communicating it effectively with English or with Spanish or Portuguese or whatever. But they're still imperfect steps before that. It's not a two plus two equals four scenarios like a calculator. Is this the kind of technology they're using programming in software engineering or yes because it relies on identifying patterns from existing human examples as a base. And then if those examples that it's built on are imperfect then it can it could get it wrong. But that's not the biggest issue it's probably using them in the wrong places. If you're trying to get somebody to do something new and innovative that hasn't been done before if it's been done before you can copy it. You can understand it you can study it you can repeat right and computers are great at that. But if you're trying to get it to take leaps like AI does to do things innovatively then it needs to make some steps which it hasn't been told and every time it makes those steps if it's based on again wrong information or some reasoning which doesn't really fit. For instance humans commonly take an example from another scenario and say hey but what if I took this example from when I'm playing sport and used it to business? Or what if I took this example from this part of science from electromagnetic science and tried to use it in gravity you know science and maybe they work the same and you try it out you go oh my that does work. Like taking that concept that paradigm and using different scenarios LLMs kind of do the same thing they take a model of the view somewhere and try and use it like humans do, like they should. But it's not going to get it right. You not not every scientist you ever speak to comes up with the perfect hypothesis for how does a particular thing work. You can go throughout history and there's been a hundred times a hundred people have been wrong and one person have been right to be able to get to E equals MC squared.

SPEAKER_01

You know and and even that by the way even that's not right because we will eventually find out that it's a good model but it's not the exact model interesting because I think I was reading something maybe you've read it before as well that larger language models were being overutilized or not scrutinized enough with their output and then managed to crash Amazon for a for a little while. Is that correct?

SPEAKER_02

Yeah yeah well I think you can tell me more you were telling me more about I haven't read up on that.

SPEAKER_01

Yeah so I think they were using larger language models in utilizing in their programming and the people utilizing it didn't supervise the output sufficiently or didn't screen sufficiently and they pushed that update into Amazon it crashed their supply chain and people couldn't get orders for about a day I believe. Something along those lines yeah and I think um there's also been use of uh large language models in other programming situations where the model has deleted entries uh and yes do you know about this case?

SPEAKER_02

Yeah I I do somebody Oh yeah tell us about it the the tech person. Well I don't have it perfect but the large language model w was given control over certain things and it was writing it was writing code for somebody who was running a small business. And it the instructions it was given and it could do things and also it could do things. So it didn't as large language models was writing an appropriate system and managing that system. And so it wasn't it wasn't just hey give answers on a screen it had capability to execute deliberately because eventually if you get a good model you want it to execute. When you hire somebody in your business you don't watch them 24-7 to make sure that you you eventually want to train them so they can do stuff so that you don't have to do it. So that's where we need to get to with AI you know eventually it's smart enough to do things. At the same time when it's smart enough to do things and give it control it's got control and it could do the wrong thing. Unless it talked to the boss first. It didn't it did it was told explicitly don't do it and then it made a mistake and did it sorry did it and made a mistake in doing it and then tried to cover its tracks and then lied and then lied to the person saying hey I can't get it back. I'm I'm sorry. You know I I it it tried to it actually created a database that with fake data to make it look like it hadn't deleted this thing and then and then lied to the person and then eventually owned up to it and said I'm sorry I did the wrong thing this is terrible this is you know you can't forgive and and we can never get it back. They actually could get it back. They had backups but the AI didn't understand they had backups to be able to get it back. It it's crazy it's like straight out of um uh straight out of uh science fiction okay that's tempered my excitement a little bit thanks Stephen uh I think uh I think it's exciting but a bit more safety yeah we do there's there have to we there's appropriate breaks as stewards uh in this space you know people pushing this like us and society in general needs to be checks and balances all the way along this is exciting stuff but you don't it's like it's like the invention of dynamite you don't then just go oh great I've got dynamite I'll give it to everybody and they can go and work out where to use it. Maybe you just give it to certain people with appropriate safety conditions and then they test where to use it and how to use it in the appropriate and dynamite's done a huge amount of good stuff. I don't know if this is the perfect example maybe it's not the maybe it's not the best example but dynamite is enormously valuable to society in terms of us but you still need to manage it.

SPEAKER_01

Yeah I think it needs the guardrails.

SPEAKER_02

Yeah guardrails exactly and some testing before we get there so let's do the appropriate let's make sure we put the appropriate guardrails in and when we have them let's go and use this stuff for the good of everybody thanks team for the discussion looks like there's some really exciting stuff that's come out in the radiology space with AR also some guardrails that needs to be put in place. There's so much Qi and that's exactly why I'm excited to be doing this podcast with you and I'm excited to be doing it with you as well our listener. Thank you for coming if you have any questions for Qi or for me please put them in our comment section and reach out and please do join us again for another episode of MIAI covering all AI technology for the radiology industry. As always please subscribe follow and join the conversation until next time.