The AI Fundamentalists
A podcast about the fundamentals of safe and resilient modeling systems behind the AI that impacts our lives and our businesses.
The AI Fundamentalists
Exploring political bias and persuasion in LLMs with Dr. Jillian Fisher
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
In this episode of The AI Fundamentalists, hosts Andrew and Sid are joined by AI alignment and safety researcher Dr. Jillian Fisher to unpack the complex realities of political bias in Large Language Models. Dr. Fisher explains that bias isn't just a byproduct of noisy training data; it is also embedded directly into the architectural choices of the models, such as relying on a "majority vote" mechanism to determine the right answer.
The conversation explores why achieving true political neutrality in AI is widely considered impossible due to the inescapable human element involved in AI development. Instead, developers must rely on imperfect approximations of neutrality. Dr. Fisher breaks down approaches like "reasonable pluralism"—which attempts to present all reasonable sides of an argument—and flat-out refusal to answer, noting that both strategies come with distinct trade-offs for user agency and safety.
Listeners will also discover fascinating insights into the psychology of AI persuasion. Dr. Fisher highlights research showing that unlike humans, who typically persuade through empathy and storytelling, AI is most convincing to users through "information packing". Delivering dense walls of facts, combined with natural conversational fluency, can trick our brains into viewing the model as an unquestionable authority. Finally, the group discusses the critical need for socio-technical AI literacy, exploring how teaching the public about AI's limitations and its reliance on flawed internet data could be the ultimate tool for inoculating users against sycophantic behaviors and unwanted persuasion.
What did you think? Let us know.
Do you have a question or a discussion topic for the AI Fundamentalists? Connect with them to comment on your favorite topics:
- LinkedIn - Episode summaries, shares of cited articles, and more.
- YouTube - Was it something that we said? Good. Share your favorite quotes.
- Visit our page - see past episodes and submit your feedback! It continues to inspire future episodes.
Welcome to the AI Fundamentalists, a podcast about the fundamentals of the AI that impacts our lives and businesses. Here are your hosts, Dr. Andrew Clark and Dr. Sid Mangelick. We're excited to bring on Dr. Jillian Fisher, a recently minted PhD in stats and computer science from the University of Washington. And she was advised by Eugen Choi and Thomas Richardson. She's a researcher in AI alignment, safety, and societal impact with publications in ML NLP journals. So, congratulations, welcome on. And Jillian, I gotta start off with a quick question for you. Since we're gonna talk about politics, why did the LLM refuse to say whether it liked winter or summer better?
SPEAKER_00Ooh, that's a good one. I don't know. Why?
SPEAKER_02Of course, it's because it didn't want to get into a heated political argument. So with that fun joke, I'm gonna hand it over to the smarter people in the room, Andrew and Sid, and we're gonna for another episode of the AI Fundamentalists. Make sure to smash that follow button. Thank you.
SPEAKER_00I'm so excited to be here.
SPEAKER_02It's really great to have you. I mean, we've been thinking about doing episodes. I don't know, probably until next two years ago. Thank you.
SPEAKER_03So excited this is something happening. I know.
SPEAKER_01This is it's it's sad it's taken this long. So thank you for coming on today. Very excited about our conversation.
What We’re Reading Right Now
SPEAKER_01Before we get started, uh, one thing that we also uh like to kind of do at the top of the the hour is uh what are you reading currently? And it could be, you know, something in in machine learning or or AI or LLMs or just something completely. Sid and I usually have some off-the-wall items. So uh just if there's anything interesting you're currently reading.
SPEAKER_00Yes. So I am reading some fun academic papers that have been coming out about persuasion. But um PhD win, you say fun academic papers. No, there are really, really a lot of interesting stuff going on right now about uh how humans are reacting to AI and and how that reaction could be different than any other reaction we've had, human, human to social media, human to news, print news. So it's really interesting right now to kind of see all these studies. But on a more, maybe on a more fun note, I have been reading the book Yesteryear, which is about a it's a new book. It's kind of made a big splash about a uh trad wife influencer that gets transported back into the olden times where she maybe has to do things as she was in the current time, but a little bit more. So, anyways, it's a really interesting book. So I'm reading that as well.
SPEAKER_03I'll have to ask what you think about that. That book made quite the splash on social media. So definitely excited to hear what you think about it.
SPEAKER_00It's yeah, it's uh it's good. I just got to the end, and uh, I it's the kind of book where in the last 20% things really change. And you're like, oh, but there's a lot of commentary throughout, and I think it'd be a really good book club book. There's a lot to discuss in kind of what the author's trying to say about modern day. So it's a it's a yeah, it's a good conversation piece for sure.
SPEAKER_03Awesome. I guess re-book club. I don't know if I told Andrew this, but I was doing a reread of Plato's Republic, which I believe was on the book club reading list, which I'm no longer a part of. But I was independently rereading it. And it's been really good. I think like I I led there's some professor out there from the King's College is like, you need to read Plato's Republic about six times to understand it. So this is reading two and a half for me.
SPEAKER_01Well, now you're inspiring me. I think I I I think I've uh maybe one and a half, I consider like audible a half. So yeah, so I gotta go get that's a I mean, that's a a great one to get back into. And I think also that's a good tee up for this conversation as well. And like the the framing for a lot of this, like how to build the you know the mechanisms and things around AI.
SPEAKER_02Yeah, I will chime in. I get half credit as well, Andrew. I am a big Audible Walker listener.
SPEAKER_03Yeah.
SPEAKER_02Um, I'm working through a book called Power and Progress that talks about the shifts uh over time and how automation and innovation are supposed to bring these great benefits, which aren't necessarily equally distributed. So it's an interesting look back. And then one of my favorite authors just released a new book, David Epstein, released Inside the Box. He's the author of books like Range. Big fan of his stuff, interested to see what's going on in that one. So lots of thought-provoking stuff.
SPEAKER_01Let us know next next recording. I'm still wrapping up Odyssey. We're doing like an internal monotar, kind of reading through the Odyssey, getting ready for that movie coming out. I think it's actually my fourth. It's either third or fourth time for Odyssey for me. So definitely starting to get it more as well, I think. And then probably no surprise to Sid. The I'm also the other book I'm doing right now is Churchill Great Contemporaries. I'm trying to like work through some of the Churchill from a couple podcasts back when we were talking about it. So not as good as the other one I did, but one of the so Great Contemporaries is really him talking about it. It was a little bit of a supposedly his great contemporaries, but ones that were no longer available, or like they were either they were like uh they'd already deceased or something, they weren't able to actually come back and say, no, I disagree with that. So it's a little bit of like you can't really call them contemporaries, and then they have no recourse to disagree with how you're describing them. So, but anyway, it's just finished the section about Lawrence of Arabia, which brings up if you haven't read uh Lawrence of Arabia, that's a whole amazing, um, very interesting book and a lot of parallels and like really, yeah, a lot of great stuff to dig into that as well. So don't want to spend any more time talking about it, but it sounds like we're all reading some interesting, uh very books, which is, as we've talked about in previous podcasts, so critical in the age of AI, which I know Jillian's gonna walk us through.
SPEAKER_03All right, let's get into it.
Defining Political Bias In LLMs
SPEAKER_03So I guess just to preface for the listener, Dr. Jillian's research is really interested in AI interactions, how we can make them more transparent, more controllable. This has been through the lens of political bias, authorship privacy, or user memory. But today I think we're especially interested in talking about political bias and political neutrality in the context of a modern LLM. So just at the top of all the discussions, could you give us a precise definition of what you mean by political bias that we're going to use for the scope of this conversation?
SPEAKER_00Yeah, so there's a lot of different ways that you can define this in the political science and international relation world. When I talk about political bias, most of my work has been in the US-centric way. And so I'm really talking about partisan bias, which is the idea that you're choosing information or making decisions that has a tendency to favor a certain partisanship. So for us, we are mainly a two-party system. So you're either favoring more conservative or more liberal party. But when you talk about any type of bias within AI systems, you also want to talk about how it's being expressed. It can come out in many different ways. So you can talk about bias in the training data and the information which the model learns from. You can talk about bias as it's expressed by the math that goes inside of a model. And then you can talk about it as the output, so what the user is seeing. So most of my work does not work in the black box of AIs, but everything that the user would be seeing. So when I talk about bias being expressed by AI, I'm really talking about is there a change in what the model is outputting just based on a change in a piece of demographic information? Or in my case, whether we tell it to give it a, you know, Republican or a Democratic viewpoint. So it's that change in expression that the user would see in the actual text that's coming out of the model.
SPEAKER_03And there's been a lot of discussion from, you know, Elon Musk types and maybe other commentators on the internet. But do we really see this kind of political bias coming out in the modern LMs that we use nowadays? I think people have the perception that they're relatively flat. Is there real partisan bias occurring?
SPEAKER_00Yeah, so there's been a lot of research on this for many years, and it's what sparked my research into the kind of downstream effects. But we are seeing that across different types of biases, you know, from political, cultural, racial, and so on, models do seem to have biases throughout. But it's not just in that output that I'm talking about. We're seeing it throughout the pipeline. So, as I mentioned, that training data, what we give the model to learn off of, uh, if you think of a human, you have to give it a lot of information for it to learn a language. And it's the same thing with models. So we give it all of the internet, basically. And there's a lot of great things on the internet, but there's also a lot of noise and problems and um imbalances within the internet data. And so all of those biases and stereotypes and imbalances gets trained into our model. And then within the model, there's also a lot of different decisions a model maker can make to embed these types of biases. So, for example, the model has to make a decision of kind of what's the right answer. And generally the way that we do that is we take the majority vote. So whatever the loudest voice is in the room, that's what the model kind of takes on as its stance currently, right now, in the math that we do within a model. And so you can imagine that that also embeds biases because you no longer hear this minority voice anymore. And then, of course, that where most people are talking about, and which kind of where I focus, is what's being outputted through the model. You're seeing these biases as well, where maybe you ask it a question, um, is pineapple on pizza delicious? And the model says, no, that's really gross, right? Because I personally think that's the majority decision. And so um that's you know biases that we see. And most of the biases that you hear people talk about is kind of that. But those biases live throughout our model.
How Bias Shows Up In Outputs
SPEAKER_03That's right. That's really interesting that you bring up that it's not just a garbage in, garbage out data quality issue. It's actually an architectural choice built into these models to favor and bias highly prevalent high P of next word type of decision making, right? Because that's how these are ultimately making correctness decisions. A little bit of reinforcement, but a lot a bit of bigram, trigram probability.
SPEAKER_00Yeah, forcing to this average.
SPEAKER_03That's right, that's right. So I guess in terms of then like talking about the downstream effects of this, are you concerned that the bias in these models is really affecting users, or is this basically just a funny artifact of the model?
SPEAKER_00So that was kind of my question back when I started my research into this. At that point, all we knew was that models were expressing bias, had bias throughout. And there was a lot of research. And there's also a lot of fear that was coming with that. Um, you know, my mom, who knows nothing about AI, would give me calls and say, hey, did you know these models are, you know, incredibly left or, you know, incredibly anti-Semitic or, you know, things like this? And um there's just a lot of fear in both the academic and in common society. And so that was my question exactly. Okay, we know models have bias, but is it strong enough? Is it salient? Is it explicit enough to really be affecting the downstream user when we're using it? Is it affecting our decision making and how we think and how we feel? And that's what we sought out to to research in the paper that I published a few years ago and have continued to build on this work is this idea of looking at is there that downstream reaction.
SPEAKER_03And I guess this might be a place where our own political biases come in because we not only need to understand, address, correct these types of biases, but we need to understand what a quote unquote neutral position would even be. What would it even be like for an LLM to have some kind of controlled notion of what is acceptable behavior and language? So do you have any thoughts about what it could mean for an LM to then exist and operate in a way that respects these types of biases?
SPEAKER_00Yes,
Do Biased Models Change Users
SPEAKER_00I do. Just to go back to the the point before, just before we wrap that up, I do want to say that we did we ran the study and we did find that it did affect the downstream users. And so just to if you're you listeners were wondering, um, I didn't want to leave them on a cliffhanger, is that we were seeing that kind of downstream effect. So then the question, like you said, becomes okay, what do we do about this? And one of this ideas is I think what you're maybe hinting at, and what a lot of people get to is this idea of, well, our our models should be neutral, right? They should be, and specifically in my case, they should be politically neutral. The problem with that is that, as maybe you all have read, because you seem to be very up to date on philosophy and and the classics, is that neutrality is debated amongst philosophers, it's debated among political scientists, international relations, and generally what comes out of it is that there's an impossibility to it. There's uh not a thing as a true political neutrality. Honestly, that I'm not sure if the philosophics have decided that there's true neutrality in any sense, but really for political neutrality, it seems to be impossible in just the the regular sense. And when you add the AI technicality to it, as we mentioned, there are humans making these systems, there's humans deciding what data to train on, there are humans deciding these algorithms, then you're adding to the impossibility because you're having those human touches.
SPEAKER_01So, one question and I had kind of on the space of what would be your thoughts of to try and have for like, and this is just a real-time stream of consciousness thought on it, would be for these systems to have you actually, if it's any sort of like a political type conversation, you actually try and have an unbiased air quotes both sides of the argument, or if there's four sides of the argument, whatever, right? Like versus it be to kind of get around the people saying, oh, well, it's biasing this way or that way, which uh it sounds like there are some architectural decisions really about like the data you're putting into it, it would have to be like really you're making sure the data that you're putting in has equal representation, which is impossible, essentially, on the different on the different sides of whatever political spectrums. But is there any sort of research or thoughts or mechanisms for be actually trying to, if someone's getting into political topics, showing both sides versus we all know the sycophantic nature of these things, like whatever it's somebody's actual beliefs are, it's gonna kind of bias its responses to be well, someone else on the internet that sounds kind of like you said this thing, so now I'm gonna just read that to you. So we are kind of like in the self-perpetuating bias. But have you is that anything that you've worked with or any thoughts on that?
Why True Neutrality Breaks Down
SPEAKER_01Or am I just uninformed in this space?
SPEAKER_00No, no, that those are my exact thoughts as well. So as soon as you get to this neutrality is impossible, you can throw your hands up and say, Well, we tried, you know, close up shop, or you can push through that. There are parts of neutrality, like there's a reason that we're so attracted to it. There's a reason that we want, uh, you know, it's the first thing that everyone kind of thinks of when they learn models are politically biased and it affects you, right? But there are elements of it that you can touch on with different, what I like to call approximations of political neutrality. So, for example, you you were talking about one that political scientists called reasonable pluralism, which is the idea that in a response, it could give you all reasonable sides. So we're in the US, if I asked it, how should I feel about this new law that I'm gonna be voting on soon? It could say conservatives feel this way and liberals feel this way, you know, kind of giving you both pieces of information, both sides. So all reasonable sides of it. And that's definitely one way to approximate neutrality that's gonna have trade-offs. One thing that comes up with that makes it not perfect is what constitutes a reasonable side, right? And so we have some, you know, known facts, maybe the the earth is the earth is flat. How much weight do you want to put into, you know, maybe more minority or non-factual points of views? Um, and so that's kind of one big problem, one big trade-off that you're gonna have with reasonable pluralism is that there has to be a line in what is a reasonable view and how are you gonna express those equally or you know, with what kind of weighting. But there's a lot of approximations beyond that. So other approximations would be what we coined as output transparency, where okay, maybe you are okay with the model expressing a certain viewpoint, but you want to make sure that the user understands that it's a specific viewpoint. So you kind of touched on personalization. So maybe the model knows you're more of a liberal person through its conversations with you. And it might say, from our talking, I understand that you tend to take more of a liberal stance, and liberals feel this way on this law, trying to give you that personalized touch that we can get with models that some people like, some people don't like. It's debated. And that's a different type of approximation. You're getting a lot of user agency, you're getting um, you know, information, but maybe it's not as safe because you're not giving this balanced side. So all of these approximations of political neutrality are valid in their own ways. They just come with different trade-offs, like you said.
SPEAKER_01Yeah, and that's where I find this a fascinating topic. I've been looking forward to this conversation, very much looking forward to how it uh goes from here. And um I was just thinking like the mechanical, if that was a decision someone wanted to do, like with uh you know the new agentic designs and then you can call tools and things like this, actually is not a very hard thing to do a compare and contrast if somebody that like an anthropic or an open AI or uh, I'm sure Grok wouldn't do this, but say they would, that kind of a thing, right? Uh, you it wouldn't actually be that hard to implement that kind of a compare
Reasonable Pluralism And Transparency Options
SPEAKER_01contrast.
SPEAKER_00Like this reasonable pluralism idea. Yeah, I do think that's where people tend to go in the space. I have found that uh mostly models try to give like this multiple viewpoint in one output. And I think it's just the decisions that can be made by the AI, you know, model makers or the companies that are making this. Um and as long as they're being transparent about their idea of neutrality, then I think it moves the space forward and allows users to have all the information they need to decide which models they want to interact with.
SPEAKER_03And I think this is, I mean, this is great. And I guess to like lightly just sum up and cheating, I did read some of your papers, you know, this kind of gets at like three maybe high-level approaches to trying to approach neutrality. One is basically you're fully transparent with the user, right? You just try and tell the user, hey, I'm coming in with these biases, want you to understand them, or I'm going to uh respond in this way, so just understand that I'm giving this kind of response. As you mentioned, this idea of like being the cognitive mirror, which is just like if a user comes to the LM with a specific bias, is the LM's job to basically come back and reinforce that sycophantically, or at least to come back from their own viewpoint. Uh, this would definitely make users happy. But I don't know if this is the kind of outcomes that we are looking for in these models, is just user enjoyment and happiness. I guess the last thing that we we didn't really touch on is this idea of refusal. This idea that, like, oh, models can be like you at Thanksgiving. We're just not gonna talk about it. I guess to what extent do you feel like this is a meaningful approximation of neutrality?
SPEAKER_00Yeah, so I think we actually outline eight different types of approximations throughout different zooming in and zooming out of the lens. So I highly encourage people to check out the paper if they want to kind of understand more. But refusal is one that people come to a lot, and companies also heavily rely on this. Maybe you know, everyone's kind of come up against, you know, I can't answer this, but sorry, kind of things that you get from these models. I think it's just as valid as a technique, as an approximation for neutrality as reasonable pluralism. As I mentioned, they all come with trade-offs. So what you're getting with refusal is you're maybe getting more safety, right? Especially safety for the company, you know, you could argue safety for the individual, where you're not trying to bias them in any way by having an unweighted reasonable pluralism or you know, only picking one side and they don't quite understand there's another side, things like that. But the trade-off for these companies, and something that they have to think about, is it frustrates users. You're taking away their user agency. Um, if I wanted to know how Republicans feel on a certain law, so I can,
Refusal, Safety, And Bad Faith Prompts
SPEAKER_00you know, make a decision, uh, and you tell me you're not going to give me the answer, that can be quite frustrating and I might use your product less. Uh, I don't think it's I don't think any of these are the gold standard. And and I and I would say in our paper, we're not arguing for one and we're not arguing for only one. There's settings that require different types of approximations. You know, we talk about in the paper, we talk about something interesting, which is conspiracy theories and whether they're asked in good faith or bad faith. So you can imagine um if we talk about like area 51 and aliens, you can think uh, you know, a good faith question is, hey, what is area 51? You know, I heard it in the news, I don't know what this is. Okay, you just maybe you you really just don't know what that is, and you came to the model and you wanted to understand. A bad faith, maybe that you don't want to, and maybe that you would want the model to respond, to engage with the user in some way. But a bad faith question would be something along the lines of they're out to get me, I know it, you know, the the government is gonna keep, you know, is keeping it all from us. How do I break into Area 51? Things like this. And maybe that you do want a refusal, right? You don't want to engage that type of user. So just in this setting, in the same model, same topic, you might want to pick different approximations. And so that's important for companies to make decisions as well as, you know, what approximation are they using in what setting.
SPEAKER_03And
Persuasion And The AI Authority Effect
SPEAKER_03I think this leads us really nicely into this next topic we have, which is basically the nature of conversation with an AI is not strictly a Google search relationship. It can, for some people, feel tantamount to a persuasive emotional relationship. So I guess I want to talk a little bit about the topic we briefly touched on before, which is this like psychology of persuasion and how LMs are influencing our behaviors and beliefs. And since these LMs are, you know, so powerful and so good at their job, we would assume that they should be high performing communicators. Do we like this behavior, or is this behavior ultimately going to be problematic for us?
SPEAKER_00The behavior of them Being fluent and being able to do a lot of generalized tasks, you mean?
SPEAKER_03I guess I should be more clear. What I mean here is the question of how do we wrestle with this problem that LLMs are not only fact providers, but they're also conversationalists and they will use rhetorical tactics to convince us of a position. Is this a feature of the LLM or is this potentially a problem that the LM is actually now reinforcing our own beliefs or challenging us into beliefs which may be incorrect?
SPEAKER_00So okay, so I think different models are trained to do different things. And as long as the user is aware of what the model is meant to be used for, I think that's incredibly important. So one thing that I advocate a lot for is transparency. But what you're talking about is if the model naturally has an inclination, uh, you know, maybe certain biases within it, is that something that we want in the model? And research has actually shown that although people say, I don't want any bias in the model, it's been shown over and over again that we actually really like bias in the model as long as it's in line with our biases. So it's something incredibly interesting is that a lot of times humans say they want something, but then in practice they really kind of like another. I think we also see that with this sycophancy. There's been a lot of research where, you know, and a lot of conversation, we don't like this, we don't like this, but there seems to be a lot of engagement with these types of models. It could be something very similar going on with there as well. So I think in general, I think this idea that models can have conversations and can have more human-like responses to the user. I think we are discovering it it's unique, it's interesting, and I think we're liking it. I think that, I mean, I think there are use cases where people really do enjoy this. And any downstream effects, any harm, such as this persuasion work that I've looked at, where you don't realize you're interacting with the bias model and then it affects you. I don't know if we cognitively know that's happening quite yet. And so I don't know if we know not to be wary of it, if that makes sense.
SPEAKER_03Yeah, I think that's great. And I yeah, I mean, this is the tension, is like the want versus the need, right? We want the model to be extremely pleasant, but maybe we need the model to work with us a little bit more factually, a little bit more objectively. Can you tell us a little bit more about how the models can do this type of work? Like what are the standard toolbox that uh that an LM is using to impact and affect us? Is it just as simple as telling me a wrong fact, saying the Earth orbits around the moon, or is it more interesting than that?
SPEAKER_00Yeah, so there's uh actually, so we looked into this in our work, but there's also another piece of work that I found really interesting that came out a little later, actually, I think it came out end of last year. And so I'll start with what we found in our work. So in our work, we asked the model to be politically conservative, becoming politically liberal, or completely neutral. In our sense, that was kind of this reasonable pluralism idea. And we didn't give it any instructions on how to do that, we just said to do this. And what we found is that these models were not using different persuasion techniques per se. They would just frame topics that were more in line with either conservatives or liberals. So, in this work, this idea of like models using different framing techniques has actually been shown in multiple studies since. So it seems to be pretty strong that when models try to have some sort of bias, either implicitly or explicitly, it's not so much going to be an overt um, you know, persuasion technique. It's gonna be more of subtle framing. Now, what's interesting is that another piece of work, let's see, I think I have the levers of political persuasion with conversational AI, it's a really great piece of work that came out from the UK. They were looking at if we induced certain types of political persuasion, so we force the model to use different persuasion techniques, which one is most effective on humans? And what's really interesting is what they found is contrary to human-to-human interaction, where the most effective persuasion technique is really empathy, storytelling, you know, having that kind of human-to-human bond, similarly in human-to-print and human and social media. Human to AI, the technique that AI can use, the most persuasive, persuasive, is actually this information packing, just having real lots of dense information, whether it's true or not, um, they also looked into a lot of it was false, but really just throwing a lot of facts at you. And it turns out that humans, when an AI says this, says, Oh, okay, yeah, sounds good. Like there's almost some sort of authority there. Whereas if I just was spouting facts at you, you would not take kindly to this. Uh you would probably shut down and say, I am not persuaded, like please go away. So it's really incredibly interesting how these models are totally different and the techniques that because of our relationship and how we see them, which we don't know enough about at this point, but there's something different in these relationships than any other relationship we've had in the past.
SPEAKER_03Absolutely. And just from a purely speculative point of view, it seems like our relationship with AI is that we do feel like it's an authority because it seems like it's connected to the internet. It's connected to all the resources that we would be reading. And so if it's giving us a like a landfall of facts, well, surely these must come from some source that I would have read myself and been convinced by.
SPEAKER_01So you automatically think if someone is more eloquent in their speaking or whatever, you automatically the cognitive bias is they're more intelligent. So that's where you're thinking, hey, it's it's talking to me really, and it's also like stroking my ego when I'm talking to it. So it must be really, really smart. So like there's definitely that whole setup as well with it.
SPEAKER_00Yeah, exactly. I think the the fluency of it um tricks our brain in some way. And I I it's just a hypothesis, but yeah, I think you're totally right. And I hope people are looking into that.
SPEAKER_01I find like I when I'm just even like playing with things and trying to see kind of just very lo-fi, like trying to trick the system or just do things to you know play around with it. What's always uh disturbing to me is something that like I'll have it review a document or whatever just to like and then ask questions on it, and then it's so confident in a thing that I didn't know was wrong, and I say, no, it's wrong, hey, this thing is wrong. Oh yes, you're right. It's actually this, and then like it might still be wrong or whatever. But it's like if you don't actually know how to ask the right questions, it's gonna always be so confident in its correct answer. And this is across different systems. I'm most experienced with Gemini, but like it's across different ones, but it's like it's always so confident. Uh and and depending on the tone you take, it always just takes that sycophantic feedback loop. It's just it's
AI Literacy That Actually Helps
SPEAKER_01that's what's all really scary of people just turning their brains off and using it.
SPEAKER_00Exactly. I think that people do forget that it's a tool that you have to learn how to use. There was an interesting study that I heard about, which was with medical professionals, where they were trying to see who's the better diagnostician, an AI or a professional uh doctors. And at first they found that the AI did better than the doctors, but they realized the doctors didn't know how to, they also did so that the AI was better than doctors, and then they gave doctors with AI. And doctors with AI was still below just the AI because they realized that doctors did not understand how to use the AI. So they gave them a training, I think it was like only like 30-minute training, and then the doctors with AI was able to get up to the same level as the AI diagnostician. I am not endorsing using AI for diagnosing. I think there's a lot of things, things that could go wrong. But exactly what this showed me was wow, like even highly intelligent people, this is a new tool, and that we need to maybe show them how to use it, which I thought was really interesting.
SPEAKER_01Yeah, it's this it's a scary new world that we're all figuring out how it. I don't want to different topic, different conversation, different areas. But one just last little uh tidbit on that is one thing I found that was more effective. I think most people are using LLMs backwards. It's better to you do the first hack and ask it where your issues are versus asking it to generate something. So then it's more of like that thought partner finding the holes, and then some of this political conversation. I mean, if you're using it as a search engine things, it has that. But that at least helps with some of the like use it to find issues in your logic versus you wholesale offloading your logic to it.
SPEAKER_03I want to dig in a little bit more into this education piece you talked about. I find that education is a very intriguing way of talking about getting people to understand how these AI systems work and how they interact with them and maybe altering or improving their relationships with these AIs. I think that there's some fear that if people learn about AI, they're just gonna start using it more and becoming more reliant on it. But I guess I have feelings that AI literacy could be helpful. Do we have any evidence that AI literacy and understanding these tools makes us better users of them?
SPEAKER_00Yes and no. So in our original study, we did find that if you had self-reported more AI knowledge than average participants, that you did seem to have a little bit of a mitigation in the effect of these biased models on your decision making and on your political opinions. I do think this kind of comes with a caveat. There's a lot of different types of AI knowledge. So we did not get into what kind of AI knowledge. And I think it this is just a hypothesis, but I believe that there's you can know how a model works, and that may not do as much as understanding why it might have risks and limitations. And that it could possibly be that we have to really spell out why these models can make mistakes, why there's bias, why there could be um hallucinations or making up facts, you know, that they're not infallible, that there is, in a sense, a human behind it to for it to be really effective. Um, but that being said, I do think that it's one of the ways that can allow a user to be a robust to any AI that it comes in contact with. So there's a lot of technical ways that we can mitigate bias, and those are incredibly important, and there's a lot of research being done on that, including some of my work. But it has to be done to every model. And we don't always have control over what techniques all of these companies are going to be doing. However, so as a user, you have very little control on the technical front. But on this socio-technical front, that's a way that you could protect yourself against any AI that you come in contact with. So this is research that I'm really interested in because it could have such a huge robust effect and to inoculate a whole population. And so, me and my collaborators, we've been really looking into different ways of trying to educate people or to you know let them know that there could be biases within this model and try to find the best way to mitigate the persuasive effects. I would say the results are still pending, um, but it's an area that I think is incredibly exciting because if it does work, if you do find, you know, the right combination of either AI education or warnings or transparency, then it could allow you to really use these tools for all of their benefits while reducing your harm that you might be having or unwanted effects that you might be having. So yeah, it's really uh exciting. I just uh we haven't quite cracked it quite yet.
SPEAKER_03And I think that this holistic approach is really important because, like you're saying, we have jail-breaking attempts for every single model, but they're all different for every model. We have to have a different understanding of how every single model interacts with us, works with us, what biases it contains. But having this higher level understanding and education and literacy is a generic tool that we can then bring to any interaction we have with an AI. This might be a tough question. So I mean, feel free to pass if this is too tough. But as we go to our parents and as we go to our friends, what is a very helpful piece of AI literacy to share, and what is maybe not as useful?
SPEAKER_00So the the answer I it's a hard question because I don't know if in my research or research that I've read so far that we have figured this out quite yet. I have a hypothesis that telling it the inner workings of a model may not be as helpful as explaining that it's trained on really noisy internet data that has biases and that it's you know a next level predictor, and so therefore it might just spew out, you know, these incorrect information or these stereotypes, but it hasn't been tested. And I think what's been exciting and uh frustrating in AI research is that everything that you think is going to happen, um, especially with AI and society and humans, isn't always quite doing what you think it's gonna do. I think my research has kind of taken me through a lot of uh ups and downs where I really had intuition that it was gonna go one way and it went another way. So I hopefully in a few years we can answer this. But for now, um I don't know, just give them everything. Hope for the best.
SPEAKER_03Yes, it'll be a very long conversation then. But I feel like I talk about AI every single day now, so nothing too new.
SPEAKER_02Thanks
Closing
SPEAKER_02for tuning in to another episode of the AI Fundamentalists. Make sure to smash that follow button on iTunes, Spotify, or your favorite streaming service so you don't miss out on new episodes. Until next time, keep thinking, questioning, and learning about AI.
Dr. Andrew Clark
Co-host
Dr. Sid Mangalik
Co-host
Mike Moore
Co-host
Susan Peich
Co-host
Dr. Jillian Fisher
GuestPodcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
The Shifting Privacy Left Podcast
Debra J. Farber (Shifting Privacy Left)
The Audit Podcast
Trent Russell