Membership Events

The Score: How To Stop Playing Someone Else’s Game, Professor C. Thi Nguyen

TRIP

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 1:22:55

On Tuesday, 20 January, Professor C. Thi Nguyen gave an online talk to TRIP members, offering a timely guide to the joy of games and the danger of metrics, and how to rewrite the rules in order to build more creative and joyful lives.

In his new book The Score: How to Stop Playing Someone Else’s Game (Allen Lane, 13 January 2026), Nguyen shows us how our newly gamified world has fundamentally captured our value systems, forcing us to prioritise what can be measured and monetized over what is truly meaningful.

About the Speaker:
C. Thi Nguyen is Associate Professor of Philosophy at the University of Utah, and a specialist in the philosophy of games, the philosophy of technology, and the theory of value. A former food writer for the Los Angeles Times, Nguyen is active in public philosophy, writing for the New York Times, Washington Post, New Statesman, and elsewhere.

SPEAKER_05

Hello, welcome. Thank you so much, everyone, for joining us. And I'd like to extend a very warm welcome to Professor C. T. Nguyen. He is an associate professor of philosophy at the University of Utah and a specialist in the philosophy of games, the philosophy of technology, and the theory of value. As a former food writer for the Los Angeles Times, Professor Nguyen is an active is active in public philosophy, writing for the New York Times, Washington Post, News Statesman, and elsewhere. Today, Professor Nguyen will discuss his new book titled The Score: How to Stop Flying Someone Else's Game. The book was published on the 13th of January and has since received excellent reviews. Congratulations, Professor Nguyen, and thank you for joining us tonight.

SPEAKER_00

Thank you. It's good to be here. Thanks for being here. Thanks to the royal tree of philosophy. Hi everyone, uh, I am gonna give you a bit of my so I have a popular book that has just come out called The Score. Uh the subtitle of the score as it's come out has been set by my marketing team. It's very uh I have a marketing team. It's very strange. Um uh this the public title is How to Stop Playing Someone Else's Game. But here at the Royal Institute, I get to give you my own secret, more academic title, which is an essay on the technologies of meaning. So, what I'm gonna do today is talk through some of the key ideas of the book and also give you a slightly more, because this is the Royal Institute, more aggressively philosophical angle. A lot of the things I did in the book I've made a little more kind of accessible, and this is gonna be a slightly weirder, more harder edge take towards some of those key ideas. So, this is the score. This is the book of mine that's come out, and it's it's about a puzzle. So half of my life has been about um working on the philosophy of games, and the other half of my life has been working on the theory of information and knowledge, and in particular, the theory of how we put information together as social beings, how we, as people that believe in science, work together and gather information together, how large-scale institutions like universities gather information. And one of the things I started to realize as I was working on this material for so long is that I had a weird puzzle. And that puzzle is that there are these things that I kept running up against, these clear scoring systems that told you exactly what was valuable and exactly how to judge yourself. And in some cases, in board games, in video games, in role-playing games, those scoring systems led to incredible delight and joy and wonder and fun storytelling. And in other situations, like in the university, in research context, in large-scale corporations that measured and gamified people's works, it seemed to deaden people's souls and wipe out what people liked about life. So this is the thing I'm trying to understand. And this starts for me at a particularly odd moment. So um uh a few years ago, I was sitting in a conference before about uh of computer scientists mostly and cognitive scientists who were building the machine learning programs that would make what they said was good art. Uh and I was I was the only philosopher of art in the room. This is one of the things I do. And um, this is before the current range of the various art-making AIs that we know now. This was about seven years ago. So I was listening to a lot of machine learning experts talk about how they're optimizing their connection as networks to make good art. And I had to raise my hand and I had to say something like, How are you defining good art? What is the target you're training your AI to hit? And what they told me was that they were training the AI to target uh whatever maximized engagement hours on the Netflix database. And I kind of lost my mind at this point. And I said something like, look, this is not, you're not actually aiming at good art. What you're aiming for is creating addictive, high-engagement, binge-worthy TV, which is not the same as good art. Uh and what they said, what they said to me was, well, show us another large-scale database that has a huge data set that we can harvest about what good art is. And what I immediately wanted to say was, there's no such thing. You're not gonna get that. You're not gonna get a large-scale data set, you're not gonna get a large-scale institutional metric that can capture what is good about art. And I've I believed that, but I didn't understand why. So, what a lot of this book is about is about why that's true. By the way, on this slideshow, what you see are two things. One is Bloom's taxonomy, which is something that is inside the American educational system that is designed to evaluate and often quantify student learning. And the other thing you see is House of Cards, old Netflix TV show. And the reason that is there is because, as far as I know, House of Cards is the first Netflix program that was heavily made with the contribution of large-scale learning algorithms. This is older than a lot of people realize. Uh House of Cards was originally made, it wasn't written completely by AI, but the themes, the director, the stars, the pacing, and the lighting were all selected by a large-scale algorithmic trawl of different programs looking to uh sorry, of different viewers looking to optimize for viewer retention. This has been going on a lot longer than we think. This has led me to think about something that I'm calling value capture. Uh, so value capture is any case where an agent's values are rich and subtle or in the process of developing in that direction, and then that agent enters a social environment that presents them with a simplified, typically quantified version of those values. And the simplified values dominate their practical reason in that domain. So, and by agent here, I mean you could be talking about an individual, but you could also be talking about a group, a department, a university, a newspaper, a government. All of these are the kinds of things that I'm worried about for value capture. So value capture includes explicit gamification, but it also includes academic assessment, research impact factors, retweet numbers, Facebook likes, numerical wine scoring, Fitbit numbers, quantified policing, and even money. Like, say, going into journalism out of a vast love of truth, but leaving caring about clicks, or starting to exercise for your health, but end up trying to max out your per-week step numbers, or trying to control your diet for health, but ending up obsessed with your weight and your BMI, or starting to learn out of curiosity, but coming to be focused on grades, or starting to tweet for the sake of truth or justice or connection or understanding, and then getting obsessed with producing viral content and upping your likes, follows, and retweet numbers. All of these are value captures cases for me. A lot of this started for me thinking about the university. So this is a third case study from my book. Um in my first academic job, I was the learning outcomes assessment liaison for my department. And what that meant was that I had to take philosophy students and I had to measure and quantify their improvement and report that up the chain for the Utah State Legislature's oversight and the Utah public's oversight. One of the important things is the kinds of data that they would accept. So here are the kinds of things that are acceptable as proof that you're succeeding as an educator in philosophy. The things like graduation rate, graduation speed, employment rate, income of recently graduated students, and student satisfaction is reported on evaluations. So I just want to note, I just want to point out that there are a bunch of things that are not on this list. For example, increasing thoughtfulness, increasing self-reflection, increasing intellectual virtue, increasing intellectual humility. There are other things that you might want out of university education too that aren't measured on here, like student community, connection, excitement, happiness. None of those are measured as student success. And this is one of the things I want to understand. One of the explanations for this is the university doesn't care about this kind of thing. That's one explanation we're going to look at. But the other kind of explanation is that they're systematically hard to measure, right? That there's some kind of gap between what is actually important, what we actually care about, and what is easy to measure. When we were looking at this list, one of the things that one of my students of mine once said was, you know what's not on this list? Actually learning and gaining skills. Which is wild, right? Because I think, let me just tell you something as an educator right now. If the major metrics for my performance, if the major metrics for my successful performance are my students getting good grades and good getting good high evaluations, student satisfaction evaluations on my student evaluations, I will tell you the easiest way to do that. Make the class incredibly easy. That also decreases student learning, decreases incentives to learn, and decreases skill development. That is, I think, a clear case where the easy measures that we have available come deeply apart from what we might care about. And the suspicion that I want to chart out for you today is that this is not just an accident of bad metrics, that there is a systematic way in which institutional metrics mute out and ignore certain key parts of life. I'm not going to argue today that metrics are bad or don't capture some truth, but I'm going to measure what I'm going to claim today is that institutional metrics, because of their basic logic, capture certain things easily, capture certain objective truths easily, but miss out on whole swa other swaths of human life. What I'm interested in is the gap between what's easy to count in a large-scale institutional setting and what's actually important. So I'm going to call this the gap. Okay. The gap between what's easy to measure institutionally and what's really important. And what we're looking for today is an echo is two things. One is an explanation of where the gap comes from, right? Why it is that institutions typically miss what's important, and an explanation of why it's so easy to be seduced by metrics, why it's so easy to let ourselves be guided by large-scale, simple, simplified metrics and not what we actually care about. I mean, and I've felt this too. I will say, like, I've had so many loops in my life where I start exercising for health and then get obsessed with losing weight. Or, I mean, these are all autobiographical, right? I went on Twitter in order to connect to people, and then I went viral twice, and then I just got obsessed with making the number go up. So that is the thing I really want to understand. So my account is going to be that the harm of value capture is that you are outsourcing your values. You're letting someone else, like Mark Zuckerberg, set the content of your values for you. But you might think, what's the harm in that? We are social creatures. We get values from the outside all the time. We learn values from our parents, from our community, from our religion, from our culture. I learned the value of improvisation from jazz culture and improv comedy culture, right? We soak up values all the time. But I think what's really important is that metrics don't just offer us external values, they offer us a specific variety of value. Something that is peculiarly fixed and rigid. And I think you might be able to feel this. There's something in bureaucratic life and institutional life that you can feel that has a smell, right? And that smell is what I want to understand. That there's a typical form to institutional metrics that arises from the logic of bureaucracy, from the logic of large-scale information collection by a large entity that misses something. So to get a better glimpse into this, I want to look at specifically at something I want to call mechanical scoring systems. I'm going to give you an account of what that is, but I think you have a gut sense, right? It's something that tells you what earns points and tells you exactly the rules for what earn points. And here's the rule in question. Because I think it's easy to look at this mechanical scoring system and think that's crap, that's not good values. But we have an important contrary data point, which is games. So I think any good explanation can explain two things. It can explain why scoring systems are often so fun in games and why they're so draining in institutional life. So here's a guiding thought. A scoring system is something that's designed to produce a single verdict. So notice, and I think this is something I missed for a long time writing about games. You can have a game or a competition without a scoring system. You can have a game, you can even compete without a singular verdict. So skateboarders, for example, kind of naturally before competitions, the before formal competitions, you can go to the skate park with your friends and compete for the coolest trick and have no official single way to settle it, right? You can actually come out with different views, right? We can go, Julie and I could go and skate with each other and a bunch of our friends. And each of us could come out with a different view about who had the coolest or most beautiful trick, right? You can have a competition without having a subtle verdict system. But in the history of skateboarding, it's the professionalization of skateboarding and the large money prizes that drags in the need for an official scoring system. It's when the ESPNX games happen and ESPNX happens, which is when skateboarding goes to the Olympics. I'm seeing the same thing happen, by the way, with a lot of other aesthetic hobbies. So what happens is when skateboarding goes from a kind of on the streets natural thing to a professionalized thing, we create an objective scoring system, and we do so by shifting what we measure. Skateboarding starts to become less about beautiful tricks, cool tricks, creative tricks, and more about clearly measurable outcomes, height and number of flips, right? So we shift from something more aesthetic and something more subtle to something more vivid and clear to count together. And if it's not clear, my worry is that large-scale institutional metrics do this for everything, right? That there's a shift towards what's easy to measure together in order to get a joint verdict. Um, so there are a lot of case studies in games. So I think classically, a lot of board games, this is one of my favorite board games, Tigris and Euphrates by Rainer Knittier. This game is mechanical all the way down. I think almost all board games are like this. The rules tell you exactly what the moves are, the rules tell you exactly what earns points, the rules tell you exactly how to add them up, and the rules tell you exactly how to win. There's no dissent and there's no room for disagreement at any point. That's a far extreme end. Uh, this is, okay, I'm gonna, I should I should apologize before I put this up. This is disgusting, but I think we need to talk about this uh because it's a really important example. Um when I was looking around for cases in which scoring systems changed in natural behavior, I found that there is a really good literature of sociologists and anthropologists that have studied pickup artists. So if you know pickup artist culture, it is the part of culture in which people, mostly but not all men, compete for success in the sexual and romantic arena. One of the things they discovered, that I discovered by reading a lot of these studies, is that pickup artists compete for very specific things. They compete for things like the number of phone numbers collected in a night, the number of sexual encounters in a night, and the top speed from meeting somebody to a sexual encounter with them. What researchers note is that pickup artists almost never compete for quality of relationship, but they don't even compete for pleasure. The sociologist Eric Hendrix says that a common theme in pickup culture is the need to sacrifice any interest in pleasure or happiness or joy in order to win. And here's the wild thing. When I started reading this stuff, I expected to find out that pickup artists were evil but at least having fun. What I found out was that they are miserable and unhappy because they have sacrificed interest and happiness or fun in order to succeed on some scale. And hopefully it's clear that the difference here between what's on the bottom and what's on the top of this list is that pleasure and quality of relationship are hard to publicly verify. Number of phone numbers in a night, speed, that is easy to publicly count. Right? So once again, just like with skateboarding, we get a shift away from what you might have thought was the important stuff to something else just because that other thing is easier to count together and in unison. And if it's not clear, again, I'm worried, this is not just pickup artists. I'm kind of worried that we are all in our various sectors education, business, journalism, becoming a kind of pickup artist. A third case, um, I think a really interesting third case is the case of Olympic figure skating. Because this is not purely mechanical. You inserted a judge. What you have is you have people that give what you might think is partially subjective scores, but then you have a s official system for picking the judges and adding up their scores, right? I just want to note this. We're not going to talk about these cases a lot, but I want to note that these exist, that you can have scoring systems that are not mechanical all the way down. Another interesting case is something like phishing. Because fishing in the net I fly fish. Uh fishing in the natural case is super interesting because there are actually multiple different scoring systems that stack, and you can rate yourself in different ways, and you don't have to pick which one is the correct one. You can aim for catching more fish, catching bigger fish, catching the most weight of fish, catching the largest single fish, right? There are lots of different ways to rate fishing. One of the interesting things is when you transition from the kind of natural environment to the tournament environment, you get a single system that tells you how to add these up. So in the standard American bass fishing tournament, the score is the combined weight of your top three fish. So notice again, in the natural environment, Julie and I can go out fishing, and I can catch more fish, and Julie can catch the biggest fish. Who actually won? We don't know because we didn't have a settled way ahead of time to put those numbers together. There's just different scoring systems that kind of float free. But in an official context, we have a single view that renders a single verdict. So here's an account of what a scoring system is. A scoring system is something, a system that offers a quantitative evaluation that creates a single verdict and enters it into some official register of record. So here's a short version. A scoring system is a social technology that engineers a convergence of judgments. If we all agree to the rule set, then we will all agree to what the official verdict is. We not personally think that that's the right verdict. So in a figure in figure skating, it's total you can totally say, like, look, the judges didn't get it right. Right? Somebody else was the most beautiful skater, deserved the highest artistic score. But if they follow the rules, you will all agree about who got the gold medal, right? Because there's an official scoring system that leads to an official register. So what I want to understand is the impact scoring systems have on our life. So a scoring system is a quantitative evaluation. So note, by the way, that not all metrics or measures have to be used as evaluations. It can be just used as data. It's one thing to use your weight as a simple piece of information, right? A mere input for a more complex decision procedure. It's another thing to treat your weight as a target and simply think moving up or moving down is the goal, and that's better in and of itself. So scoring systems are quantitative data turned into targets, turned into goals, turned into values. So to understand what that does to our values, we have to understand the nature of quantitative information. And here I think there's something incredibly interesting that we get from Theodore Porter. So Theodore Porter is an intellectual historian who's deeply influenced by the philosophy of science. And he gives an account in this book about why he thinks politicians and bureaucrats so compulsively reach for large, uh so compulsively reach for quantitative justification, even when the metrics are terrible, right? He so, and he gives us an account, and his account is not that quantitative information is bad. What he thinks is there are two ways of knowing, qualitative and quantitative. Qualitative is with words, quantitative is with numbers. He thinks each is good for different things. And the problem comes when we don't balance our approaches and combine our approaches, but compulsively reach for the quantitative. So here's what he says Qualitative ways of knowing are rich and subtle and nuanced and context sensitive and open ended, but they travel really badly between contexts. They uh require a high amount of shared background information to understand. So here's one way to put it think about. There are a lot of my students in this classroom right now. So we're going to talk about how we evaluate you. Right. So when you write essays for me, I will write rich, complicated paragraphs in response to your essays, in which what I will do is look at what you're trying to do and respond differently to what you're trying to do. If some people are trying to be creative, I'll respond about the creativeness. If some people are trying to analyze a particular paragraph, I'll respond about that. And I can use the particular language that we developed in our classroom to talk together, the language of our class, of philosophy, the particular language that we developed in our classroom. But that information will be incomprehensible at a distance, right? The business school dean won't understand it, a computer science professor won't understand it. And similarly, I won't understand your computer science professor's written evaluation of your code. I don't have that language to understand the context of what good code is and bad code is. So importantly, qualitative information doesn't just travel badly between contexts. It doesn't aggregate, right? So think about like my in a course of your lifetime as a student, you get thousands and thousands of written evaluations. There's no easy way to stick those together, and that's precisely because of their nature, because of their open-endedness, because different people are rating things in different ways on different qualities. They don't stick together, right? Quantitative information in institutions is completely different. Quantitative information, says Porter, is a case where we've identified a context-invariant kernel, something that's stable between contexts, and we all agree to keep that the same. So in the grading context, what this looks like is we all agree that what an A is very good, a B is pretty good, a C is average. And because that piece of meaning is so simple and so stable, then that information can travel between contexts. Does that make sense? So one way to put it is that quantitative information travels well because it's been engineered to travel. And the way it's been engineered to travel between contexts is precisely everything that's high in nuance, high in required context, and high in sensitivity has been removed from it. Right? So think again about the huge gap between written feedback, paragraphs of written feedback, and that simple scale of A, B, C, D, F. Right? And it aggregates well, precisely because we've stabilized because we've all agreed to keep that same scale. Since everyone can collect into the same scale, we can get all this information to instantly add up on a spreadsheet. So this is the way that Porter puts this is that information is a kind of knowledge that's been prepared to be understood by distant strangers. Sabina Leonelli, who is the foremost philosopher of data right now, puts it this way. She says, data is knowledge that's been prepared to travel to unexpected context to be used in unexpected ways by unknown people. This is referred to as the portability theory of data. And I think there's a core nugget of insight from Porter. And that core nugget is, and it's terrifying to me. That core nugget is the idea that information travels easily between contexts, and that's what gives it incredible social power. Quantitative information is made to be understood easily. And the thing that makes it travel easily is precisely the fact that it has all the nuance removed from it. Another way to put it is the kind of inhumanity of metrics and their social power come from the same source. It's they're does it make sense? They're not things that you can separate, right? The thinness is the design feature and the design weakness at the same time. It's what let lets it do the job. So um, so yeah, one way to put it is that data is powerful because it's portable, but portability requires decontextualization. And decontextualization is the thing that makes it socially powerful and communicate well and seize our kind of social attention. Um that's the first lesson from these people that study quantitative information. Here's the next lesson. This is from an incredible book from Jeffrey Bauker and Susan Lee Starr called Sorting Things Out, which is a study of large-scale classification systems. But I've come to understand reading about data and under studying people that have studied data is maybe one of the most important features of how we know things together in the modern world is hide under one of the most boring and innocuous covers. It hides under the name of a classification system. So here's the way they put it every data collection effort involves a classification system. So in grading, it's A in the American system, it's A, B, C, D, F, right? And what they say is every bucket emphasizes some information and forgets other information. So here is an easy case. The United States racial census has seven racial categories. It has white, Latino, Asian. This system remembers and emphasizes the difference between Asian and Latino, but by design, it forgets and loses the difference between East Asian and South Asian. Does that make sense? So it's got one bucket for Asian, which means that everything inside that bucket is lost. We've decided to forget it, and everything at the edges is remembered. And there's a reason for this. The whole point is that the world is too complicated, there's too much complexity, there's too much, too many spectrums, too many shades of gray. And so we need to chunk it up into lower resolution buckets in order to manage limited memory and limited cognitive resources, right? If you let everyone write what they thought their race culture and culture were on a sheet at the census, you couldn't add it up easily. You need rigid categories that cut up the world in a shared way. And one of the ways they put it is that the choice of what to remember and what to forget reflects particular interests. Whoever made the classification system draws those lines where they care about. But when we when those enter informational infrastructure, when they enter into the background, we forget it. We start thinking it's natural. One of my favorite examples is from their book. They analyze the ICD 10. The ICD 10 is uh the World Health Organization's international classification manual that codes different kinds of accidents and causes of death so that we can get large-scale statistics about mortality statistics, about what's hurting people. This is more uh when we do epidemiology, when we do large-scale uh big data collection about causes of death and what's really hurting people, this is what we use. And they notice in the falls coding, so there are different codes for different ways of getting hurt and killed by falls, there are highly granular categories for falls in urban environments. There's different codes for falling from a playground, falling from a commode, falling from a chair, falling from a bed, falling from a wheelchair. I haven't even listened to them all. There's also falling from a balcony, falling from an escalator. There's about 10 or 12 of these codes. For natural environments, they only have two categories: fall from a cliff and fall other. And this reflects the interest of the makers. Actually, everyone that made the ICD 10 is a doctor in a large hospital in a major metropolitan area, right? Like the almost everyone that contributed to this is from London, Berlin, Paris, New York, right? A very particular slice of the world. And the ICD 10 reflects their interest. And it reflects their interest. Here's another way to put it: if you have a large-scale data collection system like this and you want to go looking for whether people are in nature are more hurt by falls from trees or falls from hills, that data isn't in the system. You can't analyze that data because someone ahead of time has decided what the interests are of the system and what should be recorded. And they have to be rigid. One of the things is if you try to adapt the classification system to your local interests, then the data won't aggregate, right? So in order to get large-scale data collection, you need a rigid designation of what matters and what doesn't matter, and you have to keep it stable. So here are some lessons we've learned from Porter and Background Starr. Institutional data collection trades high context nuance in exchange for portability. That data collection system uh depends on categories which filter interests in a value-laden way. And those systems are rigidified to serve the needs of institutional stability. So mechanical, everything I've said so far, all of this, is about the nature of data, not about metrics. But in metrics and value capture, when we start treating data as our target, then they take that rigidity and pre-filtration and deconte decontextualization and they bring it into the process of evaluation. Another way to put it is all of that gets sucked into our process of figuring out what we should be doing and what our goal setting is. I mean, um I here's there's an example for all my students in the classroom from the University of Utah. The University of Utah has set this wildly insane schedule that forces students to take classes early in the morning and at night. And I think it's because that's optimizing for graduation rates and not for happiness and learning. And the reason is because graduation rates are an easy-to-collect decontextualized piece of information, but happiness and learning are high-context pieces of information that require lots of expertise to judge. So they're not the kinds of things that large-scale institutions can target. And in value capture, you as an individual might take this into your soul. This is basically what I'm worried about. That value capture takes decontextualization and places it at the center of your soul. So here's one motto Classification systems and the metrics they generate are forms of rigidified attention. Our values guide what we care about. And we build in ahead of time and let someone else set what we will and pay attention to, what we will and won't respond to, and what we will and won't target. Okay, that's the first stage. Let me get to moderate. To summarize that in one sentence, value capture leads us to decontextualize information on purpose. Okay. The next part of a scoring system is it leads to a singular verdict, right? Everyone has to agree on board. And this, I think, also leads to a lot of limitations in the kinds of things that large-scale metrics can target. So to produce a singular verdict that's generally acceptable, we need some agreed-upon procedure for evaluation. So let's focus on the simpler and more common case of judgeless mechanical scoring systems, the pure one, not like the ones with judges like figure skating. So the acceptability there is produced in significant part because everyone can inspect the scoring system. And you can check each part and see that the scores were uh were tallied up correctly. To understand that, I was stuck for a really long time until I found this book, uh, Lorraine Dastin's Rules, which I think is one of the most interesting books I've ever read in my life. Lorraine Dastin is an intellectual historian, a historian of science. And what she says is that there are different ideas of a rule over time. So let me remind you: the reason we're talking about rules is that scoring systems are shared rules for evaluation, right? They tell us what counts and they tell us to count it. They're rules for determining who won and who lost and whether we did well. So she says that um old kinds of rules, historically, the com most common kind of rule was a rule as a principle. So a principle is a generalization that admits of exceptions and grasp of the appropriate exceptions as a matter of discretion and sensitivity and expertise. It's high context, right? So a good example of a rule as principle is when I was in creative writing, people told me, show don't tell. Um, and show don't tell is a good rule, but it you don't follow it all the time. Some of the greatest writers sometimes tell instead of showing, right? Movies mostly don't have voiceover narrow narration that tells you things, but once in a while they do. And the reason is because that rule is not perfect, right? It's a rule that you can break sometimes when you understand the logic behind the rule. The second kind of rule, she says, is a rule as a model. So this is just a role model. Someone you're supposed to uh what would Jesus do as a rule as a model, right? Notice for rule as principle and rule as model, how we apply the rule is not mechanically determinant. It requires an exercise of judgment and expertise. Not everyone will apply it in the same way. The third kind of rule is a rule as an algorithm, a rule intended to be applied exactly as written, mechanically, with no deviations and no exceptions. And this has become so common that it's become the dominant conception of a rule in our time. But she says it's really recent. This is only a couple of hundred years old. Until very recently in the scape in the landscape of humanity, we barely had any algorithmic mechanical rules. Um one of the things she says that truly blew my mind is that people think that mechanical rules arose with computing machines, with computers and calculators. She says that's false. They actually arose a hundred years earlier. And they arose around an attempt to reduce the cost of labor by transferring work from extremely experienced, high-skilled, high-trained experts who are expensive and hard to replace, to low-skill laborers, right? That could just pick up the rules and follow them. So what she says is that algorithm rules are inflexible and unchanging. And they work really well. They're really efficient when contexts are stable and situations unchanging, but they do really badly when contexts shift rapidly and in very dynamic situations. So, uh, I found this a little abstract, but she gives an amazing example that explained a lot of my own experience of my life. She says, if you look at cookbooks, the look of a modern cookbook is actually pretty recent. If you look at older cookbooks, and I've looked at a lot of these, she's completely right. 20s, 30s, 40s cookbooks, they look like this. Old school recipes look something like knead the dough until it feels bubbly and alive, adding sufficient water and uh uh water and flour to keep it balanced, bake it in a hot oven until it makes a hollow ringing sound when knocked. A modern recipe, an algorithmic recipe, looks something like this. Mix two cups of flour, one packet of yeast, one teaspoon of sugar, and one cup of water until you can stretch the dough to a point of transparency, bake for 45 minutes at 350 degrees. Hopefully it's clear, right? The first is principles, and the second one is algorithms. And I think if you're used to old, to modern recipes, the first one might look primitive to you. It might look gross. Uh but if you actually cook a lot, if you have experience, the first recipe is actually better because it tells you when to be flexible. Because baking is a really complicated process. You don't actually want to add, if you're a really good baker, the same amount of flour and the same amount of water each day. It depends on the humidity, it depends on how your yeast is doing, it depends on the exact texture of the flour, right? Different days require responsiveness and flexibility. And the first one builds that in. The second one gives you an automated rule set that is inflexible but easy to follow. Also, I think a really interesting thing if you look at that first recipe, if you think about it, so bake it in a hot oven until it makes a ringing hollow sound when knocked, that might look really kind of primitive and weird to you, but actually that's a better rule because what you want from bread is holes inside, right? And listening to it is actually a better signifier. But that looks wrong to us. And I think the interesting thing for me is our sense of reality has been sufficiently captured that the better rule looks primitive, and what looks real to us is the rule that's easy to follow but less accurate. Right. So one way to put it is the upside of algorithmic recipes is that when they're ex when you follow them exactly, they give you the result to anybody. Anyone can use them, but they don't cue the adaptation and the discernment and the responsiveness of actual expert behavior. And so when you transform an old school recipe into an algorithmic recipe, you gain accessibility at the price of adaptability and sensitivity to nuance. Does that make sense? I think that's that is I think the key idea here. This is, let me just say it again, that when you transform a rule from something that's principled to something that's um uh that's mechanical, you lose sensitivity, but what you gain is accessibility. And that trade-off is essential, right? So one way to put it, that Dastin says, is that when we're looking at mechanical rules, what mechanical means is that anybody can apply it without significant cognitive effort, without skill, interpretation, or complex judgment. Mechanical application criteria make value assessment easy and portable. They make it so anyone can execute it in the same way and anyone can understand it, right? This the Dastin and the Porter accounts go together. What makes it cross-context is in part that it the rules are so mechanical. But there's a trade-off. Right? Part of that trade-off is dynamic responsiveness. So um, when I started cooking Italian food, uh, I use Marcello Hazan's cookbook and it works really well. And part of the reason it works well is it specifies specific kinds of tomatoes. It specifies San Marzano tomatoes from Naples in a can. And those are really good for cooking if you're following the recipe because they're stable. Every can tastes the same, no matter where you buy it, right? It's got the same distinctive flavor. And then I started cooking with tomatoes that I bought from the farmer's market, and my recipes started, my cooking started to suck. And the reason it started to suck was when you buy tomatoes from the farmer's market, they're actually better, they're richer in flavor, but they're hyper-variable. So to cook well, you have to adapt. When I was still following my non-adaptive adaptive procedure, right, it turned to crap. So one lesson here is that highly mechanical instruction sets work very well when the inputs are standardized and they work really badly when the inputs are highly variable and you need to adapt to them. And what's interesting is principal recipes hit the target better. Bakers who don't follow the recipe precisely actually do better when they're experienced because they adapt. And algorithmic recipes hit the target worse. They wander, they make less reliable bread, but they get you ready accessibility. Uh, you might think that the explicit version is more objective, but you have to be careful about the term objectivity. It's not objectivity because it involves accuracy to the world. Uh, Porter calls it mechanical objectivity. And what mechanical objectivity is, is a procedure that's repeatable with the same result when applied by different people in different contexts. And that's not the same as accuracy. I think that's really important to notice. This is like the, I think, the subtlest thought that comes from uh Porter and Dastin. So the legal system is a similar concept, says Porter, legal objectivity. So here's an example. We grant various rights to people, and what we're really trying to do is grant to them based on their intellectual and emotional maturity, right? Like you can't where I live, you can't vote till you're 18, and you can't drink alcohol till you're 21. What they're targeting, the reason why that line is there is because we're looking for a certain intellectual and emotional maturity. You don't want my six-year-old to vote. But 18 is not a perfect tracker for intellectual and emotional maturity. Some 17-year-olds are much more intellectually and emotionally mature than some 21-year-olds. And if I had my there are definitely 16-year-olds I know who I would trust more with a vote than some 30-year-olds. But that standard is not a mechanical standard, right? It's not someone that's repeatable by everyone consistently. 18 is an easily repeatable standard, even though it doesn't track what we actually care about. I think this is a case that makes clear why we sometimes want legal and mechanical objectivity, right? I don't want someone in the voting booth making decisions about who's mature enough. In this case, the trade-off of being objective in the legal and mechanistic sense is worth it. But the point is, does this make sense, that mechanical objectivity comes apart from we actually care about? And sometimes we might care more about mechanical objectivity, but sometimes it might not be worth it. Not every case is like the voting case. So why do we move to algorithmic rules? Dastin says the move to algorithmic rules permits the employment of cheaper labor, right? Because you don't need expertise. Elsewhere, she says that creation of clear rules serves the purpose of communicability, right? You can pass a when you have a complex baking procedure in an expert cook's head that's hard to pass. A recipe is easy to pass. And this permits, among other things, the fungibility of procedure followers. That means we can swap, okay, if you have a delicate, complicated, expert baking procedure that requires huge training, you can't just swap out bakers. If you have a clear, simple rule set like McDonald's does for cooking burgers, and you have a stabilized input set, then you can just hire and fire workers at will. Like any, you can slot anyone in, right? Because those rules have been made precisely to be consistent across different people. So here's your Scooby-Doo mask reveal. I don't even know if that term means anything to the current generation anymore. Uh the driving functional goal of mechanical scoring systems is worker fungibility. Clear mechanical rules make it easier to replace workers interchangeably while being consistent about how we're judging things, right? You if you swapped out the judger, you would not be able to consistently judge the growth in student curiosity in a classroom. You would easily be able to consistently judge. Performance on a standardized test or the graduation rate. Those things are mechanically countable. So scoring systems make it possible to replace workers easily with no interruption, with continuous functioning in the act of judgment of judgment and evaluation in an institution. So here's one motto rules turn people into parts. Does that track? Like what once we've standardized nuts and bolts, we can we can swap them out easily. And once we've standardized the rules for operation or the rules for judgment, we can swap out people easily. So if to sum up the big claim, there's this huge trade-off between accessibility and sensitivity. And mechanical scoring systems push us into the accessibility side at the cost of the sensitivity side. And when you're value captured by a metric, you take that trade-off into what you're targeting, into your values, into your soul. You start targeting what's accessible and not what requires high context and sensitivity to pick up. Graduation rate instead of wisdom. Okay, here's the big question. Is the mechanization of scoring systems inevitably bad? Um do we always have to desensitize ourselves in exchange for accessibility at scale? So here's when uh maybe some semi-rescuer comes galloping in. What about games? Because in games, you have clear rules for scoring. And Destin tells us what this means. I can quickly swap out people and they'll pick up on the same scoring system, right? Games let us jump into and instantly start evaluating in the same way, right? One way to put it is that mechanical scoring systems are like clear recipes for evaluations. They're recipes for values, right? Anyone can use them and evaluate them in the same way with the same trade-off. But there's an analogy. There are two ways to go wrong with recipes. One way to go wrong is never straying from them, because then you won't explore a space, you won't learn to cook things your own way. You're just stuck. Another way to go wrong with recipes is never to use them at all, because then you won't learn new styles of cooking. I I had a buddy of European descent who cooked excellent English-German stews and he hated to use recipes, and he decided to learn to cook Japanese about the same time I did. And when I learned, I followed a cookbook and I just let someone else tell me what to do, and I made pretty good Japanese food. And when my friend learned, he refused to use a recipe, and what he made tasted like an English stew with some shiitakis in it, right? Because he used his own instincts. He didn't change. Does it make sense? Recipes let you try something new in part because they're so accessible. So the right relationship to recipes is as an accessible starting point, but not to get frozen there. And that's a suggestion for how to use scoring systems too. They're an accessible starting point to evaluate things in a new way, but you shouldn't get frozen there. One thing to think about our games is that we should have reflective control of our scoring systems. We should pick scoring systems and pick the activities that suit our purposes and interests. One way to put it is we should choose our scoring systems. We shouldn't let our scoring systems choose us, right? And what reflective control looks like is easy to talk about in games. You can change games, right? If you don't like one game, you can go to a variation. I tried to exercise by going up mileage in upping my mileage in marathon running. It sucked for me. I hated it. Then I tried climbing. It was great for me, right? I shifted scoring systems until I found one that worked for me. You can also mod games. You can hack games, you can change the scoring system. One of my we can talk about this in QA, but one of my favorite scoring system variations is indie tabletop role players who played Dungeons and Dragons and were, some of them were like, this isn't working for us. And they looked at the scoring system and they're like, oh, you only get experience points for killing things. Maybe you want something else. So they created alternate systems in which you get experience points for creating drama in character and creating character tension in character, which is amazing, right? So here's one way to put it a distinctive danger of value capture, of taking metrics as a foundational element, is that they remove sensible, adaptable, dynamic value sensitivity from the reflective process. So scoring systems, mechanical rules are safe when they're under the control of a reflective procedure, and they're unsafe when the mechanical rules capture our core values. I mean, the simple way to put it is if you're picking your mechanical rules in games, because of a non-mechanical sense of fun or interest, that's safety, right? Your core values have not been mechanized. If you let metrics and mechanical scoring systems into your core values, then your primary method for choosing is going to be constrained by the constraint of decontextualization and hyperaccessibility at scale. And crucially, modding and changing aims to fit is possible because points aren't beholden each other. If I I'm terrible at platformers, if I play Mario Odyssey and I play the easiest person version of the game, and Julia's super good, right, and speedruns it, and then I get a hundred points on the easy version, and she gets 15 points speedrunning, and we're like, who did better? There's no way to tell, and that's okay. Because games don't have to interoperate. If you want this all in a slogan, it's that you can't house rule grades, but you can house rule Dungeons and Dragons. This is the dumbest place to end. I'm sorry, if if you you might all be disappointed, because this is the simplest possible motto, but this is where we end up. On the other side, our lesson from Porter is we have an explanation for you why you can't house rule grades. It's because institutional quantification needs to be stabilized at scale to perform its function of large-scale aggregation. We don't need to aggregate everyone's score in Mario. We do need to aggregate grades. You can't break away. There's a functional reason for that. Um, game scoring systems are mechanically clear, so we can quickly access them, but they're not rigid at scale, so we can fut with their content. Metrics, on the other hand, metrics push us towards a small number of metrics that are stable at scale. So what you can see now is the problem arises from two features. One is that the rules are mechanical, and then they're rigid at scale, as opposed to games, which are mechanical, but have offer high local variable control of scoring systems, right? You can house rule. So to sum up, with scoring systems, mechanicity blocks us from sensitively targeting a subtle value directly. But with games, you can approach a subtle value indirectly. What I mean is I can keep tweaking the way I fish to maximize meditative joy. That's okay. I have that degree of control. It's the rigidity of the metrics that ties them to some far, far away interest that blocks this goal. So the secret heart of scoring systems was worker fungibility at scale. The secret heart of scoring systems in games is what we call the magic circle. That is that every game is isolated and has a separate sense of meaning, and they're detached from each other. And there's no demand that they speak to each other and aggregate easily. Here's one way to characterize the difference, and I'm wrapping up. Okay. So Langdenwinner, the philosopher of technology, has this beautiful paper called Do Artifacts Have Politics? And he says that technologies aren't just neutral, they can change the world from their basic structure. So oral communication, for example, he says the technology of oral communication is highly decentralized. If I say something to Julia and she says something to her friend, et cetera, et cetera, et cetera, each person can change the informational content. The printing press, on the other hand, is highly centralized. And that's independent of how we use it. Even if you are, you created the printing press to be an anarchist and you print anarchist pamphlets in it. The basic logic of the printing press, that it's an expensive thing that looks different, transfers informational authority to a center, to the small number of people that are rich enough to own a printing press. One of his examples is that individual artisans, if you have individual shoemakers, they're free to improvise and respond in different ways to their improvisations. If you have a factory that makes shoes, you have enormous efficiency, but to do that, you need precise coordination. Everyone needs to do everything the same way, which means that you need an authority that centralizes decision making about what you're going to do. Does it make sense? Artistry is decentralized in decision making. Factories need to be hypercentralized in exchange for high efficiency. What about mechanical scoring systems? Games and metrics share the following: they produce a convergence of judgments for those that follow the rule set, and they do so accessibly. But games don't demand convergence at scale. They are decentralized convergence. They're a technology that leans towards decentralizing and customizing your sense of agency and your sense of meaning. Metrics are monolithically convergent. They shift control of meaning towards a center, towards whoever makes the Fitbit or the Facebook tweet. So here's one way to sum it up. Games are a technology that encourages play with systems of meaning making, and metrics are a technology that tends to centralize and rigidify the process of meaning.

SPEAKER_03

I used to be a philosopher, you know, in a department, um, and the decline of the world led to me leaving that, and I now work for the UK government um doing stuff very closely related to this kind of problem. So I work as an analyst for the government's property valuer, where my job is to help surveyors who are highly skilled, highly trained, qualified people work out how much property is worth so we can tax it. And so there's this pressure that we want to make computers do this rather than paying people to do it because computers can do it quicker and they can do it for every single property in the country more easily. Um and so I think this sort of like we really want to use this brilliant judgment that the valuers that we employ have, and we want to use data to make it easier for them to do the judgy bit rather than just the accumulating the evidence bit. And I guess the challenge is that it's not just about fungability of standards, it's this idea that when we're scaling it, we need them to be consistent, and so there's maybe this kind of middle ground where what we're wanting to do is kind of fix reference in a way that makes sense across people. So a bit like training everybody to cook Italian food in the same way, how can you like get these very abstract terms? What is the quality of a building in a way that people trained by different people in different parts of the country all coincide? So we want something that's not totally mechanical, but equally we don't want them all playing and modding by themselves, because then it's unfair to the taxpayer that their property got valued by a different method from the other people. So yeah, I was wondering like if you have any thoughts on what that middle ground looks like, because that's what I'm trying to achieve in my job, I think.

SPEAKER_00

That that's a great question. I I do think that I mean I do think it is fungibility and scale go together deeply. The reason that fungibility is important is scale and the degree of the degree of denuancing you get is scale-dependent. So one of the one of the crucial ideas for me is it's not any quantified system. You can create a very localized quantified system, and it can be quite sensitive. I think you're picking up on this. So if, like I think one of the things I've noticed is if you have five teaching assistants and one professor in a room coming up with a grading scale using examples, you don't need hyper-mechanical rules. Like you can coordinate pretty well, but that's because you're spending a lot of time with each other, you've got a lot of shared context with each other. It's really the need to scale past intimate or moderate scale context that really creates this thinning effect. Um, I mean, I my worry is that there is not in the middle ground you're looking for. And instead, what we have to admit is there's this profound trade-off. And the trade-off is between consistency and sensitivity. Um there's a neat historical example uh in uh Theodore Porter about how we used to measure land. So one of the old classic ways of measuring land, old, I mean like a pre-1200s, uh, a common unit of land uh in England was the hide. There are similar units of land measure across Europe. And the hide is the amount of land required to support an average family. So notice the hide is not a strict volume, right? A hide by a rich river is small, a hide by a uh in a forest is medium, a hide in a grassland is large, a hide in a desert is huge. It's also not quite the same as an average yield, right? So an area that's very fertile, four years out of five, but drops incredibly to nothing one year out of five, that doesn't support an average family, right? So it's a very functional measure, it tracks what we care about, but it requires rich ecological knowledge. And one of the things that Porter notes is that when we shift to large-scale centralized bureaucratic government, and this happened in England around like 11 and 1200, by the this is the timescale we're talking about. We lose the hide and we get start getting things like the acre and kilometer instead, because that can be applied by any land surveyor from anywhere. It doesn't require someone that lives in an area and really knows how that river works. So my claim isn't that this is bad, right? It's incredibly important. Uh, it's that there's a trade-off. So uh one thing I've been obsessed with, one of the things I didn't talk about that led into this discussion of metrics, is um some work I've been doing on transparency. And it's really inspired by the philosopher Anara O'Neill, who said that people think that transparency and trust go together, but transparency is actually in tension with trust, because transparency asks experts to explain themselves to non-experts, so they can't, so they have to make up stuff. Or if they don't, I think the other way to uh give the feedback is the transparency, transparency metrics ask groups of experts to justify themselves in terms of targets that anyone can understand, which absolutely fights corruption and bias, and at the same time forces experts to only to target things that are comprehensible to everybody. And what this often looks like is say arts, uh arts funding facilities, uh arts funding organizations being forced to, and this is in the historical record, target things like box office sales, right? Or things that have high-starred Rotten Tomatoes stars because those are the things that are accessible at scale. So uh sorry, if it makes sense. My view about transparency metrics is that there's a deep trade-off between trust and transparency. And if we have no transparency metrics, corruption and bias will run free. And if we have complete intrusion of transparency metrics, then we can't trust experts to do their special thing. And I think there's something really simple, there's something really similar. I see in education. Fairness, the standard of judging everyone by a strict rule everyone can see. I think I don't think that's the only value in the educational system. It's of value, but achieving perfect fairness requires that we trade away kinds of evaluations that require highly sensitive discernment. So I just don't think you can get a perfectly fair, mechanical, consistent evaluation procedure that hits things like being wiser in a philosophy paper or writing like richer creative writing, right? And so does that make sense?

SPEAKER_03

So my claim is just that we trained into exactly what I encounter in my job. I mean, so literally there's some legislation that was passed like last year on us disclosing how we valued people's properties. Right. Um so we sort of produce a thing, but that means that then we're able to use judgment less because we have to provide something that is explainable to a non-expert. So this is exactly what what I encounter in government.

unknown

Yeah.

SPEAKER_00

I mean that for me, the middle ground is balancing the demand of transparency and accessibility with the need for judgment. And the middle ground is the middle ground is not cost. Here's the technique. Right. My claim is not there, there's some I think my worry is that people have worried that the if we just aim for transparency and max out that, everything will be great. And I just think no, we have to admit that there's quantitative and qualitative, there's judgment and accessibility, and those are intention, and we just have to choose where the slider goes, and we can't optimize for both. And we're in a world that's tending to optimize for and tending to prefer highly accessible measures over every any any uh any other method of judgment.

SPEAKER_03

Thanks.

SPEAKER_05

And then you can go ahead.

SPEAKER_01

Um, thank you so much for this amazing presentation. Uh, I'm a high school senior, and we're actually reading your book, The Score for a class. So that's pretty cool. Um so my question uh is about what makes some kinds of value capture, like those in games, fun and others not fun. And um, your explanation kind of is that games can be modded and that the value you're maximizing, you're optimizing for um can change. But my kind of question with that is um take a game like chess, for example. There isn't really any modding of chess for the past few hundred years. Everyone plays by the same rules. And nonetheless, chess doesn't feel rigid. Um, it feels really fluid, dynamic, unpredictable when you play it. Um so I was wondering whether uh what makes chess fun is because of the strategies that it enables is dynamic, sensitive to different contexts, rather than its scoring system. So, like Bobby Fisher, for example, um, he really hated chess when it became a lot about memorizing fixed openings. And then later on, when those openings kind of uh when robots uh AI like came along and people realized that you know there's so much diversity in possible openings, um chess became more fun, again, it seems, to some professional players. And I wonder if something similar is going on with um value capture in uh in professional contexts. So, for example, um some you know uh value capture cases might make it extremely dull and boring to do a certain job because there's just one single dominant strategy, and that strategy is really rigid. But other cases, even when there is a very clear scoring system, like in investment banking, for example, the scoring system is quite clear, but the strategies people employ in the industry is still really adaptive and dynamic. So I was wondering if like the intermediate goals and the strategies um might have a role to play in determining what makes value capture problematic.

SPEAKER_00

Yeah, it's chess is a really interesting example. I mean, there are two levels of why games are there, there's a two-level response to why games are fun and most institutional metrics are not. One is the games have simple the scoring system has been designed and tweaked for fun and interestingness, and the scoring systems of other things have not. Like chess has been through like a millennia of evolution, where the reason for the evolution was to make the game more interesting. And I think like part of the background is what we've been optimizing for. But I think there's an enormous variation in the background that's hidden out of sight. First of all, by the way, it's just false that there aren't chess variations, there's thousands. Um, just search fairy chess or variation chess. I play tons of chess variations. Uh, one of my favorite chess variations is Bug House, which is a two-player blitz chess game where you're on a team and when you capture a piece from your opponent, you give it to your teammate and they can drop it anywhere. It's wild. Um, one of the most interesting pieces of history is that Richard Garfield invented the card game Magic the Gathering, which I think many of you probably know about, specifically because he didn't like the fact that you had to memorize openings in chess, and he wanted something as strategic. Deep, but had a luck variable opening. So part of the background is you're zeroing in on chess, but games occur in an ecosystem where if you don't like chess, you don't have to play chess. And one of the things that I predict is that if you don't have that freedom of choice, if for example, I don't know, in America, if you're in a small town and the only game in town is football, you'll have a similar experience to metrics. A few people might like it, but a lot of people, the majority of the people where it doesn't quite fit well, won't have the degrees of freedom. Where chess, as I experienced it, I really like chess. Other people don't. And in the background, we've had the capacity and the freedom and the degrees of freedom to shift to the kind of game we like, right? Like in my world, people who love chess play chess, and that's like 1% of the people. And other people play Dark Souls or rock climbing or gardening, or right. So that's I think that's really important. Similarly, for the financial world, there are a few people who I think are drawn to that world because the strategies are interest, interesting, and then there's a whole hell of a lot of people that are forced to play the money game, even though it's miserable and grinding for them, and there's no escape.

SPEAKER_05

Thank you so much. Chris, you can go ahead, please.

SPEAKER_02

Thank you very much. Uh thank you. That was a really uh fantastic talk. I've uh my mind is buzzing with all the uh implications of what you've said. I I'm from the uh I'm a doctor from the English NHS, um, and um I chair uh a number of hospital board quality committees. Um and so one of the things we struggle with in a in many ways is this sort of sense of um centrally driven metrics, um, as opposed to maintaining a sense of humanity in healthcare. So I just this is a huge field, I'm sure, but I just wondered if you could comment briefly about applications in the healthcare field and where where I might go next, as it were, to uh to to to look for uh um further further help.

SPEAKER_00

Right. There's actually actually there's uh the last third of my book focuses on health metrics as a core case. So and there's a but I'll I'll tell you some of the ideas there. I think there's a there's an easy and a tougher thought here. The easy thought is just I think most doctors, whenever I talked about this or on my doctor friends, they will immediately gripe about the same metrics. Uh, I think the one that, at least in the states context, people gripe the most about is speed of care. Like if you're being rated on the number of patients you see in a day, then you have low ability to interact with a particular patient. I mean, I think that that's obvious that the all the rest of the stuff I was talking about uh applies. I think there's a more uh spiritually threatening case. And what I'm worried in particular about is that large-scale health outcomes and public health outcomes tend to autofocus on easily countables, even really good ones. So um one worry I have is say uh public health decisions about uh say nutrition, it's really easy to measure effects on lifespan and heart attack rate, and really hard to measure effects on community, tradition, culinary joy. And so when you're evaluating something like Brie, high saturated fat food, it's I mean, my claim isn't that lowered heart attack rates are bad. It's that that tends to auto-win, right, in considerations. I think another one, here's my here's one of the most uh dangerous views I have uh in my uh political landscape. I'm worried during the COVID epidemic that we I think you can see where this is going, that lives lost and saved was very clear in things like community, mental health, child development, right? All of these things vanished from the calculus. And again, I'm not saying that saving lives is not incredibly important. I'm saying it's really weird when all this other stuff is very important that just vanishes from the calculus. So that's that's the second thing to say. I think the third thing, the to me, maybe the most intellectually interesting thing, is that I have started to suspect that health in and of itself is not the kind of thing that metricizes easily. Uh, and I found a really good explanation in a book. The philosopher Elizabeth Barnes has a book called Health Problems, which is an extraordinary book. And she gives an argument better than I have, so I just quoted it in my book about why health won't admit of metrics. She says, there's a kind of category that philosophers of language have identified of highly context-sensitive in language and interest-sensitive language. So context-sensitive language is like tall. What tall is for a mouse for a basketball player? So if you say, T T, are you tall? The question is, are you tall for a person? Are you tall for a basketball player? Are you right? Uh or in my case, are you fit for a philosopher? Yes. For a rock climber? No. I'm like above average for philosophers and well below average when I go to the climbing gym, right? She says there's a large number of terms that are interest relative, and she thinks health is an interest relative term. So she thinks, for example, her example is if is my knee healthy? That depends. If you're a 20-year-old Olympian, healthy means high functionality over the next six or seven years, long-term functionality less important. For me, my interests are more like I want to keep climbing into my 60s, and I can take a little pain. For other people, they don't care about, you know, climbing level functionality, but they care about pain-free walking. So each of these is a different. So if health is the kind of thing that's interest relative, then we it is absolutely the kind of thing that cannot be picked up by a large-scale metric. And I think the thing that I'm really worried about is that many, so the general lesson is what metrics are good at picking up are the kinds of things that are highly context invariant, longer lifespan, antibiotics making your infection go away. And the things that are really bad at picking up are things that are highly context variant and require deep sensitivity or familiarity with something like your own psychology to pick up. And so my worry is about a large-scale shift in attention towards what is shared in unison across all people, but that doesn't track what's important.

SPEAKER_02

Does that kind of answer your question? Uh no, absolutely. I mean there's there's I'm sure there's more fascinating discussions to be had there, but there's not time tonight, but that's really helpful. That's great. Thank you. Thank you. I'll get the book and read the read the last third. Well, I'll read it all, but in particular the last third. Thanks.

SPEAKER_00

Any other questions? Open all the questions, or I have just depressed everyone into despair and everyone needs to go home and hello.

SPEAKER_04

Um, thank you so much for a fantastic uh talk. Um just to give you a bit of context of where I'm from, my uh I'm a uh musician working in London. Uh and on top of that, I'm also a professor at one of the uh conservatoires in London. So I'm caught right in the middle of this um uh uh virtue capture and also the artistry side as well, and how you um how uh you know how you move between those two things. So on one side, I understand that in order for people to get their degree at the university, there has to be metrics involved. Um, on the other hand, though, I'm very aware that when they get into the business of being a uh working musician, that metrics play a part, but not the whole part. And so um what I'm trying to do is I'm trying to sort of get that balance, which talks me very interesting in terms of how I might be able to look at doing that. But um I'm slightly pessimistic about the future of the performing arts industry because so much of it is based on metrics now. And I've seen that through my own work and uh and and even popularity of certain concerts over others. Um, I suppose my question is do you see any future in which um this changes and maybe goes back towards more art uh art focused events rather than something that's based on metrics alone?

SPEAKER_00

Right. Yeah, that's a great question. I mean, it's I will say it's funny, like the stuff that when I talk about the stuff, the stuff that gets people the most excited about are bureaucratic and political metrics. But I started on art metrics. My first case was wine scoring, and I think wine scoring is a super interesting case. It's actually really interestingly connected to that health stuff I was just talking with Chris about. Uh so here's here's one of the things I found the most interesting about the wine scoring case. So wine is highly reactive with food, but it's very hard to get an objective, repeatable judgment because there's so many different kinds of food. And so wine scoring happens in a de-fooded context. It's an artificially created, uh fairer. Does it make sense? I think this is actually one of my favorite examples. It's a fairer context because how could you pick which dish to move it against? But to make it fairer and more objective in that sense, we've cut out of the judgment context the thing that was the most interesting, which is its responsiveness to food. And so the wine industry has moved in response to create wines that are much less reactive to food, but really good in the judged context of just having a few sips next to other things. So it's tended to move towards louder and less reactive wines. Um, I find that super interesting and super terrifying. Um about your case, I mean, it's so funny because there's a version of this which is like all metrics bad, all large things bad, destroying the arts, but it's not what I see on the ground. One of the things I see on the ground is it's really dependent on the kind of on the background context and the scale of the context. So, for example, you might have thought social media would destroy the arts. And one of the things, my own observation, which I think fits with what I think from the book, is the kind of arts that emerge from like the large scale, everything combined eye of Siron when we're the million level Instagram scale, that stuff is hyper uh is uh tuned towards the hyper accessible. The stuff that does really well are there are large scale, there are small scale arts communities that are and and we're talking about things in the thousand, five thousand range, where uh social media doesn't seem to have done that to them, and part of it is because I think doing well does not require um does not so I'm thinking about things like the skateboarding community. Skateboarding does really well on Instagram, and it does not seem to have been destroyed by Instagram, in part because the guiding sensibility underneath the scoring system seems to be guided by a rich aesthetic sensibility and doesn't seem to have been diluted by the demand for high accessibility at scale, right? It's like people that are in community who are hitting like and that kind of thing, and not like the gaze of everyone that doesn't understand what skateboarding is about. Um, so I don't know the uh maybe that's the best answer I can give. I'm simultaneously seeing the impact. So I I started thinking about this stuff because of social media, where I think social media is a highly scored communicative environment. And I see simultaneously the social media environment creating the possibility of highly centralized, highly accessible, dumbass videos, and also allowing for microcommunities that allow for the hyper-evolution of wild new art forms. And I'm tracking a bunch of them, like they're like indie tabletop role. They're all gonna look really disrespectable to someone that works in the fancy uh arts, but like indie tabletop role-playing is this thriving, exploding world that all lives on social media and seems to have had no particular destruction from that, but it's about 5,000 to 10,000 people. And I think that's really important. I'm not sure if that answers your question at all.

SPEAKER_04

Well, it was maybe, maybe, maybe would that be the the would the um uh the the correlation be maybe the some of the things that maybe that will in the future um be more prominent are those smaller scale things rather than say the mass projects, which might be more reliant on those sort of metric, you know, what's played on Netflix, what's played on uh yeah, video games, things like that.

SPEAKER_00

The key of my analysis is not merely the fact that there's a metric or a scoring system in the area, but the scale at which it is applied and the demand for access. That's the whole point of the games versus metric stuff, right? It's the scale of control. So if you tune, if you use Instagram to tune your scoring system to your tiny ass community of people doing some weird yo-yo tricks, this is one of my communities, uh, you're not gonna get the same problem as exposure. I suspect um, so I mean, yeah, like I don't I don't know what to say because I feel like my experience of the arts is that the central storm of social media is dumbing the crap out of it, and then I'm seeing more thriving weird corners than I've seen before. So I I I don't quite know what to do about that, except to say that the technology seems to afford both the death via scale and large and the splintering of communities into nice little weirdo local communities.

SPEAKER_04

That's all very interesting. Thank you very much.

SPEAKER_05

We have time for probably one more question if anyone would like to hop in. Um I'm thinking about when I was in my first year of undergrad, I accidentally signed up for a data science course and we were learning how to do like regressions and how to put different variables against each other. Do you think that there's any hope for kind of data science at large to become more qualitative, or will it always just fall into increasing scales of these kind of metrics?

SPEAKER_00

I don't know. I mean, the thing we're really talking about is the the social process and the authority. Some like, I mean, there is in my analysis, there's zero problem for generating a large-scale data set and carefully using it. The problem is when you just let it auto-set the target. And that's that's not a question about data science. That's a question about the large-scale reception and authority of data science in a large-scale social process of setting values. And I'm, I mean, I'll give you here. So, I mean, this is more an argument about what administrators and politicians do with the outputs of data science rather than uh about data science itself. But I'll give you my most, maybe I'll add here, I'll give you the most depressing argument I have in the book. It's I wrote it and it depressed the hell out of me. Um, the argument is uh for an effect I'm calling value collapse. And value collapse is the following. Imagine there's some gap between what's easy to measure and what's actually important. And imagine that social rewards go to people that do well on the metric, on what's easy to measure. Then you should expect the people that perform the best are the people that are willing to tear out of their hearts caring about anything rich and just hyper-target the metric. And people that still are torn between the metric and other stuff that's measured that's not as important are not going to do as well by the metric. So, insofar as the metric gives people social power and social resources and financial resources for doing well by the metric, then we will have a long-term selection process by which people rise in power if they're willing to ignore what's really important and hyper-target the narrow metric. And if such people tend to uh reinforce the grip of the metric, then we'll get a feedback loop where we just reinforce power over people that are willing to hyper game narrow metrics. There's there's a worry.

SPEAKER_05

That is quite worrying. Thank you.

SPEAKER_00

Shall we call it?