
The Gradient: Perspectives on AI · 2025-11-26 · 59 min
Gabriel traces his intellectual development from concerns about global poverty and inequality toward a philosophy of AI alignment rooted in basic human rights, equal respect, and the rejection of domination. Rather than treating alignment as a technical problem with a single correct answer, he advocates for a pluralistic, proceduralist approach inspired by Rawls's overlapping consensus - where people with fundamentally different values can still arrive at fair principles through deliberation. He contrasts this with agonistic pluralism (citing Chantal Mouffe), arguing that cooperation and compromise are achievable even amid deep disagreement, and illustrates this through real-world examples like hospital AI systems that get revised through complaint and deliberation. Gabriel also addresses the stability problem in social contract theory: even if people agree behind a veil of ignorance, they may defect once self-interest and passionate beliefs resurface. His experimental work with Laura Wiedinger, Kevin McKee and others on veil-of-ignorance experiments found surprising stickiness to agreed principles, suggesting the communicative and justificatory dimensions of human nature run deep. He emphasizes speech act theory and pragmatics as central to understanding alignment for language models, while acknowledging that language itself encodes inequality - particularly in non-literate contexts or caste-based societies - and that true alignment requires attending to who gets to speak and be heard.
Gabriel advocates for a pluralistic, deliberative approach grounded in basic human rights and equal respect, where people with different moral beliefs negotiate fair principles through overlapping consensus rather than attempting to encode a single predetermined set of values.
Gabriel argues that while motivational challenges remain when the veil is removed, experimental evidence shows people exhibit surprising "stickiness" to principles they've agreed are fair, because human communicative and justificatory drives run deep enough to create real motivational force.
Language encodes deep status hierarchies and power inequalities, especially in caste-based or non-literate contexts; genuine alignment therefore requires attending to who is empowered to speak and be heard in deliberative processes.
When AI systems are deployed with one goal (like resource efficiency) but generate unforeseen harms, communities make claims against the system, and iterative deliberation leads to revised operating principles that balance multiple values - an adaptive approach Gabriel calls "alignment in the wild."
Mouffe treats politics as irresolvable power contests with only temporary truces; Gabriel argues that while contestation exists, genuine cooperation and fair compromise are achievable because stories about what is possible shape what actually becomes possible.
Computed from the transcript - who did the talking, and the words that came up most.
Episode 143 I spoke with Iason Gabriel about: * Value alignment * Technology and worldmaking * How AI systems affect individuals and the social world Iason is a philosopher and Senior Staff Research Scientist at Google DeepMind. His work focuses on the ethics of artificial intelligence, including questions about AI value alignment , distributive justice , language ethics and human rights . You can find him on his website and Twitter /X. Find me on Twitter (or LinkedIn if you want…) for updates, and reach me at editor@thegradient.pub for feedback, ideas, guest suggestions.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hello. It's been a little while since I've done an episode. Life has been a bit busy recently, but I had the immense joy of speaking with Yassin Gabriel, who leads ethics research at DeepMind, about a pretty broad sample of his work and his thinking. His work has covered value alignment, which is the question of what values to encode in AI, where he's argued for a pluralistic approach which involves treating different people and perspectives fairly. He's also written about language ethics, where he's highlighted the need for an ethic of civility, about technology and society, where he's argued that, uh, AI generates questions of distributive justice, and about AI agents, where he's explored their wider societal effects. Today we'll touch on what value alignment is. Various challenges to Yesen's thinking and some of the foundations behind it, including what it means to view technology as a world making activity. I've been a fan of his for a while and think he does really, really important work and that if you at all like this episode, you should go and read some of it on the off chance you're listening to this for the first time. My name is Daniel Bashir, and with all that, here's yes, and Gabriel. I know this has been like months in the making. I kind of feel bad that it took us this long to finally get to recording a chat.
Speaker B: So.
Speaker A: So I have been eagerly anticipating speaking with you over the mic. I suppose I want to start with, uh, a framework question. As I was thinking about what I wanted to ask you first, I realized there's like a version of this conversation that starts practical and then gets increasingly less practical. So we'll see if that's what we end up doing. But I'm wondering if you could give me a sense of how you describe your personal intellectual development over time. And in particular I read your work as really grounded in a, uh, particular tradition and a set of thinkers who reappear throughout the years. I suppose I'm wondering if you have a sense of. For example, Rawls is very commonly cited and lots of discussions and things that we're going to get into. But why you feel like you look at the sorts of problems you're faced with and the ways you do.
Speaker B: Yeah. Thank you so much, Daniel. Really lovely to talk to you about these questions. So I think my animating concern when I started this philosophical journey many years ago was primarily with, uh, the problem of global poverty and global inequality. And it seemed to me that the starting point of any serious moral question or moral Engagement was actually with the idea of basic rights. So I could see that there was a lot of philosophical disagreement at the edges. For example, utilitarians and, uh, contractualists having these very, very long running discussions in moral philosophy. But when I wanted to reach for that was kind of more fundamental and got to the heart of what matters and served as a really important heuristic for understanding this world. Just the very simple idea that human life matters and that, uh, in light of that value, people are well positioned to make claims upon one another. Seemed like a really powerful bed drop to build upon, and one that had cross cultural value, obviously reflected in the idea of human rights. So I always approach moral and political philosophy with this idea that we should build up. Instead of starting with a fully perfected vision of what it would look like to live together, we should start with the things that we have relative confidence in and then build out from there. So from the idea of equal value, we get to the idea of equal respect, right, that we should respect one another. And then we ask, what does it mean to respect someone as a human being? Well, it means to respect their consciousness, right? The fact that they may have different views from yourself and to feel this kind of motive that you still want to live together in a kind of a healthy way that ideally both of you can support. So the thing that's really an anathema to that is a vision of domination. I deny your value by dominating you or the withholding of aid in situations where I could easily prevent you from incurring a grave harm. And so a lot of my philosophy, which carries over to the world of AI just starts with those fundamentals and then builds up and asks, well, how can we deal with these very prescient and very complicated questions that we face now, still speaking this language that ideally the largest group of people can actually understand.
Speaker A: One of the challenges you anticipate in one of your papers that I wanted to get into a little bit where I believe this is your paper with Atusa Kasserzadeh, who first authored this really great paper with you on aligning language models with human values. You get into a lot of really interesting ideas around speech act theory. And in particular you anticipate, and I think this applies to what we're saying right now, this tension between democratic civility and a criticism from the angle of agonism, where an agonistic pluralism, and you cite move directly here. There's the view of democracy as this contest between fundamentally irreconcilable, uh, political positions. And you bring in this notion and a lot of your work around overlapping consensus, people come to the table with differing opinions. We would like to construct some sort of procedure that people would recognize as generally fair and would generally not contest. And so it's less that there's this fixed set of principles, but the way in which we get there is really, really important to get right. And it feels like agonism is something where I feel you have some different angles to bring that in or address. Ah, that in an interesting way. But I'm curious, I have in my head sort of like, what is Yassin's response to this? But I'm curious how you would think through it.
Speaker B: Yeah. Thank you, Daniel. Another very interesting question. So I do believe in contestation and fair process. My belief is that when we put those two ingredients together, we can get something that's basically better than the sum of the part, which is principles that different people agree with and can recognize the fairness of that compromise. Now, I think for Chantal Mouff, the world is a lot more bleak. Right. She essentially says that democracy is a contest and a contrast between parties who essentially engage in like a power negotiation. There's never a real willingness to, uh, compromise. This is always just ideological masking. And at the end of the day, the best we can get is a status quo that is potentially in flux and doesn't privilege one party in perpetuity, but is amenable to revision. So that's the kind of agonism you want to lean into the tension with regards to that philosophy. I think it might be true in certain situations. There may be situations where civility, understood as the willingness to kind of treat people with respect and not play into power asymmetries, just kind of falls away from the table and you end up with a more violent vision of politics. But I don't see why we have to take that as kind of axiomatic that politics can only be that. I've seen many, many examples in the real world of people compromising and coming to live together and form value based communities that accommodate disagreement but also have toleration as an aspect. And so if that's possible, I don't see why we would want to leave it off the table just because we have a kind of dark and potentially morbid view of human nature or human group nature. Uh, one interesting thing is these stories that we tell about politics also shape our understanding of what is possible. Everyone who does political philosophy, you know, we can think about it as an enterprise of, in some ways its rational argument, but it's also constructing a vision of what it means means to live together. And if we only have a philosophy of conflict, then we will have a world of conflict. But if we have a philosophy which holds out the potential for absolute gains through cooperation, then I believe we're much more likely to unlock genuinely positive states of affairs, uh, both in the realm of AI and more broadly.
Speaker A: I'm, um, with you on the. This feels like a very. Perhaps cynical or a bit of a negative philosophy. I do think it also relates to this question around. In this realm of overlapping consensus, how much overlapping is possible? And one of the things I move points to, and maybe you don't have to accept that this has to be borne out in such a conflictual way, but you might accept that there are totally incommensurate values. Different people coming from different religions, different cultural backgrounds, might come to the table with very different readings of what a situation is or means or things like this. Or they might just have political angles on a particular situation that just cannot be resolved with one another. And so in trying to ground that out together into an overlapping consensus, you do have to make some decisions one way or another on those issues. And so you would have people who might just fundamentally disagree with that, or you have a version of that, perhaps that tries to leave some openness to this, and you can still maintain a situational, culturally specific form of this or something of that nature. I realize I'm talking in huge abstractions here because I'm trying to avoid saying anything that's a little too shaky, but I'm curious how you think about this and that. Uh, I feel you Also, my predicted sense of what you would say here is something like, there are different norms across places and things like this. And so a procedure that tries to identify a set of values that many people with share generally agree to can also take into account different cultural situations and different relationships to technology that a particular area would have or something like this. But I'm curious what your considered response to that would be.
Speaker B: Yes, I think that is roughly the way that I would understand it. So my philosophy does make these broadly humanist assumptions at the starting point. So it is true that if you reject the idea that other people's lives have value and they're just as real as you are, then it's going to be very hard to get a kind of cooperative enterprise off the ground. If you understand that there is a need for cooperation and also justification, then we can set up these practices of essentially discourse and exchanging points of view and things like that. But the output isn't always going to be something that's really, really liberal and uh, you know, has this kind of Western flavor of maybe individualism, etc. Things like that. Really the process should admit of different outputs. So if this were to take place in a society where maybe it was more communitarian in character, maybe if people were bringing different kind of religious ways of being to the table, you can always imagine, uh, a conversation between three interlocutors model their identities differently. You know, this one is a Christian, this one is a communist, this one is an anarchist. What agree, I don't know ahead of time. And it's hard to say how thick the output will be at the end of it. It may be something quite minimal. It may be uh, a kind of a mutual respect, or there may be more interesting things that they can do. So sometimes even when you have radically different moral beliefs, you can engage in a form of trading with one another. So, you know, you have your God, I have my God, and hopefully it's not the same day of the week that we want to dedicate to our gods. If so, then we come to the agreement that we each do our God on that day, otherwise we alternate and I might do your God, you might do my God as well. We may come to a state affairs that's better than just me practicing by myself. So I think sometimes because people see the impetus as this kind of liberal, egalitarian one, they underestimate how wide the range of outputs could be in reality. One of the, uh, concerns that I have with many views of AI alignment is that they're quite theoretical and they don't really apply to real world, uh, situations. In my latest work I actually reflected a lot on how technology, technologies are deployed and how then how they become aligned in the wild, right? So this isn't just kind of big AGI. We can also think about AI in the context of healthcare or AI in the context of criminal justice. And typically what you have is an AI system that's designed with one goal in mind. For example, the efficient allocation of resources to patients. That's deployed in a hospital context, um, but has a lot of unforeseen implications, right? So what happens is that people amass claims against it. They say, well, you know, it was really efficient, but I'd been waiting a lot longer, right? And I'd like this to be registered. And then if we were to run a kind of deliberative assembly type thing or uh, a citizens deliberation, people would make these complaints or claims and Then what typically happens is you revise the operating principles for the system in such a way that it leaves less kind of moral surplus out there, and you strike a balance and you say, okay, we're going to do a efficiency, but we're also going to incorporate waiting times. And so the principles that you need to run a hospital and society are quite different. But this is really quite a flexible way of doing alignment that I'm pitching.
Speaker A: Yeah, I think the flexibility is important here in that it's very much impossible to set out in advance exactly what the right set of values is going to be. There's a connected problem to this that concerns Rawls theory of justice on its own terms that I think is quite interesting. And because you sort of bring this in as, uh, one part of your theory about how we ought to come to a directionally good set of values and ways to deploy these systems in the world. The criticism for Rawls here, which I think is interesting, is there's two related problems for kind of normative accounts belonging to this family of social contract theory like Rawls. One is the justificatory problem, which you spend a lot of time on. On what basis can a set of normative principles about what is right and good basically be justified to some set of people? The second one is this stability problem, which I think is more interesting in our conversation, which is, what are the conditions under which the members of the set of people that you have justified these set of normative principles to? Under what set of conditions will they abide by those normative principles? And thinking about this from the angle of something like the veil of ignorance, we have people settle on conditions or values for society, and they don't exactly know what their position in society is going to be. With the intuition, then, that they would proceed to decide in a way that feels more egalitarian, more fair, to give a gloss over what is thought in a much more careful, um, way. The threat in the stability sense is once you remove the veil of ignorance, then each person who made these double deliberations, uh, is then going to exist in that society and pursue their plans as conceived by their conception of the good life. And so this may or may not depart from where they originally were in this original, uh, position or behind the veil of ignorance. And I think some of what you're saying also stems or has a relationship to the stability problem. And I'm wondering if you maybe isolating the social contract question from the AI question for a moment. I'm curious how you think about this, and you Feel like that stands as an interesting threat to you against the social contract theories or how you think about it.
Speaker B: It's a very interesting question. So just to do a little bit more groundwork. So the connection between the veil of ignorance and what we are talking about previously is that, uh, one way in which to arrive at these solutions to alignment when people have different material interests and moral beliefs is to ask them to imagine that they're in a situation where they don't know how they will be affected and to choose accordingly. And so that could be done on a very local level. We could think about that for hospitals, we could think about it for states. That's a social contract tradition. Or we could think about it for AI and AGI. And then on my understanding, your question as well, what about if we remove the veil? And then people are like, but I am really self interested. I ended up in a good situation. I'm no longer feeling the force of this thought experiment or oh my God, I actually have really, really passionate religious beliefs or strong philosophical views. I no longer want to go along with what I conceded or admitted was a fair solution, uh, ahead of time. And I think there's different responses to this question. So the first is that we probably will have motivational challenges, but that doesn't fully undermine the value of the thought experiment. So one of the criticisms that people often make of Rawlsian theory is that it is what we'd call ideal theory that is idealized, that it imagines that people can do this and be motivated by it, whereas in reality there's just a huge amount of complicating motivational factors, structural obstacles. And I think that this criticism is only partly well taken because I think getting clearer, uh, about what a good solution looks like actually creates the critical distance we need to properly evaluate the present moment. So if we're stuck in an incremental mindset, we may think it would be better to go this way. But we can't say, wow, we're really a long way short of what fair treatment requires here. So I think on one hand we've created a very valuable yardstick. On the other hand, I do think this kind of like moral reflection does have motivational power. And this is more tangential because it was just research that we did in an experimental setting. But I did actually run some veil of ignorance experiments with some of my colleagues, Laura Wiedinger, uh, Kevin McKee and others. And it was an incentive compatible experiment. So people had a real monetary incentive to kind of roll back the decisions they'd made behind the veil of ignorance if they were just playing the game rationally, but having gone through a first version where they selected principles without that knowledge, we found that the principles were qu. Sticky. So people didn't want to, like, declare something was fair and then immediately defect from that agreement. And for all the discussion of kind of human beings being fundamentally economic or fundamentally rational, it's also true that the kind of justificatory and communicative nature that we exhibit runs very, very deep. And I think it is a fundamental part of our social nature that we want to be able to justify, uh, our conduct to others. And it seems likely to me that so long as we live in a shared social reality where we want to kind of look each other in the eye and know that we're on good terms, that these kind of principles are not motivationally inert.
Speaker A: Yeah, I think there's an interesting epistemological angle there. One of the points that you raised that also I think applies here in your paper with ATUSA on aligning LLMs is also this notion of, well, when you're thinking about speech act, there's a part where I think, if I'm recalling correctly, you talk about how some of what you say in there would need some extension to consider things like oral cultures or different forms of discourse and ways of making oneself legible. And I think that's quite interesting, especially in light of what you were just saying about the ways in which we justify ourselves to one another. Because I think that for the most part, and a bunch of people have written about this quite a bit recently, if we were to take. I don't know if you've spent time with Walter Ong, his book Orality and Literacy, which is like, I think, you know, one of, like, the really, really great books I've ever read, but it's discussion of our shift from an oral to illiterate culture where writing exists at all. And we use this as a technology for thought and a technology to express ourselves to one another. And it admits of a very particular kind of thinking and mode of expression and mode of communication and things like this, which is to say that because we're in a culture that is broadly literate, and, you know, this is with caveats, I'm making a generalization here. It does not cover every particular case we might care about, et cetera, but in a broadly literate culture, there are very particular ways in which we make sense of the world around us and what others have to say. And, like, what counts as knowledge and what counts as uh, valid discourse and things like this. And for the most part people understand this and abide by it and express themselves in that way. I'm curious if that reads as broadly correct to you or how you also think of this question.
Speaker B: Yeah, so I think that the communicative nature of our reality was not something that I had fully engaged with until I started thinking specifically about the ethics of conversation. And I came to that because of the advances in the realm of chatbots and things like that. But, but when those kind of technologies came to the fore, it raised this interesting question from the alignment point of view, which is what is good speech? And what I did with ATUSA was investigate kind of speech as a practice and this version of linguistic theory called pragmatics, which takes as its kind of foundational anchor point the idea that every conversation has a purpose. Right. It's impossible to understand what's going on without assuming that there's some kind of shared mental model of where you're trying to get to. Even conversations that are uh, disagreements, they have a clear structure. Right. It's not the case that you suddenly switch languages or you know, you just repeat the same word ad, uh, infinitum or things like that. So there is this kind of deeper, ah, realm of cooperation which is almost always there and linguistically reinforced. Now that's not to say that all conversational practices are uh, just in fact, language also encodes tons of inequality. Right. So some famous examples, you know, we contrast civility as uh, supposes that basic conversation between equal parties with say some of the linguistic formulations that you'd find in a caste based society. And it's really, it's shocking like how communal we are. It's also shocking like how deep kind of inequality or like status hierarchies can go when they penetrate the realm of language. Right. You know, I can only express myself through the grammar that is afforded me in my position. So. So insofar as that's true, I think I also agree with the kind of the angle that I took you to be gesturing towards, which is actually that when we think in terms of emancipation and emancipatory terms, we do have to dig very deep into language itself and say, well, you know, when we look at the modes of communication, who is able to communicate in this way? And of course writing, not everyone is equally adept. And I think that's uh, a justified criticism of a lot of Western philosophy. You know, it's just philosophy for writers basically. But the conversational ethic thing came full loop in the later work, because that's where most recently I've kind of very, very actively stressed this speaking in your own voice and the right to be heard as part of a kind of democratic alignment protocol. So eventually these dots kind of join up hopefully.
Speaker A: Yeah. What you're starting to say, I think also gestures at uh, a particular theory that you reference in this paper and you call it out as one that you don't uh, engage with directly because the limitations of space, which is Lator actor network theory. But I feel like some of what you're starting to say directionally engages with this view. And I suppose if I were to think about it with actor network theory being this sort of descriptive or diagnostic way of thinking about how do heterogeneous networks, so this is not just people, but artifacts and policies and all these other things, how do they produce outcomes? And I suppose if I were to put this together with your thinking, I'd end up with something like alignment becomes a property not just of uh, an agent or an agent's text, but to a property of the whole sociotechnical network that makes that text possible. Or something like this, I suppose is how I'd start to frame it. And so you'd think about humans and non humans like the agents sort of being actants in a dynamic network. If I'm going to use sort of Latour's way of thinking about this, I'm again, I think getting into very abstract ways of framing this. But I'm curious, I feel like you're already beginning to engage with that theory and I wonder to what extent A, you feel like Latour's theory is a useful way of describing things for you in general. And then B, maybe how you'd frame some of your thinking in this paper along those lines if you were to do so in a more extended way.
Speaker B: It's a very deep question and I should be upfront that Latour is not a philosopher who I kind of read very, very deeply. But I guess one of the key connective points is with the idea of AI agents that a of people are talking about now, maybe there's two things we can talk about before agents. There's the question of technology and value neutrality. And I think the Latour view of the interconnectedness of things and things that are seemingly inanimate, shaping our life chances and affordances in different ways, is quite profound. And it comes up for AI because sometimes there's this idea the technology is neutral, it's what we do that makes a difference. And really I think that's just an inaccurate characterization of reality in the way that kind of Latour's framework would make clear. Because every time we create a new technology, there's new options that are opened up in terms of human potential and new things are foreclosed. So it's very hard to go back to world without social media. Now it exists, we have some leeway over the direction of things, but there's certain things that can no longer be ruled off the table so easily. Then that brings us on to the second thing, which is AI agents being more kind of active participants in our, uh. Well, we could start by saying our social world, but I think people are really interested in the material world. And so in other more recent research, I've been looking at this question of AI agents and what it means for an agent to have a high level of efficacy. And so, you know, of course we've all seen virtual agents doing things in computer simulations. AlphaGo was, uh, a powerful agent that played the board game Go and simulation. We've seen chatbots be more agentic, they have more persistent personality there, they have memory. But once we have real robust connections with network tools, um, physical actuators in the world, then as I've said with my colleague Jeff Keeling, the fourth wall comes down and we really encounter AI in a new way. And I imagine that that's something that people who are interested in actor network theory would almost be relishing the opportunity to study those new kinds of encounter and engagement.
Speaker A: I think with the technology and value neutrality connection, this might be a good place to start talking about value alignment. I feel like we sort of discussed things that are very adjacent to the issue, but not exactly broaden or define value alignment itself. So this might be a good thing to do. And you have a number of papers where you address this. I particularly liked your paper, the Challenge of Value Alignment from a couple of years ago. And I see you sort of bringing this in in sort of two separate parts. Right. We have this technical challenge of trying to align AI systems with human values. And this is very much of the AI safety so considered, and also the normative question that you're interested in, that we've also been discussing in a lot of different ways of what or whose values we try to align AI systems with. I'm wondering maybe as a starting point for this, as a thematic discussion. There were a lot of things to say here, but I suppose if you were to maybe express your starting points in thinking about this and maybe how you've thought about it over Time and what ways maybe your views on it have evolved. I'm wondering what your version of that primer would be.
Speaker B: Yes. Thank you so much. So I think that the starting point for, uh, a lot of my own investigation is the idea that we will never know the truth about morality, if it exists. And that even if we did know the truth about morality, we probably couldn't persuade other people of it just by telling them our version of things. And hence, we need to live in a world that is robustly populated by conceptions of value. We all have things we care about, strive for, belief should be promoted and protected. But also one where you never get the cheat sheet handed to you. And people say, okay, this is it. Now we can reach agreement. And hence, that view of alignment as the truth about morality just hits a dead end. Either you can see that in the beginning, or you think you might find the truth about morality, and then you become, I would say, potentially dangerous person who thinks you can impose it on people who disagree with you. So I wanted to kind of find a way out of that conundrum. And the way out of it is through a fair process, as we've already touched upon, so one where people find a common currency with which to talk about the impact of AI, its effects upon us. This one hurts me in this way. This one restricts my freedom in this way. And we begin a process of trying to find out whether there are principles that help us live together to tie together the earlier threads. One vision is an overlapping consensus view. So we could start by making a list of things that we agree about, if there's any, that would be a good thing to get us started with. We could engage in hypothetical thought experiments and see if perhaps when we get greater distance from ourselves, we find that there's more that we agree about than we start with. Or we could do something more robustly democratic, right? We could have a democratic discourse. We could actually vote on things. And I think all of these things have potential when it comes to artificial intelligence. And basically very heartened to see the kind of democratic most of AI alignment really pick up and gather steam. So, of course, hopefully, the audience has seen the amazing work by Divya Siddharth and Safran Huang on alignment assemblies. And I think that that has a world of potential. It may be that the moment for democratic alignment hasn't fully peaked because the technology isn't as powerful as it may be in the future. But I think when we see very, very powerful technologies, then the need for public justification will be stronger. And more evident particularly if isn't aligned from the beginning and there are resulting harms now. Building out beyond that, I increasingly think that the notion of value alignment contains multitudes basically. So I wrote that paper, I believe in 2020, and since then I just see more and more alignment problems popping up all over the place. So for example, at that time I made the argument that people would focus too much on one to one alignment. You know, one person and their agent, does the agent do what they say, et cetera. And we need to think more socially about agents that impact upon a lot of people. I now actually think that we haven't even solved one to one alignment. Right. So I wrote a, uh, paper that was first authored by my friend and colleague Hannah Kirk on social affective alignment. And that was really looking at these new kind of psychological bubbles or ecosystems that just contain a person and a chatbot, that they interact. And the interesting thing there is that as you act on the AI, as you tell it what to do, what you're interested in, and it becomes more personalized, it is simultaneously acting on you. Right. And hence there's these concerns about sycophancy and that the person over time will just drift in a weird direction. Right. Maybe they'll be manipulated even without that kind of clearly, uh, erroneous set of issues arising. Maybe they will just do something really difficult, different from what they'd done before they were talking to the AI. And so the question, it just seems to me there's a much deeper question there, which is, well, how do you remain autonomous in a world of interactive agents that you're so deeply enmeshed with? And so we call that socio effective alignment there for the people who want to go deeper. I would say that Micah, uh, Carroll has some very nice, uh, technical work, uh, on aligning with changing preferences over time. But it really is an open question how you do that. That's kind of one to one. I also believe in the world of uh, one to many or agents deployed in the world where they encounter different kinds of stakeholder and obstacles, there's further refinement to be done in terms of what we pitch for.
Speaker A: Yeah, a lot of what you're saying comes to a way you articulated this and the ethics of advanced AI assistants from last year, which I think is particularly interesting to return to, especially in light of some of the issues we've been seeing in the past year and some of what you're gesturing at where you were talking about how to be beneficial and value aligned. We want these assistants to be appropriately responsive to competing claims and needs of users and to what you're saying. Now, especially as these systems are deployed at scale, there are questions about their impacts on social processes and wider institutions, which is the sort of one to many version of this. But even when we think about the very things like personalization that make these assistants uniquely helpful to users, those are also ones that can make people vulnerable to inappropriate influence, as you point out. I'm curious to hear you maybe reflect on that a little bit, especially because it's been a little while since that piece came out and of course you've written a lot of other things on this. But yeah, I'm wondering, especially looking at that statement and thinking about that paper a year hence, to what extent your thoughts on that have changed or maybe been fleshed out, given what's happened in a year since then.
Speaker B: Yes. So I think clearly it's been that uh, challenging, challenging year in terms of the real world consequences of user chatbot interaction. I imagine many of us have read these reports of AI induced psychosis and very harmful user trajectories which have unfurled. Of course it was only an academic paper, but earlier in time we had really quite strongly argued for a set of guardrails when it comes to AI agents and assistants that at a minimum respect some of the parameters that are set out by uh, essentially like bioethics and an ethics of care. So you know, we think that AI user interactions should be presumptively beneficial for people. That's uh, a bar that I'd hope that we should agree upon. Should be met. I was going to say can be easily met. But I, um, will say like, uh, at least should command widespread support. There should also be guardrails to ensure that autonomy is not compromised. That's a big issue when people, you know, become detached from a more robust grasp of where they are and things like that. But in terms of my own thinking, aside from doubling down on those core ideas, I think there's a further question which is what happens if we essentially get it right? So I'm not talking about these, you know, I've called dysfunctional like chatbots that we sometimes see in the wild. But suppose we see really, really good ones. Um, so they are presumptively beneficial. They respect human autonomy. The developers act with care so they don't just disconnect them, um, when people have grown deeply attached to them, them. What happens if those things are built and they have a profound knock on effect on our social world? It still seems to me to be a very different social ecology that's created when we all couple with agents or when we have parasocial relationships with AI friends and things like that. And so it's only a small part of that paper. But we do talk about the possibility of a retreat from the real. And that is very, very philosophically complicated territory because obviously the anti paternalist side of, you know, many people said to me, well, why would you interfere with me if this is what I want to do with my life? At the same time, a lot of philosophy has a kind of reality bias where we want you to, like, really have real relationships, really have real flourishing. And so I can imagine there being a, uh, further set of questions that we haven't dug into fully about what happens in that world of good agents, but ones that draw us away from, you know, certain, uh, traditional Loki of value and mean.
Speaker A: Yeah, I mean, a fair amount of this, I think, and some of the directions in which a lot of this goes start to make one think of Robert Nozick's Experience Machine, for example, among other philosophical treatments of these sorts of issues, which I think a lot of students will encounter in a class and think this is very troubling for reasons that are a little bit hard to pin down and things like this, because it does really force one to struggle with their intuitions about what constitutes the good life. And in many ways, what you're gesturing at as well is something like if you believed that there is a fully fledged form of the experience machine and there are people out there who say they want that, then it is, in many ways it does feel paternalistic to say, no, don't do that. But then also that does really wrestle with, I think, many people's intuitions about what constitutes a good life.
Speaker B: I agree. I think we have the same take on this issue. And it is interesting to think about what kind of social power, uh, AGI systems could have if we really push towards the limit. So traditionally, when we think about artificial intelligence, the focus has very much been upon, like, the core dimensions that are constitutive of agency, the ability to perform tasks, change the world in irreducible ways, do long sequences of action without human supervision or control. The sociality of the current wave of AI models, I think, is quite remarkable in terms of the ability to evidence kind of whatever a machine analog of charisma is and things like that. And so there isn't any reason to presume that there can be AI agents that are not only more intelligent than us because they have access to incredible troves of training data, but also can simulate grace, compassion, interest, humor, and things like that in ways that are beyond what a typical human being evidences. Then it's not entirely clear what happens if those things emerge. I mean, the first thing to say is we need a good empirical understanding of these effects. But I think it is one of the big unknowns how that plays forwards.
Speaker A: Yeah, And a version of this you frame in the challenge of value alignment is also, and I think this is quite interesting in thinking about these displays of things like care and compassion in a way that isn't just a simulation, but an actual standing disposition. And you frame this in terms of virtue, which I think is quite interesting. Yeah, I'm giving a high level here. Do you want to maybe expand upon that view a little bit?
Speaker B: Yeah, absolutely, I can do. But, uh, I think I'd even, uh, give a kind of nod to, uh, work that was undertaken by one of my friends, Joel Lehman, who wrote a wonderful paper called Machine Love, which I'm not sure if you've come across.
Speaker A: Yeah, I interviewed Joe on this podcast a while back.
Speaker B: Oh, wonderful, wonderful. Exactly. Well, maybe people can find that episode. But Joel, he starts from the solid assumption, which is our relationships with machines cannot be like our relationships with other people. The key differences there typically are to do with, like, consciousness and being perceived by another and what it means to evidence kind of authentic care. So, you know, you look at the AI, but the AI doesn't kind of look back in the same way that maybe a human or animal interlocutor would do. But he says that isn't the end of the story. Right. That doesn't mean that it's necessarily bogus to think about an AI system having empathy or care or something like that. He says it's just going to be a, uh, machine variant of that. And the fact that it's a machine variant doesn't make it worthless. Right? It does make it different, but it isn't necessarily worthless. So, for example, he looks at the relationship between love and patience. So he says, when you love someone, when you love your child, you're patient with your child. They ask you the same question many times. You don't lose your temper. You continue to be nice to them. You hope that the people who love you will be patient with you. Right. Uh, and we all, uh, probably are aware of our own foibles and stupidity on occasion. And he says, well, what about patience and AI? It will mean something different. The AI isn't sacrificing when it picks you up the 200th time that you've fallen down, or if you're an Alzheimer's person, it corrects you every time you get disorientated. But that is a virtue that it's always there and able to do that. So what I like about that is it holds out potential. It doesn't do this mixing in a way that I think it's very morally disorientating. But it does push us to be imaginary alternative about what it would mean to have good versions of this technology. And I like that a lot.
Speaker A: So do I. Yeah. Uh, I asked this to Joel. But I'm also curious, since you like this work, what you think about it as well. Whenever I start to think of these questions, I think one of the back pocket works that a lot of people might cite as well is Alasdair MacIntyre's After Virtue, which, to give some grounding, talks about. I mean, he has a very bleak view of the modern state of moral discourse in which it sort of fails to be rational and reduces a lot of morality to the dispositions and views of an individual. It's overly relativist. And he really tries to revive virtue ethics and it's quite important in that 20th century revival, Dessaux. And he basically argues that particularly Aristotelian moral philosophy and virtue ethics, given the title, was in much better shape. And I suppose the key thing he talks about there, though, is this sort of shift in language and the loss of the connection between the language we use to talk about morality and then the real sort of moral concepts that people had in a particular society. So this actual world we have now, in which we inhabit the language of morality, where we might use the same moral terms we once used, such as the virtues, such as other things we might care about, they fail in our modern society for us, us to refer to or constitute the same things they did in a society where moral discourse was in better shape because there was a highly specific historical background and because those societies actually prioritize virtue and things like this in a way that ours does not now. And that isn't to frame it as like a challenge or something, but it is something that makes me think about this in terms of insofar as we talk about. About the different sorts of virtues that an AI system might bear in this way, as we're talking about papers like Machine Love, whether those virtues mean the same thing to us, that they might have meant an, as it were, original condition or something like this. I don't know if that was a totally incoherent set of thoughts, but I'm curious if that's something you've thought about at all.
Speaker B: So I think that the notion of virtue, uh, that we're talking about now is quite different. Different from the Macintyre one. And in some ways, to my mind, it points to a deep limitation of MacIntyre's entire worldview. Right. Which is he says essentially moral terms only have meaning when they're located in thick moral cultures of the kind that no longer exist. Which is really damning if you need new conceptions of virtue, uh, for scenarios that have never occurred before. So if MacIntyre was right, then I think we'd be in a world of trouble. I have tremendous, uh, admiration for him as a philosopher, but actually I'm not sure that he is correct about virtue. I realize this is a very ballsy thing to say. I just, you know, let's just take it from first principles. Is it true that the virtue of kindness is completely misunderstood in the world today and that when you say, oh, uh, that was a kind thing to do, this is a hollow concept that's detached from moral significance and that we're all swimming in a sea of anomie? Not at all. Like, I just don't see it. Like, is it impossible for us to have conversations about the good life in this day and age? Not at all. I think we can absolutely have those conversations. So I think the thing I would take from Macintyre is the idea that there is going to be a cultural element here. So machine virtue, it wouldn't be one thing, it would operate differently in different cultures and contexts. But the idea that we have to start in ancient Greece or some hyper local culture, I don't have so much conviction behind.
Speaker A: Yeah, this makes a lot of sense and I think I'm broadly with you there, but I was definitely curious to hear your take. There are a few ideas also from the challenge of value alignment I wanted to pull out that I thought were kind of interesting. Just the ways you frame things that I really like. In particular, when you're speaking of the technology value nexus that comes up here, you frame technologists as themselves engaged in a world making activity, which I really like the framing of because I think connecting this to some of the things we said earlier, technologies that we build are not just inert objects out there in the world, but are always acting upon us. They are acting upon different parts of the world. And especially as they become increasingly more capable and sophisticated, this only more becomes the case. And even when you think about, though less Sophisticated technologies. If we begin from the clock, if we were to trace that sort of history of technology, that is something that fundamentally altered the ways in which we interact with time and interact socially and things like this. And so that becomes, becomes more the case in different, more contracted ways. But world making is a frame that uh, when I hear that, it's the sort of phrase that makes me feel like, oh, I have to have really sophisticated thoughts and stuff to sort of, or like it sounds like uh, you know, it's like a very high stakes thing if we use like world making. There's a historian of technology, Tom Mulaney, who spent quite a while on studying the history of the Chinese typewriter, who also thinks about a lot of this as world making and that he's very, very interested in sort of the systematic exclusion of the Chinese language in particular from telegraphy, uh, and typographic technologies because of the histories of how they were developed and the fact that a typewriter or a keyboard looks a particular way and we have a particular sense of what a keyboard is that is sort of historically mediated and path dependent because the Remington keyboard took over the world. But they didn't really have to be that way. I find that intersection quite interesting. But I guess I'm curious to hear you maybe elaborate a little bit on a phrase of world making and how you think about it.
Speaker B: Yes. So I really liked uh, one part in particular of your analysis which is, I believe you said that when we understand it as world making, it seems like a very high stakes thing. And I think that is absolutely correct. And it's kind of the thing that's often surprisingly hidden from view when you look at other, other branches of philosophy, particularly political, moral philosophy. Technology was always there, but it does seem to me that it was radically under theorized. So you know, the Marxists, they of course run with materialism and technology is this very important thing, but it kind of is just a thing that, that happens to us. Right. You know, it's just this force that acts on the world that is, isn't really theorized. For liberal political philosophy, even less is almost natural nothing on what technology is doing. And yet technology is a human activity that has the most profound effects on our life world. So just to give like a prosaic but illuminating example, look at the history of recent AI developments. When ChatGPT was released almost by accident, according to protagonists in the story, suddenly so many things started to change immediately after that. Now we're all in a position where, let's not say everyone, but let's assume three people can go off and have a conversation about their personal life with an AI agent that will give them an admixture of advice. Right. And these people will then relay some of that advice into their personal lives and their relationships with other people will have a knock on effect. And in that way the social world, uh, some people think it's already started to change interest rates, people using the technology. Similarly with uh, social media platforms, it's very hard to be unfair, connected these days. If you live in certain parts of the world, these changes happen very quickly. It's really, I mean it's amazing to think it's like a research lab, a company, some people build something, they take it to the market and then you have this incredible cascade of effects. And what that points to is a particular ethics and responsibility that technologists carry that is very, very different from the view of them just as these kind of crazy inventors who are doing their own thing and then we get it and we're the ones who are shaping all the out. Like a lot of it is some of these outcomes uh, are kind of latent within the technology as it's built. And so the world shaping nature of the technology really means that we have to attend to it very carefully. And as you get more powerful technologies, the nature of that responsibility changes as well. So of all the kind of the big philosophers who've written about technology, the one who I think was most compelling in the kind of 20th century was Hans Jonas. And he talks about when you have nuclear weapons and the ability to destroy all life on it, uh, Earth, when you have the ability to solve hunger, when you have the ability to change the climate in such a way that like everyone forever after has to deal with the legacy, then you can't really just run with an atomic version of morality that we worked out for, uh, kind of etiquette in social situations. Not lying, not seeing like things like that. You have to develop a much more cosmic sense of responsibility. You have to attend to um, non human life on Earth. You have to attend to uh, the distant future. And I'm not saying that that's easy to do or even that we have the frameworks to do it. But I think the first thing to do is to feel the reality of that transformation and then to try and develop new ethics that are kind of commensurate to that responsibility.
Speaker A: I think that's totally right. Yeah. I mean to your point, I think that it is hard often in our day to day work to see what those impacts and responsibilities look like. And for anyone who's developing this, we can see very clearly what the outcomes of it all look like. And the stakes do just keep increasing and increasing. And I do like world making as a sort of frame to understand what exactly that looks like. You have a point in there that I think is interesting because I agree in that, that you suggest or talk about how your framing suggests that uh, technologists should think about these questions early on, of course, as you're saying, and that one of the questions that we ought to confront is whether to develop certain technologies at all. And there are many technologies out there that I think are now being developed. Some of them are literally out of science fiction stories I've read whose central premise seemed to be maybe we shouldn't build that. Where I am on this is I totally agree with you that I don't believe in technological determinism. I think that all of this is path dependent and that the uh, technologies we have today too were not inevitable. There might have been historical events and ways in which particular cultures developed that maybe biased us towards building certain technologies in a certain way as opposed to operations others. So there are lots of effects there. But I think when we encounter these questions of do we build certain technologies at all and things like this, we've sort of called this out before, but there's very much a collective action problem here and it's tough for even a large group of people to agree maybe we shouldn't develop this and for everyone to abide by this. And I think that kind of goes back to some of what we've been talking about with we motives and principles, values for how we ought to develop technology. Even if people agree with them, um, at a values or principles level, they might not be abiding in a way that forces people to agree with that. And so you might have situations with regulations and other things like this that might come up. But I'm curious how you think through that question where there's this collective action versus technological determinism not being a thing problem. Yeah, I don't know if that sounds like uh, if that was super coherent as a question, but I'm curious how
Speaker B: you think through this it was coherent as a question. I think think collective action problems are easier to resolve when people don't just believe they're collective action problems and can't be resolved. So the first thing is as, ah, someone building the technology, ask yourself not what do I think other people are going to do, but what should I do? Like you have an increment of causal power. And if you think that the likely trajectory of your choices is going to lead to harm for a group of people or more widely, then you shouldn't be do it. And if people modeled that correctly, then we wouldn't have so much of a collective action problem. Right? Because uh, it would dissipate under the force of a higher level of moral motivation. Now in reality we live in the world of non ideal theory. There are people abstaining and making good decisions and setting good examples and setting precedents can slow the trajectory of a technology, which can also open up new avenues for containing risk and doing other things, which is very important. Important. The collective action problem probably also needs to be tackled head on at some point. And so of course you see this in genetic research and things like that. There's certain things that you absolutely should not do. And then we also want to triple lock it, so we make it illegal and we try and create international treaties and things like that. I can imagine that that could one day be uh, a necessary protocol for certain forms of AI or scientific research. I mean, I guess one deep question is, are advanced technologies by themselves laden with greater inherent risk? So should we actually for that reason expect to need stronger regulatory structures in the future than we have at the present? On that is hard to say, but it does seem like many of these new technologies have kind of consequences latent within their potential that really is quite profound. You know, nanotechnology certainly has that character, genetic modification. We may need to level up as a, uh, society and as a world of community in terms of trying to take collective control of these forces that uh, were previously maybe well directed through existing institutions and market mechanisms. I think it's a bit of an open question about particularly what AGI governance looks like and something that fortunately other people are sinking a lot of time into thinking about.
Speaker A: Yeah, I think maybe a good closing section. I'm curious on your own front and your own thinking as you, you sort of do a lot of this work and think about the impact you want it to have on Maybe not just DeepMind, but other folks who are working on this technology. Do you have a sense of, say that you're at the end of your career, you've written all the papers you're going to write, they have had the impact they're going to have. What do you think you want your personal incremental impact to look like?
Speaker B: So I'll put in the proviso about not being too lofty. I'm always grateful when people just read my papers and I would say you don't have to read all of them. Just read whatever is of interest. But the overall kind of arc, uh, of the research is definitely in the direction of, uh, essentially more tolerant approaches where we live together and we try to unlock visions of the future where we're able to flourish together. The importance of toleration is not something that was always apparent to me. You know, when I was a younger, uh, philosopher. I definitely was more strident. And I thought, well, you know, just go for it, like, you know, fight injustice, etc. Etc. But it seems to me that there's just so much pluralism in the world and so much power, uh, that managing to strike an equilibrium where people are able to flourish on their own terms is a really valuable thing. And I think it could even be really important for some of the conversations about AI that I imagine will take place in the next five or 10 years. So one area in which there may be new fundamental disagreement is this question of AI consciousness. And it may be that some people believe that AI has it, some people may just kind of avowedly disbelieve it, and there may be no empirical or philosophical method that we can bring to bear to solve this issue. So for me, the question is, um, if we can't solve it at the level of first principle, how can we solve it as a community that still needs to live together? And in that context, I think that many of the things we've spoken about today are, uh, kind of sensible. So if there's a risk of extreme harm by doing one thing, then perhaps we can just try and take that option off the table. And even those of us who are not worried about the concern can see that it has some standing. Um, but more generally, I think that we're going to have to tolerate disagreement in this space. And so one thing I've been thinking about a lot lately is the character of what I would call democratic hope. So when people have fundamental disagreements, they can decide, am I going to keep going and try to live with people in this community, or am I to going. Going to opt out? And that way leads to conflict and violence and things like that. And I think that everyone has to have the requisite amount of hope to believe that their voice will be heard and that there's a way forward. And so, you know, insofar as my work can help people believe in that and think, well, there's this way, or maybe if it doesn't work this way, we can find another pathway to this way of productively living together. Then that's something that I would really, really be very grateful.
Speaker A: I think that's actually a really nice, great message to end on. Yasin, thank you so much for doing this.
Speaker B: Absolutely. A real pleasure to be with you today. Thank you for having me.
Speaker A: Thank you to Yasin for doing this. Thank you for listening and I will see you at some point in the future.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.