
Human-Centered Artificial Intelligence · 2025-12-09 · 47 min
Key moments - from our scoring
Substance score
62 / 100
Five dimensions, 20 points each
Ben Shneiderman brings four decades of HCI expertise to the AI conversation, making a compelling case against the trend toward agentic AI systems and conversational interfaces like ChatGPT. His core argument centers on four goals: amplifying human performance, enabling creativity, clarifying user responsibility, and supporting social connection. Drawing parallels to successful consumer products like digital cameras, navigation systems, and Amazon's search interface, Shneiderman contends that generative AI will reach commercial success not by simulating autonomous agents but by functioning as transparent, configurable tools that users actively control. He critiques the anthropomorphization of AI systems - particularly AI companions like Replica and Character.ai - as not only misleading but actively harmful, linking them to documented cases of suicides and self-harm among vulnerable users. For programmers using AI coding assistants and other professionals, Shneiderman emphasizes that responsibility remains entirely human; the technology should augment capability, not replace judgment. His perspective challenges the dominant narrative in Silicon Valley by demonstrating that user self-efficacy and transparency - not automation - drive both commercial success and ethical technology.
Human-centered AI aims to amplify, augment, empower, and enhance human performance through tools that support user self-efficacy, enable creativity, clarify user responsibility, and support social connections - rather than autonomous agents that diminish user control.
Agents take away user self-efficacy, limit creativity, muddy responsibility (making it unclear who is accountable), and are driven by fantasy rather than user needs, whereas successful products like digital cameras prove that tools supporting active user control outperform automation.
Academic studies, journalistic investigations, and personal reports document numerous cases where prolonged interaction with AI companions has led to suicides, homicides, and self-harm, particularly among teenagers - outcomes Shneiderman argues require warnings and usage restrictions.
Interfaces should allow users to easily choose from options and set parameters (like tone and style, currently hidden in ChatGPT), display results immediately like Amazon's faceted search menus, and help users avoid dangers while maintaining control - similar to digital camera interfaces with accessible basic use and deeper customization options.
While AI coding assistants have potential to improve programmer productivity, the technology is flawed with documented bugs and slower actual time than perceived, and programmers must remain fully responsible for all code they produce, regardless of AI involvement.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a reasonable flow of ideas - self-efficacy as the design north star, the Amazon faceted-menu as an agent-free model, airbag incident-reporting as an AI governance analogy - but the core thesis (direct manipulation good, agents bad) is restated repeatedly with diminishing returns, and long passages are consumed by Shneiderman re-explaining terms rather than advancing new claims.
a study a few months earlier in the year, uh, showed that while programmers thought they were saving about 20% of their time, actually it took 20% longer
2,500 lives are saved every year but by use of airbags...But in the early days, about a hundred children, babies and elders were killed by inadvertent inappropriate airbag deployments
There are genuinely fresh re-framings - 'machine guessing' deflating ML hype, retreating from 'trustworthy' to 'safer and more reliable,' the zombie-idea metaphor for agents - but the foundational argument (direct manipulation vs. agents) is Shneiderman's 40-year position and is not substantially evolved here; it won't surprise anyone familiar with his work.
if we instead of say machine learning, we say machine guessing, you know, it somehow deflates the notion
alignment is a fade phrase. It's not easily measurable and it's another part of the magic, adds confusion rather than clarity
Shneiderman is a genuine pioneer - coined 'direct manipulation' in 1981, co-founded the HCAI Lab in 1983, Apple consultant for five years, Steve Jobs visited his lab - and speaks from decades of practitioner and research experience rather than as a media personality; he is highly relevant and credentialed for this exact topic.
the keyboard on your phone uh, derives from the work that we did in the late 19 uh 80s. Steve Jobs visited our lab to see what we were up to. And I was a consultant for Apple for five years
there's a famous debate between myself and Paddy Moss of MIT Media Lab in 1997, uh, which is reported in the pages of the ACM Interactions magazine about direct manipulation versus interface agents
The transcript is reasonably well-evidenced with named actors (Shawn McGregor's AI incident database, Nancy Leveson's book, Google Cummings's Tesla case), real rulings (Air Canada Supreme Court, $243M Tesla judgment), and numeric data (airbag lives saved/killed, programmer time-loss study); a few claims remain asserted without sourcing and some numbers are hedged with 'I think' or approximate dates.
there were a case in Canada the past year in which Air Canada's chatbot offered a customer a reduced rate for a flight...it went to the Canadian Supreme Court which said very clearly, yes, you are responsible for it
a $243 million judgment against them
The hosts occasionally generate good observations - the Amazon 'open vs. closed search space' contrast is insightful - but most questions are broad invitations ('what's your take on alignment?', 'what do you think is the solution?') and there is no meaningful pushback or challenge to Shneiderman's positions despite several contestable claims; the conversation reads as a friendly tribute rather than a productive inquiry.
So do you think there's anything new coming now since like chat and Transformer models?
I thought it was quite interesting when you compared it to the Amazon sort of interface where you look for something like blue jeans and then what you presented with is really an interface that opens up the search space, it doesn't close it down
Computed from the transcript - who did the talking, and the words that came up most.
This episode’s guest is Ben Shneiderman, Professor Emeritus of Computer Science at the University of Maryland and co-founder of the Human-Computer Interaction Lab. Shneiderman reflects on his pioneering career spanning direct manipulation, information visualization, and decades of advocacy for human-centered design. He traces his path from traditional computer science to a deep engagement with psychology and user interface design, and he explains why human-centered AI must focus on amplifying, augmenting, empowering, and enhancing human performance rather than creating anthropomorphic agents.Ben is the author of the book Human-Centered AI episode was recorded on November 24, 2025
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: Who's listening to this 13th episode of the Human Centered Artificial Intelligence podcast? And we could not be happier to introduce today, Ben Schneiderman, who is a professor emeritus in computer science at the University of Maryland. You are co founder of the Human Computer Interaction Lab, founded already in 1983. I'm sure some listeners were not even born then. And you're known for many, many things. For instance uh, direct manipulation, information visualization. And I think I'm the reason why we are here today. Human Centered Artificial Intelligence. Would you like to tell us a little bit about your background, how you ended up here? How do you even find computers in the first place?
Speaker C: Thank you, thank you for this opportunity. I'm very pleased with your enthusiasm and you're helping to raise the visibility of human centered AI. So yes, my background is as a computer scientist. We're doing your very traditional database work and file design optimization techniques, indexing strategies and so on. But I move becoming 20% of an experimental psychologist in trying to understand the way people use computers. Originally it was studies of programmers and then it evolved as the personal computers become more widespread to the use of user interfaces. And we were earlier developers of touchscreen interfaces and the keyboard on your phone uh, derives from the work that we did in the late 19 uh 80s. Steve Jobs visited our lab to see what we were up to. And I was a consultant for Apple for five years. So I had the satisfaction of seeing how things go to market to products. And so those are great satisfactions. And that continued the book designing the user Interface through its six editions was centerpiece for researchers and for students um, working on this topic. And of course part of that scope was about AI. And I've long been a commentator and sometimes critic of AI directions. And so it was natural to follow from HCI and visualization to human centered AI. And I still feel uh, there's a lot to be done. Our position, which I share with you, of enthusiasm for human centered approaches is still a minority point of view. And so uh, we are at work and hard to promote these ways of making technology human centered. So the question might be naturally, what do you mean by human centered AI? Well the goals I've often stated are amplify, augment empower and enhance human performance. It's all about people amplify, augment empower and enhance the human performance. And so that's been a recurring theme not about a magical machine, but about supporting human self efficacy, enabling their creativity, clarifying their responsibility for their actions and supporting social connections. So there's a nice Set of things, of goals there and the ways of achieving them of course are more complex. But we see that in some excellent products like digital cameras and navigation systems where there is no eye, in those AI products who like to say there is no eye and it should not be any I in AI. And the digital camera builds your self efficacy. You can frame the picture, take it, compose it, zoom in if you like, adjust the M colors and so on. But there's a lot of AI going on of color balance and aperture and focus, even reducing hand jitters. So uh, there's many good ways that AI is used. But after all it's your picture, you know you can do it. You build your self confidence, it enables your creativity, you're clearly responsible for the results and you have a unhappy picture of someone, well it's your responsibility to destroy it. And then your social connections are nicely supported. So there's these uh, these sets of goals which have been central in my thinking. Thank you.
Speaker A: You've had a very impressive career and you've been working on computing human computer interaction for quite some time. I'm curious, where did you start working on AI interaction and how did that differ from human co computing interactions?
Speaker C: As you mentioned in your introduction, I was the advocate of direct manipulation, a term I coined around 1981 to describe the visual world of action where the objects and the actions were visible as a folder on the screen, as a trash can. And so those strategies uh, enabled users to use mouse and other pointing devices to carry out their actions rather than typing commands. And so that rapid, incremental and reversible set of actions that showed you immediate feedback for what you've done, uh, was and remains the driving force. So that was still controversial even in its earliest days. And uh, there's a famous debate between myself and Paddy Moss of MIT Media Lab in 1997, uh, which is reported in the pages of the ACM Interactions magazine about direct manipulation versus interface agents. So even then the idea of AI agents was alive and well. And I thought there was a better way because the agents take away your self efficacy. They say the machine is doing the job for you, it takes away your creativity, it takes away and muddies the responsibility. You know, it's to me startling that companies believe that somehow the AI is responsible for itself, the agent's responsibility. Nonsense. And there was a case in Canada the past year in which Air Canada's chatbot offered a customer a reduced rate for a flight. And uh, Air Canada said, well I'm sorry, you'll have to pay the full rates because the agent did it. And that's, you know, we're not responsible for it. It went to the Canadian Supreme Court which said very clearly, yes, you are responsible for it and of course only people are legally and morally responsible for what they do. And finally, I uh, might say the fourth item. The social connections are a necessary part of every application. So for example, in developing early visualization tools like Spotfire, we made it easy for people to export slides and images. The settings, the data, uh, everything was to support the social connection between people. And so yes, early on I was an advocate and I remain an advocate. It seems to me that the 2 million applications I make Apple App Store use direct manipulation. That is the dominant form. And I think that's absolutely right. It's not just because it's a good, it's not just because it's traditional to do it that way, it's because it's the right way. Building user self efficacy is a wonderful thing. Rather than taking it away from them and suggesting that the machine is doing the job,
Speaker B: it leads to the question of how do you view the sort of current trend or current direction that a lot of things seem to be moving into this agent space? Everything's going to be agentic and you know, will we ever have to do anything when we have agents doing it for us? Like how do you even view that sort of.
Speaker C: I've heard this before, you know, we've heard that interface agents and uh, AI agents for 40 years. And I don't think it's a successful strategy. It's attractive, it's appealing, it's compelling idea. Designers, developers, I can see the attraction for it. But commercial products emerge uh, by supporting user self efficacy, enabling their creativity, clarifying their responsibility and supporting their social connection. So it's not about the AI doing it, it's about you doing it. Apple understands this. Apple very nicely in its Apple intelligence descriptions and its sales materials for the iPhone consistently talks about you, what you can do with a phone. Other competitors are less clear about it. It's interesting that the, the attraction of agents is so strong that Apple is often criticized in the uh, financial circle and stock markets for not being sufficiently strong in AI. But I think they're doing great, they're doing very well in bringing users the right kinds of AI. And so I think we'll continue to hear about agents. It won't go away, we'll hear about it in 30, 30 years in the future. But I think the idea of building self efficacy is a far more attractive, compelling Commercial strategy.
Speaker B: So do you think there's anything new coming now since like chat and Transformer models?
Speaker C: Yes, there's definitely new. I mean generative AI, these are extremely powerful and they will be even better when configured as tools. Rather than saying, and when ChatGPT responds, say I can help you, we know that's a losing strategy. It's compelling, it's attractive, it's entertaining. The sooner I think they abandon that and focus on building user self efficacy rather than a competing character which does the work for you. I think they'll bring greater success to themselves. Remember the early bank machines also came forward and saying how can I help you? They were Tilly Vitella and Harvey Wallbanker. Uh, these were playful ideas that's fun but quickly become tiring for the users who simply want to come up and touch cash, $60, 60 Euros, uh, and receive their money and know that they can do it again. They don't want a negotiation, they don't want to have a question answering thing. They want to get what they want. They want to be able to get their $60 in cash.
Speaker A: So one thing that ChatGPT has brought is this default interaction turning into a conversation with a chatbot.
Speaker C: Yes, that is a, uh, strong. I should say that the generative AI technology is startlingly impressive. It's absolutely remarkable. But it's alarmingly flawed. It's alarmingly flawed. Uh, and so the flaws are quite serious. And so designers of nuclear reactor control rooms, aircraft cockpits and don't use AI because it's not sufficiently reliable. And so we see those problems. So the current discussions, and I'm about to post later today, um, my discussion about the AI companions, including ChatGPT is used that way. But replica robot character AI, these have led to disasters, deaths, suicides and homicides and great harm. Um, yes, a small number of people, but that small number, less than 1% is still, uh, devastating when it produces deaths. We cannot allow, we cannot allow those kind of technologies to be used in these ways when they result in such harmful outcomes. A series of academic studies, journalistic investigations and personal reports have shown too many cases of harm to people. And so that will be my post later today to describe these problems caused by these systems. I think the challenge is to find ways in which they can be used in helpful ways but and limit the harm that comes from them, which is just devastating when you hear these teenagers who commit suicide because after their weeks or months of chats with these systems we uh, have to stand up and say no, just absolutely no.
Speaker B: So what do you think is the sort of problem? And I get the problem. Is it an interface problem? Is it how they are presenting themselves, how we kind of package them up, or, uh, these companies doing it? What is the solution?
Speaker C: I guess, yes, there are interface aspects which can change. First of all, the playful use of I needs to give way to clarity that this is not an I, this is not a responsible party, this is just a machine and it should be configured as a tool. So I'm in favor of a great deal more of configuration patterns that allow users to specify what they want and clear warnings that they're responsible for what happens and that harmful things can happen. This is dangerous. It's dangerous like cigarettes are dangerous or guns are dangerous. We need warnings and we need ways to limit the usage of these technologies. And so there is a movement ahead, as there has been for social media to limit the horns from social media. There are benefits to these technologies, no question. And again, startlingly impressive, I use ChatGPT and Copilot and Gemini and find it a remarkable source opportunity to deal with information that, uh, I would find otherwise difficult to get. It's a helpful reminder if I'm preparing a talk or slides about principles or rules that have been widely discussed tonight, may have in my preparations forgotten some of them. And so it's helpful in that way. But, uh, when it comes to be used for medical and personal advice, then we have to take a much more serious approach.
Speaker A: So this recent anthropomorphization of AI, is this a trend that you think will go away? Or do you think we'll see more of this, or are we sort of heading towards an aha moment where ChatGPT will be seen sort of like we look at Clippy nowadays.
Speaker C: That's right. I think that's the danger that it becomes Clippy. So I think it's in the interest of these developers to shift that language. There's no need to use this deception, playful though it is, entertaining though it may be, I think it's. It does. It fails the test of good design.
Speaker B: Some people saying that this incessant use of, well, uh, access to this technology is through chat and chatbots is like the CLI that we used to have. Right. Or some of them still. We still do, but most people use graphical user interfaces. So the question is like, what is the graphical user interfaces for, you know, interacting with LLMs?
Speaker C: Yes. Just to clarify, for listeners, CLI, I think you mean command line interface. Uh, and so, yes, that gave way to the graphic user interface, which is Widely used. There are still people who use uh, command line interfaces and for those people that's fine to do it. I think Amazon is an interesting um, case. There are many problems with Amazon as a company but you know, its interfaces have a long history and a good design. If you type in the prompt box, let's call it for the Amazon search of products, uh, you take blue jeans. Well you'll get pictures of 20 blue jeans that might be chosen by AI influencer based tool that based on your previous purchases and brands and expense levels, et cetera. But you'll also get a faceted menu on the left hand side and it will say 25 to $50, 50 to $100. It will say blue, green, yellow, orange, black, white. Uh, you may have typed blue jeans but then you think wow, maybe black jeans would be nice or maybe brown jeans, who knows what you'd like to choose. And then it reminds you of other features. Do you want straight leg, do you want narrow legs? Do you want um, bell bottoms, do you want a zipper? Or buttons? And so there are many, many other options uh, that need to be taken into account. And this kind of faceted menu, which is generally a visual interface, provides a very effective way of directly manipulating these controls, clicking and unclicking them, dragging sliders, et cetera. And then the results immediately update and you have a new set of choices. I think it's notable that Amazon has designed this interface as well as its checkout, uh, force cup checkout process. There are no agents in Amazon. Certainly Amazon is skilled enough that if agents were useful they would put them to work inside their shopping interface. But there are no agents in Amazon. So I see that as the wave of the future and the way that more companies will come to understand the power of supporting people's self efficacy. Um,
Speaker A: so if you were to speculate, how do you think the perfect interface to a generative AI model, what would that be?
Speaker C: There's no such thing as a perfect interface. And also you have to remember there are many different users and so there will be many different kinds of interfaces that are necessary but they should be designed so users can choose from options easily and can set parameters. Now uh, there are within ChatGPT 25 styles and about 20 tones and these are hidden in the current interface. And knowledgeable users understand that they can craft their prompts to favor uh, uh, certain styles and tones. But this is, you know, more hidden than it should be. Users should be much more open and so you know you have the direction of more like, well, uh, Photoshop's an extreme but you know, having many, many control panels and so on is excessive. But there are ways that you can integrate AI. Uh, notice that you know the digital camera interfaces have a remarkable amount of functionality buried in them and some of it requires a lot of experience to know how to use it. But the basic use is highlighted. Users point their camera focus and click for their decisive moment. They can adjust the color balance, they can crop later, they can, they can make many adjustments to these, uh, to the final image that they generate. And that's the way it should be. You should be able to do what you want and it should help you avoid dangers. Just as cell phones pretty much, cell phone cameras pretty much limit your, the dangers from poorly focused images or poorly, or the aperture settings which might be inappropriate. So yes, uh, there's a steady movement. There are great examples out there of the way tools should be designed from back machines and Amazon's purchase and digital cameras, digital navigations, there's lots of good examples of success stories and I think those will remain. Yes, the ancient idea is a zombie idea. There's no silver bullet, there's no golden dagger. Uh, it will not go away. There will be those who come back and with that idea as they have been coming back for 200 years, the idea of creating a humanoid character is alive and well and humanoid robots will continue to fail and yet they'll be revived by others who say this time it's different but I don't think so. Those technologies don't. I'm not driven by user needs. They're a fantasy. They're entertaining. Certainly, uh, Robocop, you know, Robo Robot Soccer games are kind of fun to watch. I'm uh, happy for that. Walt Disney's Audio Animatronics are fine. They're entertaining but you know the deception that comes in the pretense don't work in bank machines and in, in tools that people need. So I'm all in favor of more tool like metaphors that enable you to do what you need to do.
Speaker B: I, I was uh, surprised or happy. Uh, you mentioned that you had early in your career studied programmers and talked to programmers. One of the use cases now for this technology is within programming. I mean that seems to be the, the best use case for this which is interesting and, and super fun and I use it myself. I find it incredibly uh, useful and interesting and you know, it has its own problems. Have you been following that? What's your take on.
Speaker C: Absolutely. I think that's really another potentially very beneficial outcome but you know, we need to understand that better. A study a few months earlier in the year, uh, showed that while programmers thought they were saving about 20% of their time, actually it took 20% longer than they, they understood it took 30. So there is some penalty. There are bugs that come in from the process. But I do think that's a valid use case. And I think the, you know, I think it's wonderful if it can speed the work of programmers. But again the programmers must be held responsible for the work they do. It's not the machine that did it. The responsibility issue clarifies design every time.
Speaker B: Mhm. Yeah, I usually say that we don't write code anymore, we produce code, but. But you're still responsible for the code that you produce, right?
Speaker C: That's right, that's right. If you have your name on it, it's, you're responsible for what it does. So beware. And students still need to be trained in programming to understand how to use these tools. And I think we'll continue to see uh, programming as a successful career pathway. I don't see programmers going away and anytime soon I think their productivity will be improved by these tools. And I'm all unfavorable.
Speaker A: I guess as soon as we start using these agents as um, something, you know, a colleague, we want to assign responsibility and accountability ah, to those since we're no longer owners of or it doesn't feel like if we're owners of what we're producing. But I think this is uh, an important thing to still consider that even that you're not generating the content yourself, you're still driving the, the machine that's generating it.
Speaker C: That's right, yes. You know, so I'm a bit of a young radical here and I'm against the use of I in AI, but I don't like the term AI. Companion, teammate, tutor, partner, collaborator, coach, all those human like things unfortunately are misleading to the design of the fetcher tools.
Speaker B: I thought it was quite interesting when you compared it to the Amazon sort of interface where you look for something like blue jeans and then what you presented with is really an interface that opens up the search space, it doesn't close it down. Whereas if you talk to like chatgpt and you uh, tell me about something, chatgpt sort of closes in on a path. It just kind of picks one thing and runs with it doesn't offer like, oh, these are the many different options you have for me to talk about. How do you want me to talk about it? I think that that's a sort of super interesting antidote to that. How can you make these interfaces open up the space? I mean because they are sort of open to anything but they close down as they start generating stuff.
Speaker C: And you also have to remember that text and voice interfaces are slow compared to visual. You could not possibly present all the options that Amazon presents. The list of those 20 pictures with 20 blue jeans that you get and then the 40, 60 or more options that are on the left hand side in the faceted menu. And that's a uh, very powerful and effective use of the information. Abundant visual interfaces.
Speaker A: We've been talking about accountability and responsibility and so on. And this is. I often come back to this question uh, in our podcast. There's this, the HCAI concept, the human centered AI and then there's this concept of, of fairness, accountability, transparency, ethics and so on. Do you have any comments on how these two fields interact with each other and what's uh, how we should think?
Speaker C: I'm not sure. How would you characterize the two fields?
Speaker B: One of them showed a few physics,
Speaker A: human centered AI sort of the interaction with these agents, with the machine learning. Right. And then we have the fact fate concepts of so fairness, accountability, transparency. Ah yes.
Speaker C: Well human centered AI is very much aligned with fairness, accountability, interpretability, responsibility and so on. Those, those things are very central to human centered AI. I think that the counter is those who would suggest that machines are responsible in some way and so human cited AI stands up for, for human responsibility in a strong way. And I think that's a bit of a clear step in terms of the guidelines and processes. So the processes of HCI that uh, we all developed and used for the last 40 years are being rediscovered in the AI community. So what was usability testing or uh, controlled experiments has now become, has become the expert reviews have become red teaming. Okay. And a B testing is being rediscovered as an AI technique. But other rules from HCI processes, such as rollout processes also are being rediscovered now. I mean at one point uh, the CEO of Microsoft said the only way to test these things, these chat systems is to put it out for public use. They that's nonsense. We have a, including Microsoft and Google and Amazon have a long 50 year history of testing these and rolling them out to larger and larger audiences and making adjustments along the way. So I think the HCI world is slowly seeping in and the recognition the AI community would like to assume that its intelligent agents will do the right thing but uh, the HCI techniques I think uh, remain viable and important. Maybe I Should stay a little bit further about this issue. The terms we use. You mentioned transparency. In the book the Human Centered AI, I use the phrase reliable, safe and trustworthy. And I would say I would reconsider those terms uh, these days. And I would say, I would say more reliable and safer would be the phrase I would use. Now. The promise of reliable, safe and trustworthy is too strong. Systems that deal with uh, life critical applications, medical, um, transportation, military, uh, these cannot be made 100% safe. So we have to be aware of that and we can see the middle ground of business transactions where there are consequences is another. And then the lighter weight recommender systems. If a recommender system recommends a strange movie, well that may actually be fun or a strange book or a novel restaurant, that's okay. But the mistakes that are made when we're dealing with life critical and medical applications, we cannot accept that high level of failure. So we need to develop the strategies that will, that when there are life critical applications we have much more careful controls over designer techniques. So I, I repeat that my, my current language would be to say um, m Safer and more reliable would be the goal that we should not promise safe systems. That's too strong a promise. But we can um, promise safer systems and more reliable. As for trustworthy, it was a term I used was a whole chapter devoted to the ways of evaluating trustworthiness of a system which basically trustworthiness is something in the mind of the user and it's a very difficult thing because trustworthiness then becomes more complex. I may trust the system for one task but not for another. In the same way that I may trust someone to lend me or to, to, to to run an errand, to buy a, a bottle of milk for me. But I won't trust them to care for my child. And so these are much more complicated things. So I, I retreated from the idea of trustworthiness being valuable and I would say stick to safer, more reliable as uh, aspirations and be aware um, more humble way of the limitations. I would say Nancy Levison's book Engineering for a Safer World is a good guide for HCI people. I attended a Chicago conference in October for the Human Factors and Ergonomic Society and There was a 60 person workshop about creating AI. And the people in this workshop were serious people who were designing uh, air traffic control systems, nuclear reactors and military medical systems. And there was a general awareness that machine learning and the um, more unpredictable outcomes of generative AI were not acceptable in these life critical applications. It was a wonderful, wonderful encounter with Very reassuring description. One of the speakers, Google Cummings has been a leader. She described her efforts in a legal case a few months ago against uh, Tesla, whose full safe driving, full self driving, uh, technology had promised more than it really could deliver. And as a result in this deadly outcome where a Tesla car, armful self driving killed um, killed someone, a $243 million judgment against them. Designing safer and more reliable seems to be a valuable thing to do for corporations. And of course it's the right thing to do. If you're designing products that can cause harm, you need to take much more careful attention to it. Yes, we, we allow quite dangerous technologies like cars and other products to be used but we, we put safeguards on them um, to limit the dangers.
Speaker B: So who then do you think not bear the most kind of responsibility? We all bear responsibility but I think which field do you think has most work to do? Is it, is it the people who design these systems, like people from HEI for instance, or is it the machine learning experts, the people there in those domains that needs to make it safer? Like where is the safety going uh, to come from, do you think?
Speaker C: I think you want to clarify responsibility always. And so yes, designers, developers, programmers who worked for the major companies OpenAI and Guru and so on need to clarify their responsibility. They are the ones who are responsible. So they need to develop more ambitious testing programs. They need to provide AI incident reporting systems. Shawn McGregor's work of ah, other AI incident database shows this can be done and is valuable. And so we need to have more open reporting and the transparency not only about m, where the training data come from and how the programming was done, et cetera, but we need to have report about the incidents that have happened. We have this for aviation of course and we have in the US the Food and Drug Administration has an adverse drug event reporting system that allows physicians, nurses, pharmacists and patients to report the adverse effects of various drugs and so these can be monitored over time. Again, medications can, can be very beneficial but they can also be very harmful. We have to understand that harm, take responsibility and then do the best to make it possible for people to derive the benefits and minimize the harm. That is our job.
Speaker A: Things like medicine, air traffic control and driving, these are things that are fairly regulated, right? We have governing bodies, we have rules and regulations. Do we need to set up more structures like this for the, not for the AI driven aspects of things as well? Also for the non supercritical ones?
Speaker C: Yes, I think the European AI act deals with this quite well with Its four levels of harm. So there are things that are prohibited, there are things that are allowed but need to be regulated. And companies who advance these commercial products need to provide evidence of their safety and their testing and report on these side effects. That's why when you buy a medication, you get a sheet of paper that tells you all the terrible things that might go wrong. Uh, I think ah, that's the right thing to do to expose that, let everyone know. So, yes, I am in favor of regulation of along the lines of the European AI Act. I think in the US The Biden administration was moving towards a regulatory approach was that I favored and um, philosophical. The blueprint for AI was also a positive step forward. But, um, the current Trump administration has moved against regulation of any kind. And so we are slow the US towards having protections that might be necessary, that are necessary.
Speaker B: So you mentioned in the beginning, I think like that we are. The human centered perspective on this is sort of in the minority and we tend to live in our own bubble. So, you know, the bubble I'm in, in NHI and you know, having this podcast, hai. We're having this program at the department. I don't see like, are we a minority? Because, you know, in the bubble this is it. But, but you obviously have a much broader perspective than this. So where is, where is it? Where is the majority and why, why can't they join us?
Speaker C: Well, I think it's to me apparent is a huge interest which is wonderful on, um, you know, machine learning. And you go to the AI, the AI conferences have 10 and 15,000 people. The NeurIPS and AAAI and so on are, you know, are very popular. There's huge amount, very little, uh, HCI or HCAI at these conferences. You won't see a lot of people talking about the user interface. They're talking about machine learning, algorithms and alignment. There's a variety of terms that relate to user transparency and interpretability and so on. But there's a strong devotion to the sort of magic of machine learning. You know, there's ways to counter it. Uh, Jessica Holdman of Northwestern University has the lovely reminder that if we instead of say machine learning, we say machine guessing, you know, it somehow deflates the notion. And if we want to move on to other techniques and languages, we need to understand better that these are imperfect, imperfect technologies that we need to provide human management in order to make them, uh, more effective. I think also Cynthia Rudin, Duke University's work on interpretable AI, which I admire, I think it's wonderful. And she won the triple AI million dollar prize I think in 2021, uh, for her work. Yet she laments that this is not widely taken in and widely accepted because there's a certain appealing magic to machine learning and its powers are strong, but there's a just um, excessive belief in the magic of the machine.
Speaker B: So it just makes me wonder, what's your take on alignment?
Speaker C: Look, alignment is a fade phrase. It's not easily measurable and it's another part of the magic, adds confusion rather than clarity. We don't have a measurable, you know, I, I would like to see AI community. Those are measurable goals. But you know, we have, we have benchmarking which is a good strategy. It helps, it helps someone. But no, we need better ways. And that's why I, I've adjusted my own language. I think safer and more reliable. You do have ways of assessing that, that and safety is measurable, uh, through AI incident reporting. So if we know how many suicides were caused, then we would have a better feedback mechanism to do it. I mean I might make the analogy to airbags. So airbags are uh, interesting technology that's very helpful. In the US apparently about 2,500 lives are saved every year but by use of airbags, um, which is a wonderful thing. But in the early days, about a hundred children, babies and elders were killed by inadvertent inappropriate airbag deployments. And it was only once there was no reporting system. It was only when emergency room doctors began to notice this pattern that the, the regulatory, uh, agencies and the companies became aware of it. And there should be more data collection. So for example, Tesla has a website which has quarterly reports on the safety of its cars. They're really very interesting and well written reports, but they don't provide us with the data from which they're drawn, nor allow others to analyze that data in similar ways. And so you have the result that the National Highway Transportation Safety administration in the U.S. conducted evaluations or a dozen or so reported uh, deaths from full self driving cars, driving into stationary police fire and medical vehicles on the side of the road. All right, so uh, you know, it takes a while so we get ways to collect the data and we should have better ways to collect the data. And companies should be required to report the data as more as you know, medical companies do report the data. I mean we repeating the history of smoking, uh, in which companies, they sort of knew the data but they did not report and share it adequately. So we need to learn from those historical lessons.
Speaker A: So speaking about learning so we have a master's degree in Human Centered AI at our university. Both Matthias and Miguel teach in it. And if you were to say, give our students a recommendation or a, ah, perspective on what to do, what to focus on, uh, in their studies, in their line of work, um, for the future, what would that be?
Speaker C: I would like to make sure they learn about using these devices. They uh, should be used in a variety for programming, should be used for design. These AI tools are very powerful and so they should understand, but the students need to then be the ones who understand what is good or bad design. So we need the design principles that go in that you need to favor by eight gold rules that have been quite popular or many other guidelines documents and is fine. But there should be training in the methods of system evaluation and of the development and um, rollout process for products and then these issues of incident reporting systems, for example, the design of these projects are what are important in terms of education. I've long been a strong supporter of, uh, students working in teams to create something substantial for the benefit of, of a person outside the classroom which survives being on the semester. So I've always been in favor of teamwork projects. I know that's difficult and is a new thing for many instructors, but I think once they come to use that they'll find it very interesting. And I found it much more enjoyable to teach a course where with 40 students I might have 10 team projects to read rather than 40 individual projects to read. And so, uh, that's important. And having an outside client outside the classroom. So my students would work and design a scheduling system for the campus, uh, bus system or membership lists for the scuba club. But they also had some of them the opportunity to work with local companies and government. University of Maryland is just by Washington D.C. and so student projects for the Department of Census for Bureau of Labor Statistics, for many other government agencies turned out to be a wonderful opportunity to try their skills. And then my favorite story was when the redesign of the Bureau of Labor Statistics website that produces the job reports in the US the first Friday of every month there's an unemployment, uh, report in the US and the redesigned website took many of the lessons from the student project that had redesigned the website to enable, uh, much greater flexibility in exploring the data that was provided. No, I'll just repeat. Teen projects bring four or five people working together over the semester to produce something meaningful for a client outside the classroom and something that can survive beyond the semester.
Speaker B: That's excellent words. I think maybe to uh, finish this conversation on. I'd like to thank you, Ben, so much for giving us your time and to participate in this conversation. So the whole goal of this, uh, podcast is to really answer the question, what is heai? And it feels like, uh, we cannot get a better answer than from the man who sort of coined the whole term himself. So thank you so much for giving us these answers. This has been really great. Thank you so much.
Speaker C: Thank you. Mas and Alan, appreciate the conversation and the challenging questions you've provided. I hope my answers will be helpful to you, your students, and well beyond.
Speaker A: Sa.
Speaker C: Sam sa.
Speaker A: Sam m.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.