The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Human-Centered Artificial Intelligence
Human-Centered Artificial Intelligence artwork

HCAI 10 - Earth-Centered AI with Ricardo Baeza-Yates

Human-Centered Artificial Intelligence · 2024-09-22 · 46 min

0:00--:--

Ricardo Baeza-Yates, director of research at the Institute for Experiential AI at Northeastern University, challenges the prevailing focus on human-centered AI by proposing an Earth-centered approach that prioritizes environmental impact alongside human welfare. The conversation centers on how generative AI's massive energy consumption affects water tables and electrical grids - ChatGPT uses 50 times more energy per interaction than traditional search engines - while most real-world problems (medical AI, banking, education) don't actually require big data or deep learning. Baeza-Yates emphasizes that impact assessments similar to those required for vaccines and drugs should be mandatory before deploying AI systems at scale. He critiques the field's obsession with accuracy metrics while ignoring energy use, non-human errors (mistakes AI makes that humans wouldn't), and the exponential challenge of evaluating language models against harmful prompts. He also advocates for humans being in charge of systems from design outset, rather than as a patching mechanism like guardrails in language models.

Key takeaways

  • →Most real-world AI problems don't require big data or deep learning - simpler algorithms like linear regression work adequately for 99% of practical applications, yet funding and publishing bias favor complex models.
  • →Impact assessments measuring risk and environmental harm should be mandatory before deploying AI systems at scale, similar to pharmaceutical and vaccine approval processes.
  • →Current evaluation metrics focus only on accuracy while ignoring energy consumption, resource use, and non-human errors - mistakes AI makes that humans would never make, creating unpredictable failure modes.
  • →Guardrails and human-in-the-loop approaches are patches for irresponsible design; instead, humans must be in charge of systems from inception to ensure accountability by design.
  • →The field's energy consumption is growing exponentially while evaluating impact too late; proactive assessment now prevents costly audits, reputational damage, and legal compensation claims later.

Guests

Ricardo Baeza-Yates

Topics in this episode

EU AI ActResponsible AIEarth-centered AIGenerative AI energy consumptionImpact assessmentNon-human errorsEvaluation metricsGuardrails in language modelsACM principles for responsible algorithmic systemsDeep learning vs. simpler algorithms

Questions this episode answers

Why should AI systems be evaluated on environmental impact and energy use, not just accuracy?

Current accuracy metrics ignore what actually happens in the other percentage of cases - including harmful side effects and unpredictable failures - and completely overlook energy consumption. An elevator that works 99% of the time would be too dangerous to use, yet AI systems with similar accuracy are deployed widely without understanding their failure modes or environmental cost.

What's wrong with using human-in-the-loop approaches to oversee AI systems?

Human-in-the-loop is a patch for fundamentally irresponsible design; humans should be in charge from the beginning by design, not retroactively supervising. It's the cheapest solution but the wrong one - like guardrails in language models that catch obvious issues until someone crafts a smarter prompt.

Do most real-world AI problems actually need big data and deep learning?

No - 99% of practical applications, especially in medicine, banking, and education, involve small data and quality problems that simpler techniques handle adequately. The field's bias toward big models and massive compute serves a few companies' interests, not the majority of use cases.

What is Earth-centered AI and why does Ricardo prefer it to human-centered AI?

Earth-centered AI expands the focus beyond humans to include environmental and planetary harm caused by AI systems - water depletion, energy consumption, and ecosystem impact. It recognizes that nature and the planet are affected by AI deployment decisions, not just people.

How should AI impact assessments work before deployment?

Similar to pharmaceutical and vaccine approval, developers should assess risks, impacts of those risks, and prove benefits substantially outweigh harms - closer to 99% good and 1% bad, not 50/50. This requires technical expertise plus domain knowledge and administrative authority to make decisions.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A78%
  • Speaker B14%
  • Speaker C9%

Most-used words

human28example27problem22today21data19energy18computer17models16understand16centered14first14model14impact14learning13interesting12problems12

Episode notes

The tenth episode of the Human-Centered Artificial Intelligence podcast, where we talk about Earth-Centered AI with Ricardo Baeza-Yates.

Full transcript

46 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Foreign.

Speaker B: Welcome, uh, Ricardo Baeza 8 to the podcast on human centered AI. We're thrilled to have you here. You're the director of research at the Institute for Experiential AI at Northeastern University. You've done very many things in computing, in AI and so on, and we'd be happy to hear a brief introduction of yourself.

Speaker A: Thank you, Alan Matthias, for this invitation. So, yeah, I'm a data scientist before data science existed, I guess. Um, now 34 years after my PhD, first I did a lot on algorithms, and then I moved to information achievable, then to data mining and then to mature learning and finally responsible AI. I have been in academia half of the time and the other half in industry, especially in Jaku Labs. And then Also I was 4 years CTO of software company. So it was really not too much of research for some time. And I have been back to academia now four years.

Speaker C: Okay.

Speaker B: I remember I started reading your IR papers years ago when I was a master's student. And I've sort of seen your career over the course of time and I've seen your works on responsible AI and so on, and it would be interesting to see. What do you think of Human centered AI?

Speaker A: I think human centered AI is very important, especially today where people don't think about the consequences of the use of the technology. Uh, but at the same time I think it's too narrow. I think we should be talking about Earth center of AI because we forget the rest of the planet in everything we do and we need to do special things to remember that. So I hope at some point we always remember that nature is part of us and we kind of basically don't think about it. Right.

Speaker B: This is quite an interesting perspective because I think we've seen in one of our earlier episodes we were talking about humanity centered AI, but this is even taking it one level higher.

Speaker A: Yeah, M. And if there are aliens somewhere, maybe we can include the universe. Like why not?

Speaker B: Yeah, I guess the interesting aspect would be how to center it on things that are sort of beyond human, beyond Earth, I guess. Better.

Speaker A: Let me give you an example. And today we are seeing the problem. For example, energy consumption of AI, especially generative AI, is huge. And we are affecting already, like water tables, um, and electric energy and so on. So we should do something anyways, because otherwise we will be in trouble too.

Speaker C: In AGI, there's this thing called more than Human Centered Design. I don't know if you're familiar with that, but it kind of has the same sort of ethics as to not just look at humans, but look at the bigger picture.

Speaker A: Exactly.

Speaker B: Interesting, because when we look at machine learning conferences, AI conferences, research in general, it's all doing all of this super heavy work. But we also know that even the good old fashioned AI algorithms perform reasonably well in comparison. They don't use the same amounts of energy, but still the trends are moving away from that. So how can we work to reverse that? How can we sort of go back to using the simpler methods for things where simpler methods are useful?

Speaker A: Now that's a great point because that's the other thing that I always uh, preach is that most problems in the world don't have big data. And we have this um, bias towards big data. Ah, a lot of computation, big models, for example, these big language models. Only some companies will be able to develop them because of the data. The amount of data you need, the amount of computing resources you need and the amount of money you need to do that. We are focusing on the needs of very few companies and we are forgetting that for example most, let's say medical applications that use AI will never have big data. And there we need these techniques that we already know and they're quite good. Maybe not as good as a deep learning network, but reasonably well for what you are trying to do. We have small data, we have quality problems in data. Not always. You can basically learn without basically have uh, labeled data, which is the case for a language model. So basically you learn from the text that is not labeled. I would say in real practical application this is the truth. 99% of the time I don't know what's wrong with us, but we are focusing on the problem that looks bigger but they are not difficult because sometimes just you put more data and you can solve that. Um, but there have been some important work on the impact on the environment. For example, the famous Stochastic Parrot paper opened Hondurabo on this problem last year. The usage of water. When you use ChatGPT a lot, I don't know why people don't do the same analysis for TikTok or Instagram because I'm sure they're using more energy. People have compute estimate that ChatGPT uses 50 times more energy in every interaction than a search engine and we are wrongly using it as a search engine. So we are using 50 times more energy in the wrong usage. Because predicting knowledge is not the same as knowing knowledge. We will not convince this big company that this is the wrong way to go. But so what I try to do is to basically be responsible and um, then try to at least be aware of all these risks.

Speaker B: Yeah, we know that most problems can be solved by simple models like linear regression and so on, but that's not what startups get funding for. That's not what you publish uh, papers in a ranked conferences for. So this is a uh, difficult thing.

Speaker A: So is linear regression AI? Good point.

Speaker C: I'm sorry, Investors it is, right?

Speaker A: Or if you have statistics, I mean it's from the 19th century, but no, no, but it is interesting because once in a meeting like an AI person said, and we're doing linear regressions all the time. I said well for you, linear regression with AI. And she said uh, yes. So for me it's interesting like that you say like for investors maybe it's AI, but for me it's not like, like you need to learn something to, to, to be AI.

Speaker B: I've seen pitch decks where people call regular expressions AI. So I think it really depends on who you're trying to convince that your solution is the most advanced. Of course statistics won't give you the million dollar venture capital. Right. I think this sort of ties. I was looking at this ACM technology policy document that you co authored which lists a few of these things like for instance, AI systems developers should undertake extensive impact assessments prior to the deployment of AI systems. And I think this is, this really fits into that because you know, impact assessments aren't just the performance of the system, but also impact, impact on environment and so on. But I kind of don't see that happening neither in research nor in industry. Do uh, you have any insights on how we can get there?

Speaker A: Yeah, this is a great question. So we did that and I was uh, one of the two main coursers, we did these new principles for responsible algorithmic systems in October 22nd. So soon we have two years and I have been very hard to push them. There are like three new principles. The first one is my contribution to that committee where I push for this legitimacy and competence principle, which is the one that is tied to impact assessment. So before you do anything, what are the risks? What are the impacts of those risks? That's the intimacy part because you need to show that this benefits society much more than harms society like you do, like we do today with drugs, with vaccines, with many other things that we make sure that don't harm people. And it's not like half and half, like uh, 50% good, 50% bad. No, it's like 99% good or even more and 1% bad. And um, there we can say, okay, Maybe this is worth to try because the benefits really are much more than the helps. And then also you needed to have the right competence, right competence, not only on the technical computing part, but also the expertise of the domain of the application. And very important because we had already some governance maybe in Netherlands, in France, where the people didn't have the administrative competence to do it, but basically they didn't have the permission to do it. Like they couldn't take the decision. If you read the institution or whatever their main mandate was. The second new principle was to basically to reinforce minimizing harm. Especially because of all the cases of discrimination overnight, which is the most common case today on models using in banking and education, hiring and so on. And then we sent the eight principles to Europe. This is a nice story. And he said, they said the eight principles, do you agree? And said yeah, we agree. But when people read minimizing harm they will read human centered AI. They didn't say it that way. But uh, because we are in this podcast of human centered AI and because of the first question, basically they said let's add a principle that reminds ourselves that we need to minimize the environmental impact too. So basically this is the last principle. And it was, sadly it's the last one because was added to something was already done because it would have been nice to put it together with the second principle that was minimizing harm. So for example minimize human and environmental harm. This could be one single principle. And that's why I started with this Earth, Earth centered computer. Because uh, uh, this comes from that discussion two years ago on, on. Um, yes, we have this cognitive bias to forget about the Earth, our, our planet. This is the first part of the answer. But I can go on, on, on impact assessment which, which is I really think is the main problem that is not done. This will be something that I will put in a regulation not for use of AI, for anything. Like today we have environmental impact assessment every time we do a large infrastructure project that's compulsory in all developed countries and in most countries in the world. So if we have a system that will affect humans and environment, why we can't do the same? This will be my first regulation that you need to do. If you can impact humans and environment in a certain scale, let's say in a large scale, but through Internet you can affect 5 billion people, you need to do that. And this is what is not done today. So the question is asking why it's not done today. There are many reasons. One, there's no economical incentive to do It Second, many people don't know how to do it. Third, measure risk, if not trivial, that's the PT of the EU AI act that measures risk in four categories that few of them don't really exist because we are trying to find three categories in something that is measured by a, uh, continuous variable risk, where you put the limit between low and high risk. It's completely arbitrary. Like many other things we have done in society, like dark skin, clear skin, that's also another continuous variable. We invented these categories for other purposes. Power to simplify regulation in this case, uh, and in most cases only complicates our lives. Now, even if we can measure the risk, it's not too easy to measure the impact. It's also really hard to measure that because I think, um, it's not the best metaphor, but I think it's a metaphor that also explains the problem. We have problems with certain technology in the past, like chemical weapons, atomic bombs. And we have learned the hard way to not use them with AI. It's more like uh, like a cluster bomb. Through Internet you can get to hundreds, thousands of people that are very hard to characterize because in some cases maybe almost like random, not like, uh, these are, uh, African American women or whatever. Random. And then we only know the problem until we know the impact. And this is what's happening today. Maybe language models have between quotes learn a lot of bad shit in the sense that they will output wrong stuff. Uh, but we don't know what are those things, but we don't know even how to ask. So it's like a, uh, double black box to me. Not only, uh, it's very hard to understand what the black box is doing, but also it's like that's another bolt inside that, uh, we don't know how to query, to answer, to know what are the dangers there, what is coming out of the Pandora box. We don't know

Speaker B: who else is supposed to understand this, um, than AI researchers.

Speaker A: Yeah. But for example, last week I was in the 50s, um, Latin American research Computer Science Conference. I was invited because I was a former president of all the computer science departments of Latin America. And I gave a keynote there on responsible AI. And also the. Another keynote was by Don Knus. I'm sure you know Don Knus, I think one of the best computer scientists that exist. He's 86 years old. So he did an online Q and A session and of course people asked him about AI. So what do you think about AI? And something he said was like,

Speaker B: was

Speaker A: so interesting how he put it, he says something like, I will try to paraphrase, I never thought about uh, using an algorithm that we don't understand. Think about it like, it's like saying, I never thought about eating something that I don't know if will kill me or not. Wow, this is pure wisdom. And he said other things that were very interesting, that this was uh, one that really for him the idea of having algorithms that are written by uh, other algorithms because at the end models are written by algorithms. So this is like Animeta algorithm have these issues uh, that he never thought about.

Speaker C: Because you are talking about, you know, computer science, you're talking about Donald Proof. And also you co authored the handbook on algorithms and data structures in 91, right? So you kind of know this stuff. And I'm sure in that book you talk about like one way to assess an algorithm is by doing order notation. Look at the complexity, time complexity. So I'm thinking like, is there something next in time complexity that should have been in that book to assess whether it's not obviously probably a global and humanic impact or risk assessment, but is there some sort of measure that you could put that would be similar to order notation?

Speaker A: Yeah, great question. Haven't, haven't talked about that book in many years. It was my first book and uh, it was just because my PhD advisor needs help to finish the second edition. But I helped. I think today I uh, would add, not only at that time we were worried about memory and time and memory, but I think energy is something that we should add. Uh, and this ties with the environment. So what is the usage of energy? But now that you have that. I remember other things that I usually mention in some of my talk is that for example we are even forgetting about time dimension. You see typical papers where you have two or three models and then you have one typically one, the vertical axis is some measure of quality, accuracy, whatever. And then the horizontal axis is some measure of uh, for example data files, how much data can you process? And then you have variance curves. And then okay, the one that is in the top, that is the best one. And um, you see this in thousands of papers. But these graphs don't take into account how much time you took or how much energy you took to reach the position in this graph. And um, probably the ones that are on lower use M M much less resources, much less time than the one on the top. And you are seeing this in generative AI these days. But people seem to not care. And if you think I'm old enough that I, the first computer I use had 64 kilobytes. Can you imagine something working in 64 kilobytes today? You need, I think you need at least 64 gigabytes to have a reasonable cell phone. Or maybe computer will use only like mine has only. Yeah, 32. Some symptoms using 32 gigabytes in my Mac. So we have, we, we lost this, this obsession to do things better, to do things faster, to do this, this. We're using less resources and only cell phones and maybe sensors and maybe other small devices. Uh, need to worry about that because of mainly energy, battery, not necessarily power or things up. I wish Don Knuth had more influence because I'm sure he will remember us. Ah, all these things. I mean like, you need to be worried about doing things responsible in the sense that the solution is not to buy a bigger computer. The solution is to design a better algorithm. And I believe that's the right way because it's part of the ethics of our profession. Do the best with the less possible resources.

Speaker B: That's an interesting uh, perspective because if you look back at the uh, say 15, 20 years of research, uh, in AI, it's all been deep learning and it's not been making algorithms more efficient. Mainly it's been throwing more compute at it, better GPUs, more GPUs, a bigger infrastructure. But it's only just recently that we start seeing this type of reasoning around. Well, the energy use. I think there was a, uh, keynote at Chi this year. I think Kate Crawford was talking about water usage with ChatGPT and I think ECIR, the European Conference on Information Retrieval, in their opening this year they had a slide where they listed estimates of energy use for papers. This is 2024. This is the first time we, we really see this type, or at least I saw this type of thing. So I guess you could maybe claim it's showing the field maturing, maybe, or it's maybe just showing us waking up.

Speaker A: Can I say both? It might be short. No, but think about that. For example, the first paper that talk about bias in computing systems, the landmark paper in the middle of the 90s, the topic was not discussed seriously until 20 years later. And I think the same is happening here. There have been papers that talk about uh, responsible use of energy. And now when we're using like a significant fraction of our electrical energy on this, like when this maybe when it's too late because already we're reading in one single digit and um, estimation for 20, 30 are in two digits really like I Don't know how much, but are uh, we using 5% of the energy on that we produce on procrastination. But I don't know why we always think about this thing when we're already in trouble, not in advance. That's why it's so important to have this impact assessment. Because let's think about it before, not after. So let's think about it now and then we avoid an audit in three years that also will affect our reputation and also will be uh, a lot of money on compensation, for example to people. So but people decide, no, this will not happen, let's take the risk. And three years later they uh, find the problem. And we have been involved in audits, for example for health insurance companies that had the problem already and now they don't want to repeat it. But they are having the audit because they had the problem, not because they were thinking, let's be wise and think about this now if you want to continue in this line. I think the main problem today also in AI and let's say human centered AI is evaluation. I mean we are evaluating everything. Wrong. And I just gave a keynote at Sigma where I started to put everything I had stored. And I'm writing an article about the limitations of data, machine learning and us. Because the main, the main problems come from us. But one of the main problems in machine learning is how we evaluate these models.

Speaker B: Yeah, we've been talking about in IR and Rexis and machine learning about how evaluation is broken for as long as I've been involved in the community, but never have we spoken about things like energy use.

Speaker A: Oh, but I'm talking about accuracy now. I'm not talking about energy. Even if we're not evaluating energy. That's the problem. Yeah, we are using accuracy. That's the main problem. Yeah, yeah, because accuracy is measuring success, um, and harm, um, measures the opposite failure. I don't care. And uh, the example I use is that if you have an elevator that says it works 99% of the time, most uh, say common sense people will not use the elevator because it may fail and you don't know what will happen. But if the elevator doesn't work 1% of the time, when it doesn't work, stops, you will use it because you know you're safe. This is different between knowing the side effects of the drug or not. But today we don't know the side effects. People say, oh, this works 90% of the time, but what about the other 10%? What happens? We don't really know. And um, this is worrisome. When you have, uh, autonomous car killing people, or when you have, uh, models discriminating people, or you have, for example, models m favoring people too. So they have the two sides of the problem. That's why I'm working with some of my PhD students in measuring a harm. Try to find what are the key errors that the system can do. And I call these non human errors. These are errors that AI does that the human being will never do. For example, even if you are, let's say, drunk, you will recognize a woman in a bicycle at night because you are trained to recognize that. Or you will recognize a dog from a cat. But AI systems sometimes make this mistake. They don't recognize people at night, or they confuse a dog with a cat. And those two things, maybe you are in danger. And also because these are errors that you will never make. These are really complicated because people are not expecting them. It's like, I can never see that. Right? And then suddenly you see problems. And I think this class of problems at least are the ones that are more complicated for you because they are outside humanity. Now here is only humanity, not theirs. But they're outside our comfort zone. And suddenly we see some stupid mistakes. Sometimes, let's say funny, but other times are not funny.

Speaker C: So you're talking about the problems that humans wouldn't make and that we don't anticipate machines to make. I'm wondering if. I think you're also talking about that in your center, about a human in the loop. That experiential AI is human in the loop.

Speaker B: Right.

Speaker C: Is that not solving that problem? Like when you, you have a human basically able to react to whatever the machine is trying to do?

Speaker A: Yeah, I, I would like to, to ban human in the loop because this was a way to say, okay, let's have someone overseeing the system. But the reality is that we shouldn't be in the loop. We should be in charge. We should be in control. Uh, and that's the right way to do it. Like humans in charge. So humans think the loop is a patch. It's like, we have the problem. Let's put someone looking after no, humans in charge means that it's irresponsible, design by design. So, uh, you have that condition from the beginning in the system that you are in charge. And then if they're in charge, nothing wrong will happen because humans are in charge. But basically today people are doing things like AI is in charge. And then to make sure that there's no mistake, you would rejuvenate them is not the right solution. Right. I mean, think about it. It's the wrong way to do it, but it's the easiest, cheapest way to do it. And that's why we are doing that. Like guardrails in language models, guardrails catch all the obvious things until someone else put a prompt that is smarter than the guardrail and you get how to do an atomic bomb. And also there you have a huge evaluation problem, because the number of possible prompts, if you don't put a limit to the prompt, even if you put a limit to the prompt, it's exponential. It's huge. How large has to be the test set to basically evaluate that? Ten thousand prompts, Hundred thousand prompts, One million prompts. Doesn't matter how much you put, uh, the metaphor I use there is that you go to, let's say to the Baltic Sea, closer to you. I take a, uh, water sample, you analyze it, and you say, oh, there's no bacteria in the Baltic Sea. We are safe. You just took a one cup from the Baltic Sea. But the Baltic Sea is huge. And, um, we are doing the same today with language models. So in evaluation, we are fooling ourselves so badly.

Speaker C: I guess, theoretically, you think that you need to have a bigger model overseeing the smaller model. And it's a smaller model that you then put into a product that you release. But then at least you have a bigger product model that is overseeing a smaller one. And that's always.

Speaker A: Yeah. The other day someone said, no, we need to have a small model overseeing the large model. I see where this is nonsense. Of course, this will not work because has less parameters than the large model, so we'll never be able to predict anything that will go wrong.

Speaker B: This makes me think, uh, of the papers on reinforcement, uh, learning with AI feedback, where we've removed the human in the loop, but we've replaced it with more or less the same models. So that doesn't really fix evaluation. It doesn't really solve the problem. It tries to look at it from a different perspective, but I'm not sure if it, uh, actually is the right way to go.

Speaker A: Well, this is like, let's see. That's a good paper. Well, in the sense that what they were trying to do, I don't think it's too sensible. But this will be like, uh, I guess the best metaphor will be here. The typical example of searching your keys under the light.

Speaker C: Yeah.

Speaker A: So, okay, we can fix everything we know, but really, we cannot fix everything we don't know. And this has happened because these models will fix themselves on things that they know that they shouldn't do. Although you should know this is dangerous because we are humanizing them. But clearly we go back to what I said before. I mean, we don't know how to query them in a way that we can find out all the wrong things that they could output because that may be very large. And second, it's not trivial to come up with just the right prompt to find out. And I think the best example is last year when someone said, put a prompt that repeated three times the same word. I don't remember which was the word, but it was a short word. And after repeating like three, four times the same word, which doesn't make any sense in any language, the output was a private record of the person. Oh yeah, so was a data privacy breach by just a silly prompt. So if that works, imagine, uh, how many other things will do something similar. It's unbounded and we seem not to care. We said no, we can mitigate it, we can control it. But this is like controlling, I guess, I don't know, uh, a power plant that doesn't really have a dam, but more like have a neck and water is going out everywhere. So yeah, you fix one, but you need to fix many more. I don't want to sound negative. I think if we don't make people aware of all these problems, they keep being too positive.

Speaker B: Yeah.

Speaker A: So I try to, I have a negative bias on purpose because there are too many people with positive bias.

Speaker B: So to try and frame it positively, the awareness is rising. We have things like human centered AI becoming a thing. We have this entire fact movement looking at fairness, accountability, transparency and so on, trying to figure out what AI systems do. But it's also difficult to figure out how to create a catch all system or infrastructure for these things. Somehow we need to test how we can use these systems and how we can develop these systems. But the difficulty lies in making sure that they don't go out of bounds. And that's not a simple problem.

Speaker A: Well, but I think the first part, the first thing is not about testing. The first thing is about how to start the design. So typically today we have some assumptions and we quickly do a proof of concept to see if they work. And if they work, 80% of the time we're happy. But the whole thing was the wrong assumption because we should be happy. It's 100%. But for example, we didn't talk about. And um, there's a lot of pseudoscience here. For example, there's no scientific paper that have proven that data from other people can predict, uh, a new person on average. That may work, may work, but I'm sure that many cases doesn't work and you don't know in which cases will not work. So if you don't start your design talking to for example the users, making sure your assumptions are correct, ensure that in most cases you are forgetting things, you are not taking some important things in account. Even if there is a perception of discrimination, you need to address that. Maybe you just need to uh, change something on the user experience. But today we are not doing that. We are not basically having all the stakeholders from the beginning involving the design of a key system for society. If you had done that, I think we will remove 99% of the problems we have today. So that's common sense. If you include everyone that can be affected from the beginning, the probability of having a problem will be hundred times lower.

Speaker B: And is this a matter of training the people who do these things? Because I mean as a computer scientist training was on algorithms and machine learning and so on. It wasn't really on thinking uh, of how a user will uh, think of my algorithm. Right. But in essence it should have been right. So is it a matter of training or you know, rethinking how we train the people that work with AI and machine learning and so on?

Speaker A: Yeah, I think uh, part of the solution, uh, is training for example on ethics from the beginning. And then you do it the first year of studies and then in every new uh, subject you do one example where ethics is involved and you get training in ethics because ethics is contextual. So you need to see many problems to understand the trade offs between, for example harm, autonomy or justice. But it's more than that is to realize that we are not building computer science systems. We're building systems that are beyond computer science, that are social systems. So for that we need them, which is humanity. It's like you need to have the sociologist, you need to have the ethicists, you need to have the psychologists, maybe you need to have uh, ethnographer in some cases also you need to have people that is very good at design. And uh, yes, we need to also have the computer scientists and engineers to do the technical part. But it's not something that we can solve alone. And I think that's part of our arrogance that uh, we think that we can solve all these things alone, which is not true. These are systems that are much more complex than anything that can be trained. So it doesn't make sense to train a person in all these things that I mentioned. It's too complicated. Also you need a different philosophy in each case. You will need to study 20 years to be, I don't know whether computer scientists, sociologists, psychologists and so on. Some um, people have done it by choice, like having more than one PhD. But the truth is that we need to say, okay, this is not a computer science problem, this is a multi civilized problem. And approach it that way.

Speaker C: It's quite obvious it's not enough to have a, uh, computer scientist if we are talking about technology. Even Don Klute says he doesn't understand, so he doesn't understand. No other computer scientists can understand it. And so we need more.

Speaker A: I think, and I personally will claim that it's not that we don't understand it because they will say no, no, we understand exactly how a neural network works. And that's correct. That's true. I think what we don't understand is the output of them because they have so many parameters that we cannot really even uh, have a max tool to. And you can try to predict that in all cases we do that. It's like a small model trying to understand a big model. And this is like saying okay, yeah, ah, so we can have a small model understanding you for example, a person. But of course the understanding will be much less complex than what we are. For example. And in this analogy maybe we can think that we are the most for people that think we're like machine learning. I don't believe that. But some uh, people think that we are same kind of intelligence, although I don't believe that. But we can think that we are today we are the top model, like the one that is in the top. And then we're building the large language models and then we have other models that are much smaller. So we, in that sense we are the model that have to control everything because we have larger number of farmers. So we are the human in the loop at the end. But we need to be in charge.

Speaker C: I think, I think, I'm thinking because I've been thinking a lot about what you said that Don said and I'm thinking about it as, I mean we obviously know how to uh, sort a list, right? And we understand the output of that sorted list. But it's as if we didn't understand the implication of having a list sorted. We wouldn't understand that it means that we can now find things more easily or that the order has extra significance. But here in this case of for instance LLMs, we know that it's going to predict the next word. And we know that if we run it several times, going to give us some sort of output, but we don't really understand what those answers are going to be, what the actual recursive behavior of that would be.

Speaker A: I think the thing happens in inflammatory retrieval. So I think information retrieval in that sense is the precursor of AI because we build ranking algorithms for billions of documents. We get the ranking list, but you can't really tell why the order is like that, why this document is better than this one. Well, the algorithm defines that, but sorry, I can't tell you. Uh, and some human might say, oh, it's wrong. I think this is better. And maybe the human is right. But going only to the algorithm, we cannot explain because why, for example, BM25 put this. Because it's too complicated. You need to. This is where machine learning is so good. Understanding huge amount of that. No, not understanding, sorry, processing lot of data and coming up with something that makes sense. But we don't understand why.

Speaker B: I think it's an interesting analogy, especially thinking of ir because now there's so much talk about human centricity in AI, but IR is really the most user centric or human centric system there is. Right. But I can't really recall reading anything about human centered information retrieval, to be perfectly honest.

Speaker A: But I think that's because, uh, everything in information retrieval is around relevance and relevance is human. So I think if encoded in the measure that we are trying to model. Relevant. Well, not the measure, but the concept we are trying to measure relevance. And this is human. This is the first thing I teach in my classes of information retrieval. That relevance is also so human that there is no single answer for you. The order will be different than from other people because you know more, because you have a different need, because you analyze things in a different way. So there's so many factors that can change so relevancy in some senses. And it's interesting because IR is, I guess it's the only before AI is the only computer science field that you need to live with this uncertainty and you have to embrace it and accept it to be able to do research. For example, we don't know how to really measure relevance, but we don't care. Let's do it. Today we are in something very similar with machine learning, but the consequences are much worse because there's no harm in having a wrong ranking most of the time. But now if you apply the same idea, uh, to every problem in the world, the dangers are much, much larger. And um, I never thought about this metaphor with ranking. I just bought it from, um, our conversation. So thank you for inspiring me.

Speaker C: I think we're getting close to the end here, and it seems like we are perhaps able to end on a pretty good note because what this podcast really is about is trying to figure out what hai is.

Speaker B: Right?

Speaker C: I'm talking to different people. What is Hai? Um, and perhaps what you've said now in the end is basically hai is just AI that is relevant for people. Is it perhaps that simple? You know, is that it? It's just AI that is relevant to people?

Speaker A: I think this could be way to put it, but I think maybe, uh, relevant can, can be interpreted in so many ways. Yes, I would say that, uh, and something I always say is we need to think about AI as uh, uh, another kind of intelligence, Computational intelligence that help us to be better. So it's a compliment to us. So I would say for me, if we talk about human centered AI and not Earth centered AI, Human centered AI should be the way to empower people. And this is, I think a better word for me, like how we empower people so they can be more, I mean, better. Well, being more productive, if that's the need. And uh, more happy. Surely that's the need. Because if we do this well, AI in the future can do everything that we don't want to do. And I love this, uh, quote from a Polish woman, I don't remember her name, saying, I don't want AI to write for me, and then I have time to do my laundry. No, I want AI to do my laundry so I have more time to write. This is what we want. So imagine a world where AI does everything that we don't want to do. It's not the case today because most people in the world are, uh, not lucky like us, and they're doing things that they don't like to do. And um, they're surviving in some sense. But if we have achieved that, imagine that world. That world is the real renaissance. You can develop the potential, the things you want to do. This is my utopia. We're not going that way today, but I think we can go that way if we change the incentives, the market incentives, the social incentives, and we improve the training of, of people on ethics and well being and maybe philosophies like uh, Ubuntu or other things that worry about more about the group than the individual.

Speaker C: I cannot think of any better way, uh, to end this episode. Those were very wise words. I was a little bit worried we were going to have to rename the podcast Relevant AI. So I'm glad we don't have to do that. So, Ricardo, thank you so much for joining us and taking the time to talk to us. Thank you.

Speaker A: Thank you for inviting me to the podcast. It was very interesting.

Speaker B: Mhm, It. M.

Speaker A: M.

Speaker C: It.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Enterprise Software Deals Now Include a Vendor AI Model Explainability MandateB2B SaaS Talks with Fexingo · on EU AI Act94 / 100
  • How Fortune 500s Use Procurement to Manage Vendor AI Training Data RightsEnterprise Tech with Fexingo · on EU AI Act90 / 100
  • Shadow AI: 7 Out of 10 Workers Use AI Their Company Can’t See | Ravi Soin, CISO SmartsheetCXO Spotlight · on Responsible AI87 / 100
  • AI You Can Trust, Audit and Keep with Russell Moore, Co-Founder & CEO of Amotivv | Episode 494Leaders In Payments · on EU AI Act85 / 100
  • The Trust Gap in AI: Why Agents Need a New Certification Model ft Rajiv Dattani & David Meyer @ AIUCSecurity & GRC Decoded · on EU AI Act76 / 100
  • Hiring Candidates Based on Practical Skills, Not Just ResumesLock it Down Podcast · on EU AI Act73 / 100

More from Human-Centered Artificial Intelligence

All episodes →
  • HCAI 13 - Human-centered AI with Ben Shneiderman82 / 100
  • HCAI 12 - Democratizing AI with Mark Coeckelbergh
  • HCAI 11 - Ubiquitous HCAI with Niklas Elmqvist
  • HCAI 9 - Human-Centered AI-mediated Communication with Mor Naaman
  • HCAI 8 - HCI and HCAI with Mikael Wiberg
Explore the best B2B AI & Data podcasts →
All Human-Centered Artificial Intelligence episodes →