
People of AI · 2025-07-24 · 53 min
Clement Farabet brings a rare combination of deep research credentials, entrepreneurial execution, and infrastructure expertise to this conversation about the current state of AI development. Having worked with convolutional neural networks 20 years ago, built a visual search startup sold to Twitter, scaled GPU clusters at Nvidia for self-driving cars, and now leading research platforms at Google DeepMind - including Gemini, Gemma, and AI Studio - Farabet offers perspective on how the industry has evolved from siloed research teams to integrated model-product development. The episode explores how scaling laws and compute availability have made the Transformer breakthrough feel inevitable in retrospect, why large language models opened new product possibilities that self-driving cars couldn't sustain, and how foundation models like Gemini now function as centralized bottlenecks integrating visual, audio, and textual modalities. Particularly valuable is Farabet's framework for understanding how research and product development have become inseparable - models are no longer transferred from research teams to products, but jointly trained and continuously deployed. This is essential listening for AI builders, infrastructure engineers, and product leaders trying to understand how to structure AI development organizations and which capabilities matter most for the next wave of agentic systems.
Farabet spent 20 years in neural networks starting from PhD work on convolutional networks and real-time image understanding, built multiple startups (including Madbits sold to Twitter), worked at Twitter scaling GPU clusters and at Nvidia on self-driving cars, and currently leads research at Google DeepMind. He describes himself as atypical because he's full-stack across neural net development, custom hardware, and product innovation, rather than narrowly academic or purely applied.
Research and product are no longer separate - they jointly train core models like Gemini that are continuously deployed and called via API by products. Previously research breakthroughs were transferred to product teams who trained their own models; now there is no separation and tight feedback loops are required between what researchers optimize for and what products actually need.
Farabet viewed it as largely inevitable - Transformers made scaling more efficient and probably accelerated the timeline by 2-3 years, but similar results would have emerged anyway given that compute doubles annually and the industry has access to exponential data and hardware from companies like Nvidia and Google.
While self-driving cars represented an important frontier, they required very long development cycles and sustained investment over years. The breakthrough in large language models opened a new realm of agentic system possibilities that could be deployed faster, and he wanted to work with people he admired like Demis Hassabis and Cora Kwiatkowskyj, a former lab mate from NYU.
Gemini acts as a centralized bottleneck integrating data from visual, audio, and textual modalities trained end-to-end to predict across all modalities seamlessly. The better it performs at this core task, the more it serves as a foundation that can be deployed into any application requiring real-world understanding and action.
Computed from the transcript - who did the talking, and the words that came up most.
Join hosts Ashley Oldacre and Christinia Warren as they kick off Season 5 of the People of AI podcast with their first guest, Clement Farabet, VP of Research at Google DeepMind. They discuss the evolution of AI, from early neural networks to the latest advancements in large language models and AI agents. Learn how research and product teams work together to bring developers the best tools to build with and are setting the stage for an agentic future. Resources: Clement Farabet → "Attention is all you need" Paper → Google DeepMind Models → AI Studio → Agentic systems → Watch more People of AI →
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hi everyone. Welcome back to season five of the People of AI podcast. I'm Ashley Oldacre and I'm joined by a new co host, Christina Warren. I'm so excited to be co hosting with you, Christina. She is a veteran developer and podcaster and I'm so excited to have her join the podcast because we have a special theme this season which is focusing on builders, people who are building with AI.
Speaker B: Yeah, I thank you so much, Ashley. I'm super, super excited to be here and I'm super excited to pod with you. And as you said, we will be talking about builders this season and so that includes the people who are building applications, people who are using tools, the people who are also working on the research. And on that note, we have a really great first guest to help get us kicked off, right?
Speaker A: We do, we do. He is a VP at Google DeepMind. His name is Clement Farabee.
Speaker B: Yeah, and Clement is fantastic. He is the, uh, VP of research at, as you said, at Google DeepMind. And we thought he'd be a perfect person to have on at the top of the season because in my mind he's the consummate builder. He comes from an entrepreneurial background. He's also a researcher and engineer. He's such an interesting guy and I think that he really, uh, has his finger on the pulse of what's happening in the industry, what's changing, and he's also playing a big role on the tools that developers can use, uh, to build things themselves. So really, really excited to talk with him.
Speaker A: Absolutely. So welcome back everyone. We're so excited to have you here. Let's jump right in. This podcast is sponsored by Google. Any remarks made by the speakers are their own and are not endorsed by Google. We are so excited to have Clement Farabet join us. Clement is a VP of Research for Google DeepMind where he works on building Google DeepMind's Genai platforms which include the development of Gemini, Gemma and AI Studio. Before joining DeepMind, he was a VP of AI Infrastructure at Nvidia. He has a PhD from Universite Paris and his thesis focused on real time image understanding, introducing multi scale convolutional neural networks and a custom hardware arc for deep learning. He also is an entrepreneur and the startup of Co founded Madbits, sold to Twitter in 2014. Welcome.
Speaker C: Thank you so much. Pleasure to be here.
Speaker B: We're so excited to have you here. And uh, thank you Ashley for uh, pronouncing the French so perfectly.
Speaker C: I would have wished this was very impressive. It never happens, you know.
Speaker A: Well, uh, uh, I can demonstrate my French skills. This is great. Um, but we're very excited to have you here. Um, and you describe yourself on your own website as part engineer, part entrepreneur, part mad scientist. We'll get into the mad part hopefully of building AI, ah, at Google DeepMind. I, um, would love you to tell us us about each of these three parts and how they fit together or not. Maybe.
Speaker C: That sounds really great. Yeah. So I sort of, you know, I think I have an atypical background. Um, I started getting into neural nets and AI around now, 20 years ago. Uh, and I really got into this through pure passion. You know, I had this sort of, I was extremely drawn to the idea that we could build and train and scale these neural nets to solve arbitrarily complex tasks. And I had a very full stack approach to the problem. So I started fiddling with neural nets, training them to achieve tasks that were seemingly extremely hard back then. Um, but I was also obsessed with what does it mean to scale these neural nets, how do we run them? So I started building custom processors for these neural nets. So I was a bit all over the place in terms of interest. Uh, I became very full stack, uh, very deep on the neural net side. So I always thought of myself as a bit of a mad scientist in that I'm not, uh, a strictly academic person. I don't do things by the book. Um, and then I have this separate threat which is I'm drawn to products, uh, to innovation, to how we turn this technology into useful products. Uh, hence the startup, you know, uh, and I started a few of them. The initial ones were failures and then the third one was sort of the
Speaker A: charm, you know, and that was the one that ended up being sold to Twitter.
Speaker C: Yeah, exactly. So, you know, towards the end of my PhD, uh, I would say, uh, the technology was sort of maturing enough. Like we were actually able to build these neural nets sufficiently well such that they could actually solve very impressive tasks. Uh, and the one I was focusing on back then was essentially these neural nets that were trained to look at pixels from video and imagery content and sort of magically describe the content of these videos. And so we started building this small startup that was focused on essentially building a search engine for visual content. And it was quite new back then. Uh, most of the search engines looking at videos and images on the web were mostly looking at metadata, uh, whatever folks would add around the image or around the video when they would sort of upload them. This enabled us to actually look at the content and so it had tremendous appeal for many companies in the Bay Area. Uh, and so we pretty rapidly sold the company to Twitter, uh, and then helped Twitter essentially build an early version of this AI stack, looking at all the media content that was uploaded on the platform.
Speaker A: So already early on, you're merging kind of this, not just the research side, but also putting it into production and testing, not just from a research perspective, but from a product perspective, how things could work. So you're sort of in an original in the space from that perspective.
Speaker C: Yeah. And that's what put me on a kind of odd, uh, track. I remember many of the colleagues I had back then in the lab. Uh, today, they all grew into fairly successful research scient, and they're very deep in the sort of like, you know, introducing new architectures and, uh, the models we train. I was always a bit downstream of that. I was really interested in, you know, what types of applications we're going to unlock with this technology. Uh, what does it mean to essentially, you know, when I showed up at Twitter, what does it mean to bring these neural nets so that they look at all the content that's produced by everybody on Earth, right? Like, everybody's posting tweets, producing content. What does it mean for these neural nets to interpret this content and help us make sense of it? Um, so I became very applied very quickly. Um, and it was a fascinating journey because I ended up working with pretty amazing people. Early on at Twitter, we built one of the first GPU cluster to help scale the training of these models. That's how I ended up meeting with Jensen, who runs Nvidia, uh, and meeting many other people in the industry who are quite applied as well. Um, so, yeah, this is so interesting.
Speaker B: You were, um, when you were doing your PhD, and then when you were working on, um, the startup that you sold to Twitter and the stuff you were working on with Twitter, this predates just, I guess, to set the stage for our listeners. This predates Transformers. Right, so this was before or right before the current wave of what we've kind of been experiencing in AI. Is that correct?
Speaker C: Yeah, absolutely correct. So it predates Transformers. Back then, we were mostly using convolutional neural networks, which were invented by Yann lecun in the 80s, uh, during his PhD thesis. Uh, so he's been at this for a very long time, you know. Um, and so we would train these, you know, these very back. You know, I mean, in retrospect, compared to what we do now, extremely small models, uh, and it would take us a considerable effort to make them work. I remember training these systems on, you know, this sort of distributed CPU cluster, uh, for multiple weeks to get it to barely sort of like, you know, recognize basic things in images. Uh, and today we sort of do the same thing. You know, Transformers were a very impressive sort of like, uh, module in these neural networks that replaced some of the ways we used to do things before. But it's fundamentally the same problem. We would sort of expose those neural nets to a lot of data and teach them what they should actually predict instead of what they predicted before being trained, uh, and keep going through this process. So we do the same thing today that we did then, except with, you know, a million, 10 million more compute, you know.
Speaker B: Right, right. Which I think is so interesting. Right, so, yeah, because you're. You said you're solving the same problems, but you're able to do it at a completely different scale. I am curious just, I mean, you know, with, with like, I guess, you know, 2020 hindsight was, um, I guess Transformers and kind of, uh, you know, this evolution. Do you feel like that was inevitable or did it surprise you as much as it surprised, you know, outsiders, observers like myself, who. It seemed like, you know, one day we, we were doing things a certain way the way that you were, the way you described how things were being done. And then suddenly it was like there was this, this new breakthrough and it. Even though the problems being solved, um, were, were the same and it might have, you know, done in some similar ways, it was being able to happen at a scale that was exponentially different. So I guess that's like my long winded way of asking, like, do you feel like that breakthrough was inevitable, um, or did it surprise you as much as, uh, it surprised people like me?
Speaker C: Yeah, I think so. To me, it was, I think it was inevitable. Um, and it was not super surprising. In fact, I sort of. Like when Transformers came out, the initial paper came out, I sort of barely registered it. Um, and all I saw was that, you know, things were getting a little simpler in terms of scaling these models. But if Transformers hadn't come out, I think we would have ended up doing the same exact thing, maybe delayed by two years, three years. Because the way to think about these innovations is that some of them make these systems more data efficient, more COMPUTE efficient. So you're going to need less data, uh, less compute to achieve the same thing. But because compute keeps scaling every single year by, uh, 2x or so, with the world that Nvidia has been pouring into this or tpus@google even faster sometimes. Uh, I think it's inevitable and things are just sort of on this path to happen. I think what transformers really unlocked is they sort of probably accelerated the schedule by a few years and then suddenly they really made this technology, uh, possible. You know, it was very clear that through scaling these neural nets you could actually understand uh, natural text and start building these systems reliably. Whereas before that it was quite hard.
Speaker A: It almost feels like you have all of these layers that allow you to interact with the computer. And transformers was just one additional layer that allowed the speed of. It took all of the manual work of having to train the neural nets out and you had the transformers doing all of this work for you that allowed you to then focus on other things. Um, so yeah, this is fascinating. And so then you said you went to Nvidia and started working on self driving cars. So a little bit of a different, different direction.
Speaker C: Yeah, exactly. So what happened is that when I was at um, Twitter, uh, we were building these, scaling these neural nets, building GPU clusters. So we got very close to the Nvidia leadership team. Uh, I met Jensen, uh, I think at a cvpr, a computer vision conference in Las Vegas, I think around 2018 or 17. 17 I think. And it was very clear to me that Nvidia was going to be at the center of this entire revolution. Uh, Jensen was building the entire tech stack, uh, all the sort of like accelerating um, compute platform at the hardware level, the software level with cuda, uh, the stack was very impressive. All of us were using it. You know, there was just no other way to train these, these models at scale. Um, so yeah, I got very excited to go work with him. Um, and very rapidly. One of the things we did was to build this sort of scaled platform to train the neural nets that with power self driving cars. The best example of that technology is what Elon ended up doing at Tesla. This was a very exciting journey. The thing that sort of like ended up pulling me out of it is, is, was the pace of it. It turns out it's very hard to actually build self driving cars. It's a very long journey, a very long investment. You've got to keep at it for years and years. Um, and so you know, when these large, large scale, you know, LLM started to work came out, um, I got drawn to you know, going back to, okay now we uh, with that on the web and the barrier is sort of like lowering. So I got excited to get back into that.
Speaker B: Okay, so it's so interesting you say that so it was kind of the lore of LLMs and the problems those could solve. Is that what you said kind of led you to go into DeepMind? Because that was one of the questions that I had for you was that you joined Nvidia at, ah, it seems like a really opportune time. You were working on really interesting um, projects there and you said you seem to I guess feel like their ascent and kind of really strong role in AI. Um, it was obvious to you meeting uh, Jensen, which I think is really cool to hear, but, but I would like to know, I guess more from you, like what did um, draw you, I guess into, into joining DeepMind, uh, and what were some of the um, problems or I guess opportunities, uh, that you know, made this attractive to you?
Speaker C: Yeah, I mean it's a great question. So as I said before, I've long believed in this sort of scaling hypothesis and that the more we're going to scale those neural nets, the more we'd be able to build amazing applications. And so I've always been drawn by, based on the state of the technology and what we can do with these neural nets today, what are the types of products we can enable and so where should we bring the technology? And so back 15 years ago, I think we were at the beginning of this journey. The types of products were fairly limited. We were able to sort of, uh, automatically annotate images on pieces of text. And this was the beginning. So that's what I tried to do. Uh, then there was this next wave of it was pretty clear that we could train those neural nets to do fairly advanced computer vision tasks. And so self driving cars or robots with visual perception seemed like a thing that could be the next frontier. And so I got very excited to go do it. And it turns out that it was partly true, partly, I think it was going to be a very long sort of journey. And I'm very impressed by some of the companies like Waymo, uh, now who have deployed those services at scale, uh, in multiple cities in the US Partly on a sort of AI vision stack, partly on other modalities, but they really got the service to the sort of like launchable state. Uh, and then I feel like in the past two to three years, because we went through these models being scaled enough that they could actually understand natural language, natural modalities, at such a scale that it just opened a new realm of possibilities in the types of things, things we could build. And so I just wanted to be part of that. Um, and the other aspect is, you know, I've known Cora and Demis for a very long time. Cora, uh, used to be my lab mate at uh, nyu. And so uh, you know, there was working uh, with folks that I really appreciate and admire as well. Was one, one big aspect. Um, and so yeah, it's just, I think the next frontier now is the types of things, these agentic uh, systems, the sort of like products we can build with that level of AI technology. These are the things I really want to focus on basically.
Speaker A: Yeah. Well, coming back to this experience that you bring of having done research, having worked in product, and now seeing sort of two big revolutions between the transformers and now the large language models. The research has led to these productizations and we're in a stage now where we could do these, uh, but we couldn't do it before. And I'm curious as to your perspective on sort of merging this research and this product aspect of coming together and what has kind of, as you said, like a couple years ago it was really just research focused and now the products have exploded around AI. So I'm curious to learn more about what has changed in that regard as well.
Speaker C: Yeah, I mean it's a really, first of all it's such a complex problem to, to even think about, right? Because on the one hand we're sort of like driving the technology entirely bottom up. We're scaling the models. We're continuously looking at, um, these sort of like fundamental evaluations of how well the models can predict content, uh, how well can they sort of reason, how well can they carry out tasks in the real world. And this is very much driven by a constant analysis of how the model is training, uh, refining the data set, refining the EVAs to understand what the models are doing, are doing. And separately, as we deploy these models into products, uh, we're really trying to understand what the product experiences need to get better, uh, and then feedback those signals into the core model. And so it's this constant dance of pushing the state of the art on the research side while really understanding what the product needs and the opportunities of products now is really wild because we're looking at very diverse things. Uh, on the one hand you can build this sort of, um, agentic assistant type applications that use the models to sort of like connect to your data, look at your documents, your emails and help you sort of like automate tasks. On the other hand, you're looking at using the same models to power robots. Uh, so like Gemini would look at the visual world and take actions in the real world. These are very different, but at the end of the day, they're all pushing data into the same core model. And so that's really fascinating to go through.
Speaker A: The more you can do with the models, the more you can then translate that into different products.
Speaker C: Yeah, yeah. Essentially we have. You can think of a model like Gemini as essentially this sort of like, huge bottleneck where we're integrating data from the visual world, the audio world, the textual world, and we're training the model to basically be able to predict across all these modalities completely seamlessly. And the better we get at that, the more you have this foundation to enable applications across essentially anything. Right. Like you can have this model that you can throw into a robot, you can throw into a video game. The model can sort of like natively look at the content of pixels in a video game and start taking actions. Um, so it's a bit of a miracle, actually.
Speaker B: I did have a question for you kind of on this sense. Um, obviously there's been a ton of research that has gone into then building these models. And then obviously, ah, as you were saying, products can then come out of this. Right. Sometimes it feels like things have been in silos where there's been the research about. These are some ideas we have and some assumptions and then maybe product get things after the fact. Um, but I guess how closely, um, uh, to those two groups, I guess work together. And have you noticed a change in that at all?
Speaker C: Yeah, I think there's been a pretty fundamental change. Uh, so first of all, there's always been this sort of gradual, uh, process from fundamental research to applied research research to embedded research, like teams that do research within product teams, uh, to engineering and product. It's this sort of like. And it used to be the case, you know, 10 years ago that a research team would come up with a breakthrough to, you know, train models more effectively or. And then that breakthrough would be get. Would get picked up by a product team and the product team would actually implement and train their own models. So these would be fairly separate worlds and you would have teams in the middle doing the tech transfer, if you will. Things have changed quite a bit now in that we are not transferring, uh, knowledge anymore. We're literally training a sort of core model that actually, you know, runs the products. And so there is no separation anymore. It's, it's. You have to build this model, you have to have it be useful for the products. The products are literally calling an API into the model. Um, and so it's a very different world. And I think it's made Things very challenging for a lot of people to wrap their heads around. How do you even operate something like that? We still want to have the foundational, uh, more advanced research. We want to enable that to ultimately keep disrupting what we do in the mainline. But then we want this sort of core Gemini model to be trained continuously and deployed to the products. So it's a very interesting moment in time, I would say.
Speaker A: Yeah, definitely. Previously products or technology used to be sort of at the helm of governments and they were the ones that would use it first. But with the explosion of large language models, the people have now been like developers and users are the ones that are actually starting to use it and develop with it first. Um, and so what's also interesting is this kind of feedback loop that kind of changes. I think the products used to go from sort of big, big corporations, big government, you know, kind of the military, and then down into, you know, like the Internet for example was sort of down into regular, regular people's lives. But now it's changing and it's almost like sort of going the other way. Like with large language models, it's like it's going in another direction. I don't know if like you're seeing this too, Christina. I'm curious to know more about like this feedback loop like is because research also used to be very much implemented and picked up by product, as you said. But now I guess the question is, is research being, picking up what product is doing and incorporating that into their research as well?
Speaker C: Yeah, I would say so. I think one of the reasons that research has moved so deeply into industry now is the scale of things again, like, because we've realized that the more scale we're throwing at the problem, the more compute we're throwing at the problem, the more access to uh, data coming from all sorts of sources, uh, we're throwing at the problem, the better these things work. And there's very few places on earth where you can afford to build something of that scale. And it turns out that even governments are sort of struggling to keep up now. Uh, and so that's one aspect. Um, now I would say, you know, what's coming from products is mostly a grounding in reality. You know, like what do you want these models to actually do? And I think the core capability we're unlocking right now with these models is that we're teaching them to achieve fairly advanced end to end tasks. Uh, for instance, we're going to show a model how to use an email API to fetch emails and process them and Then summarize them and then spit out some kind of result back to the user. And then we're going to tell the model this is good, this is bad, and the model is going to learn to actually do that better. This is very sort of end to end and deeply integrated into the real world through these tools and through these interfaces. So there's really no sort of separation of the model from the world we wanted to sort of operate in. Uh, and so based on that you have to have this super tight interaction between the two. And I think what we're seeing now is that you have startups like Anthropic, OpenAI, um, ourselves on the Gemini side, we're making tremendous progress when we can sort of do these things very jointly, uh, look at what people really want to do with these models and go optimize for that and then feed that back in the model. While separately you're continuously trying to push these models to get generally better at everything.
Speaker B: As an enthusiast and a technologist, developer or whatever who's not obviously, uh, anywhere close to as deep as you are in these things, seeing the results of this and then having the APIs, having the tools to then build out new things has been incredible. Um, and it's been really exciting. I think it's just, it reminds me a lot of when the mobile took uh, over. There's been kind of uh, this is a similar kind of breakthrough I think, in terms of creativity for people who are building things, um, on the product side and obviously for end users in terms of just the experiences that we can have and the things that we can do that did frankly seem like science fiction a decade ago that are now either already here or are on the cusp of getting here.
Speaker C: By the way, it's also why uh, uh, we built AI Studio and the reason I was so excited to sort of push for AI Studio providing an API to developers, because we're learning so much through that, we don't really know what's feasible with the state of the technology. And each time we improve the models, some new capabilities get unlocked. And what's really exciting is to just put the model out there through an API and let people actually build with it. Um, in the past few years, you know, you've had these startups like Cursor and uh, other friends, you know, sort of like enabling these models to do fairly end to end tasks to help people to code. Uh, you know, there's micro breakthroughs happening every other week and the state of things keeps changing. So it's really Exciting to sort of build these models very close to the way they're being used. And I feel like that's really how we learn and move at pace, you know?
Speaker A: Yeah, absolutely. What are some of the things that you have seen that have really impressed you from the developer community? Uh, from either the use of AI Studio or Gemma models or Gemini models. What are some of the things that you're really impressed by?
Speaker C: Yeah, there's so much. So, first of all, the use of these models for code is really inspiring and impressive and sort of confusing at the same time. Because in the limit, it's sort of becoming almost unclear what we're supposed to do with these workflows. I think we're still in a very brittle state. You know, folks get excited about vibe coding. The idea that you're sort of describing some task at a high level, and then the model sort of dazzled the grunt work. The reality is it's not quite there. So I think we need to figure out this intersection of, you know, what the model spits out. How does a human interact with those things to actually produce reliable software? But that entire body of work is very sort of exciting. Uh, in December we released this live API, which essentially enables people to stream content into Gemini. So visual content, you can share your desktop, uh, screen with Gemini. Uh, there's real time audio as well. This, um, is very exciting because you end up seeing people, uh, who build assistants. For instance, you're coding on your screen, Gemini looks at the screen and you can ask it questions in real time. Uh, and Gemini is sort of like helping you modify what you're doing. Uh, so these things are quite inspiring. And we're just at the, I mean, we're deep into it, but we're also just at the beginning because the capabilities keep moving so fast.
Speaker B: You know, I mean, honestly, like the, the last. I guess it's been four years at this point since we've had, um, uh, large language models, um, in kind of a proper sense, assisting with, with coding and development. And it's just got so much better. And, and it is hard honestly for me to like, remember what it was like before we had AI assistance, um, for those things. It's one of those things if I'm on a plane and I'm trying to code something and if the wifi goes out and then I don't have Gemini or whatever, I'm like, oh, how do I do this Again? Not to say that it's becoming a bad crush. It's just becoming, I think, almost an expectation. But I think you're right. There are still these brittle parts where we need to figure out, um, how everything works together. But the opportunities are great. And I think that we're seeing that based on how quickly these things are being adopted. Um, kind of on that note, I guess, kind of to shift things a little bit to talking about, since you mentioned things like cursor and kind of these agentic systems. I guess really in the last, I guess year, uh, 14, 15 months or so, we've seen a year of agents and agentic workflows and toolings. And this is something that you've been interested in, uh, for a while. Uh, I think on your website, like you have, like that you've been interested in and how to build more general AI agents that can learn on their own and figure out how to acquire optimal data sets to solve tasks. That's me quoting you, which I think is well said. Um, so with the growth, um, and development around agents, what do you see happening in the future and how do you think that these types of, uh, systems will continue to be adopted?
Speaker C: Yeah, it's a fascinating topic. Um, I'll take a step back and try to, you know, define agents a little bit.
Speaker B: Thank you.
Speaker C: Uh, there's a very sort of like blurry line between models and agents. In my mind, we, we sort of like start talking about agents when we have this notion of agency, meaning that we have models that are able to, you know, make decisions on their own. So the relationship you're going to have with this model is, you know, I'm going to ask you to do something, uh, like go book me a flight ticket to Paris, uh, and make it under, uh, a thousand dollars, you know, and the system is literally gonna go off. Uh, start browsing the web, open the, the website to, you know, a travel agency. It's gonna go look at the screen, start clicking on buttons, uh, it's gonna pull your credit card from the file you've saved it in, it's gonna put it in. And by the way, I've tried that recently. These things actually work. Uh, it's a bit scary because, uh, when you start actually giving agency to the systems, there's a huge question of who's responsible for the outcome. So I think we're literally in this phase now where real end to end agents are being built. Uh, the technology makes that possible. There's huge questions around how do we actually deploy those things effectively. Um, I'm very interested in the idea of treating these agents as initially my world was interns. You're going to hire agents into the team, they're going to do internship level work. Uh, and so we need to build the scaffolding so that we treat these agents as virtual workers and they're going to produce team, you know, work for the team. Uh, what is the accountability system? If you have interns, if they make a mistake, are you responsible? Is the AI responsible? So I think we're really, it's, it's critical that we start actually building these things for real because instead of the technology is there, we need to understand how we're going to deploy these agents, how we're going to create a reward structure around them. I like to think of coding agents, for instance. If you have developers in your team and they can spun up 100 or 1,000 agents doing additional work, they should get rewarded for their success if they produce 10x more work. But they should also be accountable for their failures. If the agents start producing code that destroys the company.
Speaker A: Right.
Speaker C: So there's lots of like crazy questions around how it's going to actually work. Um, but I think we're very much in the, in the, in the, in the, you know, thick of it now and so, yeah, just super exciting topic, you know.
Speaker B: Yeah. I mean, look, that is so interesting though. But I mean, but even to your example, right, Because I think you're, I think that you're right, like to a certain extent, like who is responsible and accountable because yes, I could spin up a thousand coding agents and have things do these tasks for me. But, but a, to your point, who's responsible if something goes wrong and you know, I make a git commit that uh, destroys everything, um, or that you know, leaks data or does something else. But I think the secondary question I would have with that is how do we think about the fact that, okay, we, because we can do this at scale, right? You mentioned we want to treat these things as interns. Most typical companies have, you know, um, a set number of interns every year, right. And, and um, you know, the teams have. And that's something that you could have a, ah, human kind of manage them and look over their work. But when we're at a point where you could be deploying hundreds of agents or potentially, you know, quote unquote, you know, recruiting thousands or tens of thousands of interns, how do we, I guess maybe like even, even check the work, so to speak, right? Like if all these things are happening autonomously, um, at scale, like what can we do, I guess other than maybe deploying other agents to check the works of the agents. What can we do to make sure that there are those checks and balances and that whoever needs to be accountable is accountable?
Speaker C: Yeah, it's a very open problem. I sort of think of it as, to me, it's the same way we structure teams of humans. You have a manager and you have folks actually doing the work. Uh, and it turns out that in many cases the folks doing the work, like the ICs in the team, uh, can be much better at something than the manager. They can maybe be moving 10x faster because they know the domain so well. The job of the manager is to create the sort of conditions for the workers to actually row in the right direction. Uh, and I think it's not so different. The only challenge is that it's the speed of it, the fact that these agents might be able to carry out tasks maybe 10x100x1000 times faster than we do. Uh, and so you can't really have a human in the loop to actually check the work anymore. Um, but I think you're gonna have to figure out how to put the agents within guardrails so that they produce economically valuable work in one direction, uh, but are not able to sort of just get into everything and have the ability to destroy everything. Uh, so I think we're taking this fairly incrementally now. We're trying to figure out, uh, even questions of what's a good OS for an agent. Maybe it's an OS that's actually not read and write. It's an OS where everything is additive so you always have a trace, you can always go back. There's fairly fundamental questions like that. Uh, and so, yeah, I mean, really deep topic basically.
Speaker A: I think for me the question is, you know, it comes down to trust too. Because, you know, if I like, even with, you know, when I use Gemini, there's a tendency for these models to hallucinate too. So if we're also building agents that also have a tendency to hallucinate, what is your perspective on how to build more trust around these, these models? And it could also be a comparable question to, you know, how do you trust a human? And I think a follow up question to that is, you know, I'm, I'm a pro, I'm a program manager and if I want to have an agent that's going to be developing code for me, I don't know anything about codes. I have no idea. I can't check if the assistant is going to be accurate or not. Um, so what are your thoughts, I
Speaker C: guess around that Yeah, I think there's two questions. One is a question of delegation and trust and another one is a question of alignment. And on both questions I think we have the same problems with humans. By the way, uh, when you hire someone, and I like the analogy of an intern because it's someone who doesn't quite know the work yet, uh, they have enough skills to get going, but there's a notion of, you know, they're going to make mistakes and you have to create a framework where they can go try make these mistakes. It's not consequential. They learn and then they get better. And I think it's how we're approaching these agents as well. Um, the other aspect to make this whole process improve is that we want these agents to have the ability to introspect so you know, produce a, uh, hypothesis or make a, you know, make a try on something but then look at the results and actually self critique and then improve from that. So that's also by the way, how we're training these, these models right now. Like we're, we're having them self critique, uh, produce some results, look at the results, try to pick the best one and, and, and keep iterating like that. The other aspect is that of alignment. You want the ability to have that agent able to understand what you wanted to do and, and, and do it and if it's slightly off, uh, have the ability to give it feedback. So it's a very sort of like in my mind like similar relation we have with humans, you know, uh, except again the notions of like guardrails and ensuring that the sandbox within which these agents operate is sort of bounded and understood. Uh, so that's really what we're trying
Speaker B: to build now kind of on that question yourself, I mean, you mentioned that you tried out having an agent book of light for you and it worked. Even though it was kind of, kind of scary, which I completely understand. I have, I have the same uh, thought process where part of me I'm like, I would love to do that. And there's another part of me that's like, oh, I'm too much of a control freak. Right.
Speaker A: You know? Yeah.
Speaker B: Even though, even though I would like absolutely trust a travel agent in theory, you know, to do the exact same thing. So like, conceptually it's not that different. But um, but I am kind of uh, curious like how much are you, are you using um, agents yourself and how confident are you in the results that you're getting?
Speaker C: Yeah, I mean right now I would say that I'm much more focused on trying to build them than really using them. But we use like fairly sort of limited versions of these things on a day to day basis. Uh, yesterday I was having a meeting with the team and you know, the whole, the goal of the meeting was to improve how we work together. Uh, uh, and so step one was to have everybody in the room write down, you know, four or five bullet points of, you know, what's dysfunctional, what's not working. We immediately pulled in Gemini and asked Gemini to sort of recap the whole thing, help us understand what's common across the different persons, what's unique to each uh, person in the room. And so that's pretty incredible because you have this sort of virtual entity that's looking at ah, what's going on and helping you extract insights. Uh, that's a very sort of limited form of an agent in the sense that the agency we give it is fairly limited. We're asking the system to essentially spit out uh, an answer. And so the consequences of a mistake are not very high. Right. Uh, but in my mind it's a very early form of an agent. Um, and then as I said we're trying to build the scaffolding for these much more end to end systems like go book me a trip. And in this case we're mostly trying to look at what's going on, look uh, at the failure modes, um, more so than actually using them for real.
Speaker B: That's really, really interesting to think about. I do have a question too. I mean when you, because you do kind of work on the research side, um, do you think about using you know, agents to assist in research or to simulate behaviors, um, that could be used in other types of research? You know there was, there was a recent paper that Stanford published I think looking at the role of these systems, um, in social science. M. Have you thought about that at all? Or again, are you still more m focused on building these things?
Speaker C: No, we actually, we actually have a, ah, pretty interesting experiment going on right now which is um, this sort of research, uh, agent that's consuming research papers and producing new ones and continuously sort of like looking at what's been produced. And so I don't want to sort of reveal too many details of that, but these systems are sort of starting to work. And what's interesting conceptually is that on the one hand you can sort of brute force this. Like if the system is producing loads and loads of papers, uh, maybe one out of a thousand will be relevant. And so then the question is how do you improve this whole thing so that it produces more relevant ideas and less waste? Um, and I do think that the challenge we have now is there's a body of work, uh, that's sort of like testable completely digitally. You know, if you're doing research, uh, uh, on algorithms, for instance, you can sort of like produce hypothesis, run the algorithm in a, you know, Python interpreter and then look at the results. And so it's completely testable. There's a body of research that's more like in the physical world, you know, like, uh, coming up with drugs or in which you sort of have to produce the thing in the physical world and then test it out. And so in the limit you could think of, the AI plugged into a wet lab and sort of producing the things. And so I think we're still a bit far from going to that level. But, um, but in the digital world I think we're definitely seeing some interesting results already.
Speaker B: That's fascinating. Do you have a white whale of what you want to be able to accomplish with agents or AI? Um, it's the things that you want to build, um, things that we can't do right now. Is there something in your mind? You're like, this is the thing that I really want to be able, uh, to accomplish.
Speaker C: Yeah. I mean there's a few directions. One is based on what we discussed before, self driving cars. I would love to have an end to end system that we literally throw into a car and it's looking at maybe one or two cameras and can drive and it just does it. To me that's a fairly sort of concrete test and we're nowhere near that actually. Um, there's really no like, right on. The way we build these systems is extremely decomposed. You know, we extract perception signals, we feed them into another AI. Um, but in the limits we should be able to achieve that. And so in our robotics program at Google, we have like early results of that where Gemini is essentially looking at the camera streams and it's actuating the robot and you sort of like show it a few examples of, for instance, picking a banana and peeling the. And it sort of learns it in like three or four examples. So you have these finite examples of success. And if we keep going down that path, to me, like having a robot where you show it one task once and then it can sort of like pick it up end to end and do it, that would be pretty amazing. You know, uh, I like these tasks in the physical world because I feel like they really test Things out very end to end. Um, in the digital world, I think it's a bit, um, simpler. And yet we're not there yet. But like, you should be able to ask your, you know, this virtual assistant to go, like, pull data across all your modalities and help you be on top of things. And, and even that actually is very brittle today. So it's interesting because on the one hand, we're very advanced.
Speaker A: Yeah.
Speaker C: On the other, it's like not, not quite completely down, you know.
Speaker B: Right. Well, no, because what, what it feels like is that. Because we're getting closer. Right. I mean, I think this is how progress always happens is that you get a taste of what is possible and you want to immediately go there. Um, but. But then, you know, the, the reality of, of what is actually, you know, capable Right now you kind of run into that. But, you know, things, things can change, um, quickly, as we've seen in the last few years. But, yeah, no, look, I'm with. I would love to have a personal robot kind of assistant that I could show something to, and then immediately it can do all my cleaning or cooking or whatever tasks I want to throw at it. For me, I mean, that's literally out of the Jetsons. But I would be very into it.
Speaker A: Yeah. I'm still a little bit skeptical. For me, the trust is a big thing. I would want to be able to have a better trust, understanding and make sure that there's more reliability in the results that, that these models are making. But. Yeah, interesting, interesting topic to.
Speaker B: Uh.
Speaker C: What I like about robots is that even though they might seem more scary, it's actually easier to bound them in a way you can have. You have fairly sort of like, you have fairly clear constraints on what they can do.
Speaker A: Right.
Speaker C: And so it's sort of easy to push for fairly advanced intelligence and have them be bounded to. Okay. It turns out they can only do the dishes or pick things up and they can't do much more. Whereas a virtual assistant that's living on the web and you give it access to the web and your credit card. Just think about it. Right? It can go places.
Speaker A: Yeah, exactly.
Speaker C: And it's very hard to contain them.
Speaker A: Yeah.
Speaker B: When we were talking about booking the flight for you, I'm like, oh, I would love this. And then I'm thinking about, okay, but what if the agent feels empowered and there aren't guardrails in place? I'm just being, uh, ridiculous right now. But these are the things that keep me up at night. I'm like, okay, what if There aren't guardrails in place. And it says, oh, I'm going to just. I know that you will love this. And so I'm just going to buy all these things for you because this is what you always buy anyway, and this is what you want. And then the next thing I know, like, I have like, a $20,000 credit card bill, um, for, you know, in a fully one year in Hawaii. Absolutely. And I'm like, did I. Would I enjoy this? Yes. Is this a thing that I was intending, you know, to happen? Maybe not. Right. So, yeah, yeah.
Speaker C: And it's. It sent an email to your boss saying you, you're resigning to take a year off.
Speaker B: Oh, my God. Yeah. O m. Yeah. I'm with you, Ashley. I think this is, um, something we have to, like, feel like we have trust in. And I agree with you, uh, uh, uh, Clement, that I think I do like the physical kind of, uh, options more. Because you're like, okay, well, I can, I can take the battery out. Right? Like, I can.
Speaker A: Yeah, it's tangible.
Speaker B: Where we are right now, like, that definitely seems like, okay, to your point, we can have these more bounded by tasks versus the, you know, kind of unended possibilities.
Speaker C: Yeah. My mental model for this is that the upside is so high, uh, if we really unlock these agents and they can go carry out, um, they can actually do innovation. For instance, they can sort of read papers in one particular domain, come up with a hypothesis, come up with something that's new, test it out, and somehow you bound the system to this sort of like these constraints, and it's producing novel things. It's incredible. It's like now you have this machine to actually do more research than we can do alone. It can innovate, uh, invent drugs. I don't know. It's pretty endless. The downside is to have these systems sort of break havoc and sort of start doing things on their own in the wrong places. But I think we can control that. Uh, these are not, uh, completely abstract entities. These are bits. They're programs we run. And so I think we can decide in which constraints we want to run them and what we want them to focus on. You know, like, it's, uh, a fairly concrete engineering problem in my mind.
Speaker A: Well, and also, I mean, there were concerns, there have always been concerns around this technology ever since it's come out with models coming out with, like, open models. You know, this whole discussion about, do we open models, do we not? Um, I think at least one of the things that helps me sleep at night is working on a team like Google Deepminder, especially at Google, where, you know, these concerns came up and then the AI principles came out and the fact that you and your team are clearly thinking about all of these aspects, you know, sort of the, the long vision of the potential that this sort of agentic area can, can take us into. But at the same time, you know, we have to bring everybody along and we also have to ensure that there are these guardrails as well. And I think that also beautifully fits into this theme of, I think the research also plays into helping to develop these guardrails and the product aspect sort of helps to, to check it. So there's this, this sort of collaboration between research and project that has really evolved. Uh, is uh, really, uh, for me that's really exciting.
Speaker C: Yeah. And by the way, that's also why we've had this hybrid approach. We've been building closed models at the frontier, that's Gemini. But we've also been building Gemma, which are open weights models that we sort of release to the community. And it really, I love this approach because it lets us really understand what the community can build with these models. And at large, if you release open models in the wild, you get a lot more creative return. Right. Because anybody can try out new applications and build new things. So you really understand the state, the limits of what can be built. And so I've really been loving this approach of doing open models on one side and these larger, more capable models behind an API on the other.
Speaker A: Definitely. On that note, I think we should turn to our rapid fire questions. Christina, you want to start with the first one?
Speaker B: Yes. All right, so this is a thing where we asked everybody kind of a similar round of just rapid fire questions. So just uh, uh, come up with the first thing that uh, comes to mind for you. What was the last thing that you automated for yourself?
Speaker C: Yeah, I think it's literally what I said before, which is this, this meeting we had with the team, uh, and it was basically using Gemini to help us really understand who thinks what in the group and what are the common themes. And it's been incredible to do that because you save time, you get insights that are maybe not, you know, popping up with the group alone. Uh, so that was the last one.
Speaker B: Awesome. Uh, what was the last thing you asked Gemini?
Speaker C: Uh, yeah, I think the last thing I asked Gemini was probably, uh, the weather in London. I still tend to ask that all the time. So these models are all connected to search in the back end. So when you go to the Gemini app, and you ask it a question about live news, weather, or anything like that. Uh, it actually fetches, you know, search in the. In the background. And so it's a really simple way to sort of, you know, interface with search in a way. Also, the thing I do all the time is in the Gemini app, you have this live mode, uh, which is really great. And so when I commute, when I drive, I just turn it on and I start asking questions about history, for instance, you know, like, who was the king of France in the early 19th century? And you just have a chat with Gemini while you're driving, and it's pretty great, you know, like, it's very natural. Um, so if you want to learn about things, you know, like, just do that.
Speaker A: It's such a great learning platform. For sure.
Speaker B: I was going to say that. That's a really fun pro tip. That's a really fun pro tip. Yeah. Yeah.
Speaker A: Um, what could you do now that you couldn't do six, I don't know, six months ago? It feels like a long time. What could you do now that you couldn't do three months ago?
Speaker C: I think. I mean, I can tell you I've sort of, like, stopped. I used to code a lot myself, up to, you know, six, seven years ago. Uh, and then I really, you know, now I'm sort of managing teams. I lead projects, so I code a lot less. And the thing that blows my mind now is vibe coding of, uh, specifically, like, web applications. I love to build UIs for things so that, you know, you can sort of rapidly experiment with things. I can do that on my own now in, like, five minutes. You know, I've been building these applications for my kids to have them sort of like, uh, so my daughter loves to draw, and so I built this app for her so she can draw something. And Gemini tells a story. It turns the sort of, like, drawing into a story. It write this out, and then she can read back. And so it's a very cool, like, educate. You can build education tools in, like, five minutes, uh, on AI Studio in the build page. And, you know, this would have taken me a lot more time, so I find that pretty amazing.
Speaker B: You know, that's great. You kind of answered this before, but I'll ask you again, like, what are kind of the next set of problems that you're trying to solve?
Speaker C: Yeah. So for us right now, like, a lot of my focus is to figure out what is the right, uh, platform to sort of turn Gemini into an agent and bring that to essentially the whole of Google. We want all the teams across Google to sort of very rapidly be able to build application layers on top of Gemini and dramatically simplify this problem. Uh, it turns out it's a very complex problem because the way we deploy Gemini, uh, in different products, the way Gemini interfaces with, with the rest of the world through tools is quite complex. So we're trying to really make that super simple. So it's a very sort of like, internally focused problem at Google, uh, but very exciting for us because we're trying to help all these different product teams innovate with Gemini.
Speaker A: Final question. What is your biggest learning in the past year about the future of work? And we've talked a lot about this in the podcast already, but to wrap
Speaker C: it up, yeah, that's a great question. I mean, I think to me, it's really back to this agent discussion. It's the fact that because we're building the frontier of these models, we're building frontier capabilities, we're on the cusp of enabling these very autonomous agents. We have this responsibility to actually build the scaffolding for them to be treated as this sort of virtual interns that I mentioned. And I feel a lot of pressure to get this right. And I think we. This is the biggest difference from now to, like, a year ago. We're really getting into this era and we have to get it right.
Speaker A: Definitely. Any last thoughts, questions, Christina, before we wrap it up?
Speaker B: No, I mean, I think this has been such a fantastic, uh, like, conversation. I've been. We'd love to talk to you. I, you know, again, in the future, too, to kind of see where we are, um, and, and what's been accomplished.
Speaker A: But.
Speaker B: But no, thank you so much for, for taking the time to talk to us. You've had a very fascinating career. You're working on very challenging problems that, um, I'm very excited we have people like you working on. And so thank you just for being here.
Speaker C: Fantastic. Thank you so much.
Speaker A: It's an honor. Thank you. Thank you.
Speaker C: Thanks a lot.
Speaker A: Thanks for joining us. For more information about our guests today, please check out the description.
Speaker B: And for more great conversations like this, hit that subscribe button. Until next time, thanks for listening.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.