
People of AI · 2025-10-16 · 58 min
Key moments - from our scoring
Substance score
51 / 100
Five dimensions, 20 points each
Bibo Xu traces a decade-long career in voice and dialogue systems, starting from MIT Media Lab's human-computer interaction concepts through Google Home's foundational far-field speech recognition breakthroughs to today's LLM-powered conversational AI. She explains how word error rates dropping from 20% to 5% made voice practical for home use, but the real inflection came with large language models enabling true multi-turn dialogue rather than brittle command-filling. This shift manifests in products like Gemini Live - where users can have genuinely open conversations with an AI that understands context and nuance. Xu details three major LLM-enabled evolutions: natural dialogue mechanics including back-channeling and interjections, expressive speech generation with tone and emotion, and real-time speech-to-speech translation (demonstrated at Google I/O in Google Meet). She notes unexpected user behaviors emerging from better systems - like wanting conversational interfaces in Google Maps for navigation - and highlights multimodal translation's particular power in communities where spoken rather than written language dominates.
The shift from command-control to dialogue-based systems using LLMs fundamentally changed user expectations - people stopped trying to break systems with unexpected inputs and instead started engaging in open-ended conversations, leading to unexpected requests like wanting conversational navigation in Google Maps.
Cascaded systems convert speech to text, translate the text, then generate speech in the target language; multimodal systems take audio input and produce audio output directly through an LLM, preserving the speaker's voice identity and emotional tone without transcription delays.
Key improvements include natural back-channeling ('uh-huh,' 'mm-hmm'), smooth interjection capabilities, expressive prosody and emotional tone in speech generation, and the model's ability to understand context across longer conversations rather than expecting specific command structures.
The physical and contextual difference mattered - people felt awkward speaking commands in public or while holding a phone, but naturally used voice when devices were stationary in private spaces like kitchens where voice was the most convenient hands-free interaction.
Astra is built on true conversational dialogue capabilities rather than the command-and-control model, allowing users to have back-and-forth conversations where the AI understands context, interjections, and naturally human speech patterns rather than expecting explicit structured commands.
Our reviewer’s read on each dimension, with quotes from the episode.
Contains genuine technical substance - WER history, the cascade vs. native audio architecture distinction, full-duplex/open-mic research goals, and latency targets - but the useful signal is heavily diluted by host restatements, filler affirmations, and conversational padding throughout.
what we've been calling a cascade system. And now it's multimodal. You clobber the entire thing into one system and it's just audio in and it's audio out
it was something like 20% which is 1 out of 5 words at the time...it got to like 10%, uh, 8% and then for Google Home, at some point that was like, okay, this is like completely usable, 5%
The Astra memory bug anecdote and the user demanding Astra inside Google Maps navigation are genuinely fresh practitioner observations, but the overarching thesis - LLMs enabling human-level fluid dialogue - is widely circulated and no contrarian or first-principles argument is advanced anywhere in the episode.
it remembered from a past conversation that it couldn't see
she's like, this is great. I want this in Google Maps for navigation. And I'm like, what? Like, that doesn't make any sense
Bibo Xu is a genuine deep practitioner - original Google Home PM, now lead PM on Gemini Multimodal and Project Astra at Google DeepMind - who has demonstrably done the work at scale across a decade; the conversation does not fully unlock her technical depth but she speaks credibly from the inside throughout.
I was the original pm, uh, for Google Home, the first version of a lot of these assistant y kind of technologies
I joined Ah, DeepMind as the Gemini audio modeling PM
The WER trajectory from 20% down to 5% and the 500ms latency target are concrete and useful, and named projects (Project Mariner, Android Computer Control, Gemini Live API) ground the discussion, but the vast majority of evidence is anecdotal with no user-research data, adoption metrics, or rigorous outcome numbers.
word error rate...it was something like 20% which is 1 out of 5 words at the time...it got to like 10%, uh, 8% and then for Google Home, at some point that was like, okay, this is like completely usable, 5%
the goal is to get under 500 milliseconds for human dialogue
The hosts ask reasonable scene-setting questions and occasionally identify a useful thread (interruptability, politeness norms), but they habitually restate or amplify the guest's answers rather than probe or challenge them, and the rapid-fire closing section produces no new insight.
No, I mean, that's so cool. And that opens up so many different avenues for so many different people. Right.
That's amazing. You're getting your MBA, you happen to kind of go into this linguistics kind of class and then it ends up kind of changing, is it fair to say kind of change the trajectory
Computed from the transcript - who did the talking, and the words that came up most.
Bibo Xu is a Product Manager at Google DeepMind and leads Gemini's multimodal modeling. This video dives into Google AI's journey from basic voice commands to advanced dialogue systems that comprehend not just what is said, but also tone, emotion, and visual context. Check out this conversation to gain a deeper understanding of the challenges and opportunities in integrating diverse AI capabilities when creating universal assistants.
Transcribed and scored by The B2B Podcast Index.
Speaker A: On the communication front, I think the um, you know, the learning that we're having is basically how to build these systems that really can communicate with you in a super fluid way. Um, and that I think, you know, LLMs are the answer and I think it'll take us to essentially the human level dialogue system and I think that piece of core technology will enable a lot of things. Like I have no doubt in my mind that um, an AI that can communicate with you but has access to all of the world tools and knowledge, um, can make a lot of different kind of products, um, in the out in the world.
Speaker B: This week we have an amazing conversation with Bebo Xu who is a pm, uh, uh, lead of uh, Gemini Multi Modeling and she also is a PM lead on Project Astra. And, and we had a really great far reaching conversation about kind of the history about how these, these voice systems have evolved from going back the last decade from like the early days of Google Home and to where we are now where we can have full on um, back and forth conversations with AI. And I think that evolution has been so fascinating.
Speaker C: I think what's also fascinating is seeing how the interaction with computers essentially is evolving and how voice has, is playing a role in that evolution and opening up accessibility, opening up different ways of learning, different ways of interacting with the AI systems, with the computers. It's fascinating how this has evolved and it's an incredible conversation. We can't wait for all of you to hear it. So let's jump right in. This podcast is sponsored by Google. Any remarks made by the speakers are their own and are uh, not endorsed by Google. Welcome Bebo. So happy to have you. Like we do with all of our guests, we're going to start off by reading your bio. So Bibo Xu is a product manager at Google DeepMind where she is a lead PM of Gemini Multimodal Modeling covering audio and vision. She works with AI researchers to train models with new capabilities such as native audio dialogue, real time video understanding, expressive speech generation and translation. She also is a lead PM for Project Astra, a research prototype exploring new capabilities of a universal assistant. It is an AI agent built on Gemini capabilities with the goal of helping users with a wide range of tasks. Welcome.
Speaker A: Thank you. It's great to be here.
Speaker B: We are so happy to have you here. Okay, so um, Ashley just read your bio, which is uh, incredible. I mean work on a lot of things with surrounding uh, voice. And so that's kind of one of my first questions for you. How did you become involved with voice?
Speaker A: Yeah, It's a good question. So, uh, I've been working on, I think I've been saying I've been working on things you talk to for about 10 years or so. Um, but actually the interest started when I was in business school. So I think like, as many PMs at Google come, uh, from an engineering background and then went to business school. I went to MIT in, uh, Sloan, um, and I actually took a bunch of classes when I was there at the adjacent school, which was the MIT Media Lab. Uh, it's an amazing place. Um, and that's what I first learned about hci, the concept of human computer interaction. And the people at the Media Lab, one thing that they were doing is they're talking about how you can communicate with computers using your entire body. So it's a bit weird, right? Like evolutionary from evolution perspective, where these, like, we stand up, we have hands, we are, uh, super expressive, we have our voice. But then, like, all the computers that we work with are either like this or like this, right? So basically they were interested in that. And that totally just like, blew my mind. I'm like, yeah, like, that's, that makes sense. I would want to do that. Um, and then I ended up in basically what became the original Google Assistant team. At the time, it was voice search and voice actions. And that excited me because it was something different. It was a different kind of modality of inputting information at the time was mostly just about inputting what you meant to, uh, a computer that understood you. Um, and then through that project, basically, um, I ended up working on Google Home. So I was the original pm, uh, for Google Home, the first version of a lot of these assistant y kind of technologies. I also ended up meeting, um, a bunch of people that were called linguists. And I never met linguists before. Uh, you know, it was only at Google that I met them and I was like, oh, you know, trying to understand what they did, which was linguistics. And I went to take a, funny enough, an audio class, audible class on what linguistics work. And I remember in the first part of that book, you know, the, you know, someone basically said, hey, like, I think humans are the only, um, species that we know of that have language. And what does that mean? Uh, a lot of animals can communicate. I mean, even bees communicate with each other based on, you know, the dances of where the pollen is, whatever. But he's like, think about this phrase. He might as well have done that. Just think about the depth of that phrase and what that means. Yeah, it's. There's so much packed in there. And this is just something you will casually say to a friend in that, just a couple of seconds. And then I was. Then there was, that was my like, okay, like language is that HCI mechanism that is truly different for us as human beings and especially spoken language and spoken dialogue, which I wish we do more of. And it can really help create understanding between each other. So that was the thing that I think ever since then I didn't want to work on anything else kind of, uh, so dialogue and conversation and language has been, uh, my passion for, for the last 10 plus years.
Speaker B: That's amazing. You're getting your MBA, you happen to kind of go into this linguistics kind of class and then it ends up kind of changing, is it fair to say kind of change the trajectory of what you worked on as an engineer?
Speaker A: That's exactly right. It's a combination of basically the MIT M Media Lab, which I'm very thankful for, and actually the people that I met at Google, uh, which opened my eyes to the fact that you can build things that allow us human beings to more naturally communicate with computers. Um, and that has completely shaped my career.
Speaker C: And so in the evolution of voice, um, I think the brief history is that it started kind of in Bell Labs, first of all with being able to identify, uh, pronouncing numbers. And then I think in 2012 there was a big breakthrough with Google Research. Where were you in that? Were you like behind the scenes, like right as it was gearing up to be, not to date yourself, but you know, where were you within that circuit? That was one of the biggest breakthroughs I think, in voice research, which has then led up to the expansion of how voice has evolved in our technologies today.
Speaker A: Yeah, I think this was kind of right before I got really involved so I started getting involved, uh, with uh, at the time what Google Home was, uh, before it became a product. Um, I ended up working a lot with folks in the speech team who were the pioneers of a lot of this stuff, uh, for far field speech recognition. So instead of holding your phone, Far field speech recognition means that instead of holding your phone and you can speak into the mic and you're very close to the microphone, uh, and it understands you. Essentially the thing got solved with essentially a lot of these home Alexa, uh, type of products is that the device can be sitting much further away from you and it can still hear you. And that was combined with echo cancellation, which was the ability for the device to unhear itself so it can hear you. Uh, so that combination that Started to really become a thing that I was at least aware of in uh, 2014, 2015. Um, and that was really the origins of Google Home as a product. Um, at the time actually uh, we didn't know what this thing was and I was working in voice search, voice actions at the time and I was trying to get, I was the music pm right. So I was trying to get people to say play music into a text box and I just could not succeed. It's just like a different, this was Google before Google Assistant, Gemini, all of that. It was just search and the user mental model just wasn't there. No matter what I did, people didn't want to do it and if they didn't want to do it then you know, the music, uh, APIs don't really care about it, you know and the whole experience just kind of sucked. So I was struggling. And then this far field speech recognition thing, I was like, oh, it's interesting capability. We didn't know what to do with it. But then we basically stuck at the time an Android tablet onto a tray that we stole from the micro kitchens and the, and the lunch, and the lunch cafes. And then we like, we literally hard glued the tablet and a far field microphone array like on it and then we had basically the crux of what Google Home was. And this was running far field speech recognition. And, and then I brought one home and I said I can't get people to you know, say play music into their phone, but as soon as I put it on my counter I was using it all the time. I was saying like play this, play that, stop. You know. And I was trying to set timers and all kinds like okay, like that's a, that's a thing. Like that's a very different product. I don't know what that is. Uh, it ended up becoming Google Home, ended up becoming beginnings of Google Assistant, uh, which you know, this was a big part of. Um, but yeah, that's like 2015, 2016. So it was still early days. And also I remembered speech, um, recognition got good, really good around that time. Uh, especially in English I think. Like international languages have always been like hard as well. But it was getting to like, you know it was originally I think when I started word error rate, which is like how many times it makes a mistake, um, when you're speaking to it. Uh, so lower is better. Uh, it was something like 20% which is 1 out of 5 words at the time, which was not really usable. And then it got to like 10%, uh, 8% and then for Google Home, at some point that was like, okay, this is like completely usable, 5%. You basically can, you know, can, um, can get away with it. So, yeah, those were, those were the early days, was a really fun project. Uh, I learned so much and I really m. I still work with some of the same people from those days.
Speaker C: It was such a big hit. I mean, I remember when it came out, it was like the, the holiday gift that I gave to all of my family and friends, you know, because they had the big Google Homes and then the mini ones and they were, I mean they are very, very good.
Speaker A: Yeah, thank you. Thanks for, uh, your support. But yeah, we, it was a, it was a very fun project and it did a lot of things, uh, very, very well. It took, we always had, I think, essentially all the ingredients to build a product like that. Um, and, you know, it ended up becoming. I still have multiple, um, in my house right now and I use it, I still use it every day.
Speaker B: But I think this is so interesting, you know, kind of thinking about how like the different modalities of how these systems work and how comfortable people can be with them. Like, you know, you, uh, were having a hard time getting people to speak into a search box. But then as soon as even your prototype of what kind of became an early version of Google Home was in your kitchen, you were, uh, more comfortable talking to it. Do you think there is. And we've seen some similar things with that, with LLMs, whether it's first through text, and I think now it's getting that way through voice as well. Do you think there is something about, um, I guess maybe kind of a UX element or something about the spaces and how people interact with things which change what they're willing to do and how they're willing to experiment. Um, with voice.
Speaker A: Yeah, I mean, voice is a very different input modality and it has its huge flaws as well. Right. So people don't feel comfortable talking to machines in the public, um, even at home. It's a bit of a challenge sometimes when, you know, when someone is there and they're talking, like it's kind of awkward to be like, set a timer, you know, like, that's still weird. Um, so, but then it is very, like, it turned out, I think it's been proven to be really useful when you're essentially shooting off commands like very quickly and you're in situations where you really don't want to pull out your phone. And this is why at the, in the home, I think, um, there's no, there's no competition because on the phone there's, you have the phone, you can literally just open up an app and do whatever you want with it. Uh, but in the home, in the car, uh, those are situations where you don't really want to take out your phone. Uh, you sometimes it's not on you, uh, in the home, I know mine is always stuck behind a couch or something like that.
Speaker B: Right.
Speaker A: So in those situations you just really want to like a quick thing which is like, hey, set the timer. What's the weather? Can you play the song? Like um, that is just really convenient. The other thing that is really starting to come to fruition in LLM world, which is I think very different, um, is actually dialogue. So.
Speaker B: Mhm.
Speaker A: In the, in the era of the home products, um, that was pretty brittle and that was not a thing that really worked fundamentally because you had a concept of back and forth. Right? So set a timer for how long? Five minutes. Like you had that kind of thing, but you were effectively like form filling. M. Like you had like you, you had a couple of things you have to like form fill and then the moment you say something a little bit off or a little bit random or something that it was just not expected in sort of the natural language processing aspect of that, like the whole thing would, would break. Um, right. And that is the fundamental shift I think what we're seeing now. And when you're starting to use audio system, like essentially LLM based audio systems for dialogue. Um, this was what brought me to gdm by the way. Uh, it was very clear that you're like this, this idea of like a free, free flowing dialogue that and then you have a very smart thing just, just kind of understands everything that you're saying. That's I think the step change that we're seeing right now. Um, and I think that is, that is gonna, that's gonna be a very meaningful uh, change. Because before it was like you know, brittle systems, command and control. That works in a lot of context. Um, but free open dialogue is a thing that you see in the Gemini live the product. Right. Like you can just like talk about your problems, like get some ah, advice about something and it's really a really uniquely human experience, I would say.
Speaker B: Yeah, no, and I think that's so interesting that you bring that up kind of, you know, the change between kind of these brittle kind of, you know, like command control, uh, versus having a real dialogue. Do you think that that changes? Because from my observations, even just watching um, individuals in my life and just, you know, even. Just. Just outsiders, um, speak with things. I noticed when some of the first home products came out, um, that people immediately were trying to unintentionally break them, right? Because it was comfortable. Like, it would play your music, it would set your timer, it would tell you the weather, and now you want to go further and then it breaks and then you get a little bit frustrated. But it was very clear that once people got over that initial hump, being comfortable talking to something, whether it was in their home or their car or whatever, they wanted to continue, and then the systems just weren't ready for that. Now that I guess the systems are more ready for that and are doing things. Have you noticed, um, any changes, I guess, in how, um, people are using these systems and maybe the distances and what they're doing with them that maybe you didn't anticipate? Um, you know, as a researcher and engineer and pm.
Speaker A: Yeah, it's very different, I think, um, the. Because, like the previous, as you said, the previous era of these products, like the Google Assistant, like kind of the OG Google home products, we called it an assistant. But the mental model was not really, uh, like, we didn't live up to that, right? Because it meant that you can actually talk like a human. And I think this is like. I think now the change that we're seeing, so we see people use it in all kinds of ways that are, like, unexpected. But I think there's like an overall mental model shift. So I'll give you one story. Um, so we had a uxr. We did user research with people who were using an Astra prototype. And, uh, one of our, uh, researchers, Tanya, she was like, you know, she got feedback from this user, said, hey. She's like, this is great. I want this in Google Maps for navigation. And I'm like, what? Like, that doesn't make any sense. Why would you want this for navigation? Like, Google Maps is a great navigation system, you know, very. It's perfected over the years. And she was like, well, it doesn't really listen to me, right? Like, it just gives me stuff. It gives me instructions. I don't get to be part of this, like, oh, yeah, thing. And I was like, oh, my gosh.
Speaker B: Interesting.
Speaker A: So she's basically, what she's saying is that, like, I now, like, because I can't. I seem to be able to talk to it about anything, it seems to understand me, so why can't I be part of the navigation experience when, you know, maybe it's saying, turn left and you're like, I don't know about that, that, that exit like that, you know, that one's a bumpy one. And she expects now that the Maps product can basically say oh let me just go and change that route. Right. And you apply that to a lot of products that you use, which is like kind of a one directional thing with the exception of punching buttons effectively. And you can really change the way you think about technology and applications in that way. So I think that's like a, that's a thing that was absolutely not possible before and I think it's possible now. Uh, so it's very exciting.
Speaker C: So I'd love for, for us to go back to. You said things are really changing with LLMs and I'd love for you to dive deeper into what that really means and what the oncoming of large language models onto the tech scene has really done for unlocking voice and taking it to the next level.
Speaker A: So there's, there's a lot of pieces. Let me see if I can like create sort of buckets of that. So yeah, first is I think the big shift that we discussed was essentially um, the capability of fully fluid human level dialogue. And I think we're almost there, it's getting pretty good now. But I think there's a couple of things that will make it really, really human. Um, so for example, the ability for uh, you to interject very, very smoothly. Uh, human conversation is often not turn taking which our systems are still kind of like turn taking. There's also the mhm and the. Yep. And the fact that we, we constantly like use. Exactly we, we use voice to confirm, uh, we call that back channeling all of the huhs, uh, and the yeses and like all of these things. And that's like super normal for all of us. And the ability for us to express much more naturally is also now very, very different. If you hear now some of our audio models both for dialogue and for generation of speech, they have like, they, they sound like um, someone who's expressive. It's not this like robotic like blah blah, blah, blah blah. But it's, it's actually very um, there's
Speaker C: tone and pitch, there's tone like inflection which helps as you said, with the communication.
Speaker A: Yeah. The other day I told uh, you know I was using the live API, I was like telling a joke and kind of like, like laughed a bit and I was like, you know, but that, but that's what happens when you're like giving it the data like be expressive and that's how we communicate as humans. So I think there's just so much in essentially like naturalness, the quality of the dialogue, um, that is truly different. And now yours just sounds like you're talking with someone who's just really, really smart. And that's, that's really the goal. But then to be able to communicate properly, we still need to like really focus on sort of the dialogue mechanics and as well as kind of the, really the quality of the voice and all that. The other aspect is uh, you know, tone and the emotion of speech. So when you say something like how you say it is sometimes as important as what you've said, what you're emphasizing. Um, so when you say, hey, I got my report card today versus Hey, I got my report card today, like that's like a very different, very different thing. So there's a lot of really cool stuff that just on conversation alone, in terms of the dialogue alone, uh, that is really interesting. Um, the other area of voice that we're starting to see, um, you know, a lot of opportunity is actually translation. Um, so translations has been part of these large language models in the tech space for a long time and they're very good at translating. You just give it a doc, it'll translate to a different language. Uh, but actually like translation actually in real life happens more in spoken language. And in audio and traditional systems what you do is you take the user's speech and transcribe it in text first and then you throw text into a translation system and then you get the French version of the thing I just said out in text and they speak it out. So this is what we've been calling a cascade system. And now it's multimodal. You clobber the entire thing into one system and it's just audio in and it's audio out. And it's ah, amazing. The entire experience is just like the model is translating for you live in audio. So uh, one of the cool things we showed at Google I O this year was uh, live speech to speech translation in Google Meet. And that basically uh, employs a LLM. It's uh, a small one, uh, that is live translating what someone is saying into a different language. And it's very difficult to express how amazing that experience is. It doesn't really come across even videos or anything like that. But I think I would highly recommend uh, trying it. But you literally hear, you know, this person is speaking Spanish like, but you hear their voice duck and you hear exactly the same voice come back in English speaking to you and your Mind is just like, you're looking at this person and of course, like, it's a little bit delayed, doesn't match what they're saying. M. But you're just like, oh my God, like, that is crazy. So there's a lot of really stuff happening. Yeah, it's really cool.
Speaker B: No, I mean, that's so cool. And that opens up so many different avenues for so many different people. Right. My thought immediately goes, okay, if we have this here, what if we have this on our phones? What if you have these in other places? You know, if you're traveling around the world and you're trying to communicate with someone, especially people who maybe primarily communicate, uh, using voice rather than language. Which is. Which is true for many, uh, parts of the world.
Speaker A: Exactly. Where.
Speaker B: Where the written language rather is maybe not as expressive or is harder to
Speaker A: translate, uh, as accurately, by the way, on that. So one of the kind of interesting capabilities that I. Is probably one of my favorite that's coming out of audio is, um, uh, in a lot of part of the world, the language is actually generally is. Is natively m. Like a mix. So Hindi and English is like your, you know, there's like, you know, I'm Chinese, uh, we speak Chinglish, which is like combination of Chinese and English. Um, but there's a bunch of other languages that are kind of even like, um, someone's telling me about Korean. There's just like a lot of like English words that's made into the Korean vocabulary. Especially the younger generation just. We call that code mixing. Right. So think about sort of the. Let's say you're having like a dialogue system or a translation system or any generate like anything that has to understand what the user is saying. So one of the coolest things actually, if you experience, uh, the. Our, you know, our models in the ally of API or Gemini Live or Astra, is that you can kind of like speak whatever language at whatever, in whatever mix you would want. And the AI responds in whatever language as well. Like it basically chooses what language to respond. And so it's kind of like a crazy thing. You're some. You're speaking to it first, like, let's say English, and you just change into Spanish and it just changes into Spanish with you in the same voice. So that's a function of native audio models. Because it doesn't have to. It doesn't care. It's just audio. Right. So it will just.
Speaker B: Right.
Speaker A: It will say, okay, well I'm hearing Spanish now, so I'm going to start thinking.
Speaker B: So it'll immediately be able to process that. Right. It's not having to go through that initial like translation step because it already is strained and it's just getting the input as native audio and is able to output however you would want it set to do the. That's so interesting. That's so interesting. So you talked a little bit about what brought you to GDM, Google, DeepMind. Um, but you work on Project Astra and you work on some of the other um, uh, the Gemini multimodal stuff uh, with audio and vision. Tell us a little bit for our listeners who might not be aware, tell us a little bit about what Project Astra is. Because it seems like it's a nice culmination of what you did at the beginning of your career at Google. And then as the technology has evolved this is kind of uh, like the almost you know, like it seems like more of kind of the perfected version of maybe what the, what some of the initial ideas were.
Speaker A: Yeah, it's a. Astra is very fun project. I kind of fell into it. So uh, I joined Ah, DeepMind as the Gemini audio modeling PM. Uh, and Ah, you know we were always chasing this concept of a human fluid, human level, fluid dialog agent. And we were like, I was like okay, I want to go build that. So uh, I joined DeepMind to do that. It was working uh, with a lot of really, really smart people. Um, and I am in London. Um, so I ended up uh, meeting a team that was here in London. Uh, the lead of that team, um, the lead researcher is Greg Wayne M. Who is just amazing. So they're here in London and it's a small team that's trying to build this thing called a situated assistant. And what that was, that's basically what Astra, uh, what became Project Astra. And the reason that I met him was that we were in a meeting, we're talking about dialogue and he's like okay, I need dialogue in this thing that I'm trying to build, which is basically dialogue with the camera on. So instead of just talking about anything, you have the camera on and you're talking about the thing that is in front of you. Uh, this has been uh, an area of research uh, that DeepMind has been working on for some time. And in this new Gemini era it was very clear that we should be able to do something pretty amazing because image understanding, video understanding and dialogue were all kind of like possible in these model. They're coming together in this model. And then I started working with this team in London. Um, uh, and I'm still actually here in the, in the team room at the moment. Um, so, so what Project Astro originally was is the concept of having an assistant that sees the world with you and able to help you out based on what it sees. That's why it's called situated, because like, it's idea that you and the AI are like together and you're situated together, uh, and seeing the world, uh, maybe it helps you cook something, maybe it helps you identify. You know, talk about this, like painting that's in front of me. Uh, and you know, one of the most interesting applications that we also, um, uh, announced at Google I O this last year was essentially helping a blind user navigate the world. So it was, I thought it was a really interesting, uh, concept. Um, and it was a very interesting team and started getting, essentially working with the team as that was the primary sort of thing where you, we have a Gemini model that is able to understand the world in front of you and you can talk about it with anything. And that experience is amazing. And that has really. That was essentially what we announced at Google I O in 2000, like one year ago, 2024.
Speaker C: Right, right. So you have, as you started to say with the evolution of AI, what we're starting to see is, as you said, these, these different modalities that are coming together and working together to enhance the experience. Um, what are some of the, um, challenges? Like? We're looking at this specifically from a builder's perspective, from a developer perspective, because this season is about what our community is building. And what are some of the challenges that you have seen with integrating both of these multimodalities together to create this virtual assistant.
Speaker A: Yeah, so many every day.
Speaker C: I can imagine.
Speaker A: So there is the audio visual part of things and then there is like actually a bunch of other capabilities. So Astra is, you know, the goal of Project Astra is to be the sort of what Greg calls a concept card for the universal assistant. So we're playing with a bunch of capabilities, throwing it together and see if it works together. Um, so along with that is things like memory, um, for example, which is actually really difficult to build, debug and understand what's happening. Uh, one of the craziest bugs I have ever seen was that we had this bug at one point where um, Astra basically says, I can't see, although the camera's on, I'm not able to see. Okay, that's weird. We're like, what's going on with the model? Immediately we jump to the model, we're trying to troubleshoot like what's happening. And okay, so the wow moment number one was that what happened was um, essentially the part of the technology that was responsible for uh, accepting and understanding the visual element was down, that tech stack was down. It was actually right that it can't see. It was telling you a debug information in dialogue. Then it took us a while to like oh my God. Actually it was right that I couldn't see. And what was crazy was that for someone on the team and it was not consistent, um, they still, after everything was back up they were still reporting this issue that it says I can't see when you have the camera on. And the reason was that it remembered from a past conversation that it couldn't see.
Speaker B: Oh my gosh.
Speaker A: And then that was stored in memory and then it's now.
Speaker B: So, so that was like one of its priors. So that was because. Right. It had been, you know, basically I guess trained itself to be like okay, I can't see these things. And so that's how I need to respond to this when I'm asked these things. Maybe. I don't know.
Speaker A: Yeah, so, so basically the previous con, it's a, it was, it's a, it's a really fun bug. It's um, it's one where I, yeah, it's just kind of those like mind blowing bugs that we get, we get that all, all the time. Um, but usually like uh, you know, when you're bringing a bunch of these capabilities together is where, which is required I think for mhm for the goal to have a universal system, for the path to AGI. The integration of a lot of these capabilities together in the same model creates issues. So it's, you know, it, it doesn't really know what to do if this memory is adding a bunch of random things here. Um, it needs to be proactive so it starts to talk more and then that combination for example like throws out more bugs. So uh, I think that it's a very challenging project to try to integrate a bunch of these, what we believe are critical for humans to be an assistant into a single system. And I think that has um, it's a very hard challenge uh, that we're trying to deal with every day.
Speaker C: Yeah, definitely. So what are some of the things that you have noticed that are like good pieces of advice or good, good learnings that actually can uh, help developers with addressing these challenges?
Speaker A: There are a lot of things that basically I think the live uh, API that developers would use um, would actually do pretty well, at this point, so I think the basic concept of having a coherent dialogue is like where I think we're, we're. You can always be better in these things, but I think it's reasonably good. Uh, with uh, sort of search grounding. Like, you can also ask like pretty like questions like what's the weather today? And like all of these things I think fundamentally work that alone, I think can create a lot of interesting products, including maybe a voice agent for your company that helps uh, your users. Um, the added aspect of vision creates a lot of um, new use cases. And I think a lot of companies, uh, enterprises, even folks in manufacturing are really excited about that. The idea that you can point your camera and something, trying to figure out how to troubleshoot something. I think those use cases are uh, pretty good. Um, and because just because the vision ability of these models are really good, you pretty much can point it at. I remember early days like you we pointed at like a, like a very famous painting and be like, who, what is this? Like, it'll say, you know, oh, that's uh, you know, Michelangelo's whatever. But then if you go to like a second tiered one, it basically doesn't work. And now a lot of these like true vision capabilities are solved. So I think you can actually do a lot with that. Um, the other one is that you can screen share. So it's vision, right? Vision's vision. So it's not just the physical world, it's your digital world which we prototype and uh, is now in Gemini Live. And it's really cool. Like it helps you with a thing that you're looking at on the phone. Um, once I had, I was like reading an article I was testing basically screen sharing with, with Astra. And then uh, which basically is the same kind of API that you get in Live API. And then I had a word that I had no idea like how to pronounce and it was, I forgot what the word was. I think it was like the name of a Greek philosopher or something like that. And I just like seen that word so many times at this point.
Speaker B: Like, I don't.
Speaker A: How do you say that word? And it just like circled it. I'm like, how do you say that word? It told me, I was like, oh my God, like, wow, this is awesome.
Speaker B: See, that's great. That's so cool.
Speaker A: That's cool.
Speaker B: And honestly, that's like open up in my mind. I'm like, okay, these are things that I should play around with and do because those are like previously things that you know, you would search for and then like you wind up on like the various sites that like have like the, the playback that may or may not be accurate. I would, I would actually love that if while I'm reading something be like hey Gemini, how do I pronounce this? Talk to us a little bit about like we kind of mentioned this before when we were talking about dialogue and, and the importance of being able to have like the back and forth thing. But one of the interesting things here, and I think that for me this was a really big moment when uh, we started to see this a few years ago was interruptability. Um, when it comes to dialogue. Can you talk to us about um, that feature and maybe how that has progressed and what the potential um there is and why that's important for these to exist, whether it's in Astra or in any kind of voice driven products.
Speaker A: Yeah, so interruptability has always been a pretty important aspect of all dialogue systems. Um so historically even we had barge in systems um, that essentially detect if the user is speaking and then will try to kill the audio um, while the user is speaking. So then the user can speak. And I think the challenge is that a lot of those systems you have to know how to differentiate between the users actually trying to interrupt you versus some other background noise like a car sound or like a background speech chatter. And then there were usually a bit slow and I, and I. And, and that speed matters. By the way, the thing that matters uh, in all things dialogue is latency. Latency. Latency. It's the, it's the one thing I
Speaker B: was gonna say that's, that's the thing. Right. Because yeah you could interrupt the systems before but it could take some time. And I think that this is, this goes back to I think even to your point about the translation things before too. Right. Like when we can reduce the latency. That's when I think people are more willing to use these systems. That's when we become more comfortable and that's when they become more useful because m. It feels more natural.
Speaker A: So I think the biggest innovations uh, actually is not just the LLMs themselves but essentially the infrastructure on low latency serving of LLMs. Uh, and there's still a lot of good stuff that we're working on that's coming out but the ability to do it fast and the goal is to get under 500 milliseconds for human dialogue. Um, but basically we're kind of there if it's just a turn taking thing. But things like Bargen does require Very fast latency. Now we're working towards getting to a place where uh, there isn't a separate thing that's trying to see whether was a you who are trying to barge in or not. Um, and it's just in a single system. Um, so a lot of these conversation mechanics essentially will get a bit smoother where when the model is constantly listening it understands um, what it's saying. So one of the things that we're dealing with right now for example is that once you start speaking you like, the model can like basically just doesn't hear you anymore. Right. So that's what a turn taking thing is. What we want to do is like it actually continues to hear you um, as you're speaking. Uh, and that's how we work. Like we don't shut off our ears the moment we open our mouth. So it's where right, we're listening to what's happening in the external world, what we're saying all at the same time. So that's been a goal, uh, for research called mixtape, um, or full ah, duplex dialogue for some time. Uh, so there's still a lot of really cool research that is basically headed in that direction. And that's when we're going to start having what I believe is supernatural dialogue. The other part that is related to Bargen is actually um, so the modality change by the way, if you remember how voice search actually works today, you press the mic, you say weather in London and then the mic closes and it does something. Or even if you do it like uh, if you were in the old Google Assistant days, you're setting timers, the mic opens, closes, opens, closes. And what we're doing right now and what the Gemini live experiences is that Mike is open the entire time. So the challenge with that is basically well, if there's background noise it's really hard and it's still I think one of the biggest challenges. But we're now able to build models that understand the difference between you know, background noise, side chatter. Is this intended for you? Is this not intended for you? And that enables the experience of uh, what we call open mic. So there's no mic on, off, mic on, off. We just basically have it on all the time. It's constantly listening to you. Um, yeah, so there's still a lot of work honestly in that space and it's going to take uh, but I think we have like some really uh, really amazing folks in the industry that's working on this right now.
Speaker C: Yeah, this is extraordinary. The evolution. I mean I remember again back in the day talking to Google Home and it's sort of being this very traditionally robotic experience, you know, where you're interacting uh, with, with sort of a separate machine. And now, now with as you said that the, you know, when talking to like the live API, it is um, it's incredible how the tones and everything and, and I think one of the things I'd wanted to sort of shift and ask you about is in a way there was a comfort with all of the bugs and the sort of robotic voice or the m. What to say like the machine, like conditions of, of what we were interacting with. Right. Because there was a distinction between talking to a human versus talking to a machine. And it was almost like you adapted, you know, the way that you speak to different people. You had started adapting it to a machine. Now what's happening is that those things are blending together and there is like I said, a comfort in knowing, you know, when, when I do hear like mess ups or, or things that are not working exactly, uh, as intended, there's almost a relief being like, okay, I'm still interacting with a machine, you know, I'm not interacting with a human. But as this technology is evolving and it's becoming so much more human, like there is a little bit of a concern of being able to distinguish, to distinguish between is this a human that I'm talking to? Is this a machine? Is this something that you think about? Is the goal to become, go into a space where both of these things are merged? Or is the goal to really, um, you know, is the goal to go in the direction of creating an assistant that is very human like, or is the goal to create its own unique version of and being an assistant, but with human like qualities, but very still obvious that this is the computer?
Speaker A: Yeah, it's a, it's a great question. It is a thing that we're definitely thinking about. Um, you know, the sort of, the ethics and safety side of these voice technologies have actually, they've existed for some time. But I think you're right that we're in a different phase now and we're trying to be super careful with this and just making sure that it is very clear that, you know, it's an AI that you're speaking to. But you're right that like our goal is to create a conversational AI that is able to communicate at the ability of a human being. So that means that it probably can be mistaken if not careful. Um, but the user perception of this is an interesting thing. So before, like it was at least understood to me that no matter how much I. No matter how I talk to my Google home, it would never understand, for example, like, that I'm angry.
Speaker B: Like.
Speaker A: Right.
Speaker B: I mean, I curse at mine all the time. I still do.
Speaker A: But, like.
Speaker B: But it does. But it doesn't under. It might pick up the words but like, and understand that if I use like, uh, you know, expletive and stop, that it needs to stop.
Speaker A: But.
Speaker B: But it's not understanding that I'm angry and that I'm, like, upset that it's doing something. Yeah. Ah, totally.
Speaker A: Yeah. That's, uh. You know, it's funny that you.
Speaker B: You still.
Speaker A: You still do it, and I do, too. Just like, why isn't this thing working? But, um. Um. But I think once you're like, it'll take some time for the user expectation to be changed when you realize, oh, it does understand that I'm angry at it. And it just. But it takes a few of these experiences for you to realize, okay, this is. This is very different. Um, we were once, I was asking some of my, um, the PM's on my team and the researchers, I said, if we were to put one of these voice agents, you know, and compare that against a person, um, and just ask someone to, like, interact with it, and they're supposed to guess, like, how, um, you know, like, are they speaking to an AI or are they speaking to a human? Like, what do you think the chances of are? Like, that they actually know that this is an AI.
Speaker C: The Turing Test of voice.
Speaker A: Yeah. Yeah. I mean, that's kind of like where I think some of this, like, from a capabilities perspective, we're probably gonna get there. And.
Speaker B: Yeah.
Speaker A: And I actually was thinking, well, one of the only ways I know. I knew today that I'm talking to a person is this fluencies in speech. So, you know, your little pauses, the fact that I said yes, yes, a couple of times, and a little bit of, like, these unnatural, uh, breaks and so forth. Um, our AI models don't have that. Is that a good thing? Is that a bad thing? I. I don't know. I think it's. It's one of those things that we're still trying to figure out. But our goal is not. Our goal is very much to not to fool a user that it's a. It's a human. But we want it to be able to communicate with it and just being able to, like, you know, just say anything the user wants to say and be able to have a really good conversation to help them understand Something that they're researching or help them do a task. Like, that's fundamentally our goal, but it is a pretty important, um, area of, like, safety research and ethics research.
Speaker C: Yeah, that was one of the things that I, um. Like, you just brought up this example of, you know, when you're thinking about an assistant. Because if, uh, a person is hearing that you're angry, their response might be to, you know, I'm gonna step away and wait until you calm down. But, like, do you want. Like, if you're driving the car and you're saying, you know, hey, I'm frustrated. I'm in traffic. Help me get out of here. You know, you want the assistant to say, hey, I hear you're. You want it to respond in a way that's not gonna get you more angry. That's gonna understand you're frustrated, but it's not gonna be like, I'm go give you a timeout so that you can reset yourself, and I'll come back in a minute when you're feeling better. You know, like, that's. That for me, is like a big. One of the advantages of having it being a machine, in a sense, is that, like, it's like, I hear you're angry, and I will try to help you as best I can. Here's what I understand the problem to be versus, like, sorry, you need a timeout. When you're feeling better, I'll come back.
Speaker B: Or worse. Or worse. You don't want it to yell back at you.
Speaker A: I know you do.
Speaker B: That would be kind of funny. But you don't want that. Right? Because. Because that's the thing. Right? Like, I had kind of the same thought, which is that as these things get emerged, I was curious, like, do you. Have you noticed in your research and based on how people are using these things. Because we've read things about how people are, uh, you know, there were, like, rumors that if you're more polite to system your assistance, that you'll get better responses and whatnot. Have you seen that people, um, are more polite, maybe, uh, in the languages that they're using, with things like Gemini Live versus maybe what you were seeing, uh, in terms of how people like myself might speak with their Google home. Yeah.
Speaker A: So, um, we haven't seen people being more polite, um, in our data, but we do see people being kind of, like, more demanding in some ways,
Speaker B: because
Speaker A: we say that it has memory and it does. Um, but also, I think just the concept of this fluid dialogue thing, that sounds way more human. One of our, uh, trusted testers uh, told us that he basically spoke to it for hours and hours and hours. And I'm like, would you speak it? Would you say what'd you say to this thing? And he basically said, I. I really carefully try to craft it to the assistant that, that I want. Um, so we, you know, I have no idea what he said to it, but I think it's along the lines of like, you know, what, you know, what are his, his favorite colors and what are the things that he likes? What are the things that he doesn't like? Um, you know, how we want to be addressed, you know, how short or long to the responses from the assistant be, which apparently is always kind of like uh, one of the challenge or challenging things. So people are expecting, they are expecting more out of the AI um, and I think they will, they will behave differently. One of the things I actually really want to do with the, um, you know, knowing that the user is like frustrated or angry is actually saying, okay, so what have I done wrong?
Speaker B: Right?
Speaker A: Like, uh, they're obviously angry at you. Maybe for, I mean, in the case of your driving, maybe you're just generally angry. But like, yeah, uh, I mean it
Speaker B: could be unrelated, but it might be related to the system.
Speaker A: It's probably related. You know, they're just like saying a thing like again and again. It's just like not doing the thing and then they're like getting more and more frustrated. But then we, in the traditional systems, like, we don't do anything with that information. And now it's like, can we learn like, okay, what did I do wrong? And even if you can't do the right thing at this moment, just knowing and say, hey, you sound angry, how can I improve? Um, I think that would be amazing capability, uh, to build towards.
Speaker B: Yeah, no, it would. Right. Because that would in some ways be like a great way of like, you know, training back. Right. Like, it's not just so much, you know, saying, oh, was this a good response or not? But knowing. Okay, well, this response elicited, you know, uh, based on the tone of the voice, you know, the person was happy with response versus not, which might then be useful in terms of like broader training stuff in terms of how it's going to respond and what it's going to do.
Speaker A: Exactly.
Speaker B: This has been such a great conversation. Um, and um, we uh, wanted to, uh, we have a section that we do at the end of every episode where we kind of have some rapid fire questions for you. So we ask you some questions and then you just give us like your quick answer or whatever comes to mind, and if you don't have an answer, that's okay.
Speaker A: But we just, this, uh, is just
Speaker B: sort of a fun thing that we have. So, um, what was like, the last thing that you automated for yourself?
Speaker A: Yeah, uh, I mean, I don't know if this counts, but, um, I basically, like, fed, uh, Gemini a lot of reports that I don't understand about, about, like, all kinds of topics. And I just said, I just want to talk about this. I just want to just like, give me, give me a basic, uh, so I, I, I guess I kind of now use it as like a, an automation system for learning about new things.
Speaker B: Yeah. Do you, um, type or do you speak more with, With, Ah, with Gemini,
Speaker A: I do a lot of both. Uh, I do a lot of both, and I actually, my favorite thing to do, uh, this is a, uh, this is a trick that I don't know if it's available to everybody. But I'll tell you what the think my favorite thing is. I use Gemini Deep Research on a topic, uh, and usually something pretty complicated. Uh, financial, medical stuff. I'm like, I don't understand this thing. Just go and figure this out for me. And it creates a doc PDF. I make it into a PDF and I share it with Astra, or basically I give it as context. I'm like, here's the thing about this topic, and then I just talk to it. So I do this actually a lot now, uh, because to me is the perfect combination of I want deep research high, but, but I don't want the output to be. I don't like reading a lot versus talking about.
Speaker B: Oh, I love that. That's so smart. That's so smart. So, okay, take advantage of the deep research, but then interact with it using voice. I love it.
Speaker A: Yeah.
Speaker B: Okay, so what was the last thing that you asked Gemini? Whether it was by voice or text?
Speaker A: Uh, a lot of questions about taxes in the uk. There's a lot of changes for, uh, UK taxes. And I talked to my tax accountant two times, and I still don't understand after our phone call, like, what does that mean? Um, I don't know. It hasn't been a great experience. I would say so. Finally recently, I just, like, I'm gonna just, like, sit down in front of Gemini and I'm just gonna ask him about these questions that I don't understand.
Speaker C: Was it helpful?
Speaker A: Yeah, I think I finally understand it now, the first time.
Speaker C: Wow, that's awesome. Awesome. Um, so what can you do now that you couldn't do six months ago
Speaker A: I think the most interesting thing that I would want it to do is effectively let it code and use tools to do something on my behalf. Uh, we're starting to dabble into this and I don't think you can do that. 6 months ago um, I think as an assistant this can become very powerful. Uh, that would be a new thing that I think is now becoming more possible just because there's a lot of stuff I think that could be done with code that we're excited to experiment with.
Speaker B: Okay, so this is a big one but what are the next set of problems that you're trying to solve?
Speaker A: There's a couple of different things in the audio and vision uh space. There's audio vision and there's assistant. So maybe I'll pick a couple of interesting bits in the audio space. I think the thing that I'm still very passionate about is basically the human level dialogue system. I don't think we're there yet but I think we're going to keep going and we have a path towards a supernatural super fluid conversation that is fully duplex. Like, like we are so um, very excited by that. Um, in the astro space, um, the area that I think is unclear how you make the biggest amount of progress is actually the execution part of being an assistant. So this isn't actually has to do something stuff for you. We focus mostly on helping you do stuff which is talking to you about stuff. But doing stuff for you I think is neat. Like absolutely needs to be solved. And there's been a couple of different approaches. We actually showcased what we called uh, Android Computer Control uh at our most recent uh Google I O video. And the idea is that an AI understands your, your computer, your screen. Uh, so Project Mariner from Google is in the same direction where it just like it does stuff for you and it like you know it can like maybe like today uh, I was talking to someone who is talking about this like really terrible website to book tennis. Like nobody wants to actually really wants to use that website to do it because it's really hard to use super slow. But wouldn't it be great if your assistant can like go look at this website, you know and like they'll tell you when is the available tennis court and actually book it for you. Um, so that, that like basically being able to do a lot of stuff for you at scale using either um, APIs or some sort of automation, agentic automation system. I think that's going to be the biggest, like the biggest piece from a universal Assistant perspective that we, we need to solve on the research side.
Speaker C: Yeah, well, and with the advancements of agents definitely heading in that direction.
Speaker A: It's definitely heading direction. I'm excited for that.
Speaker C: Yeah. Yeah. Well, what is your biggest learning in the past year about the future of work?
Speaker A: Of work. Work.
Speaker B: Maybe let's reframe it this way. Like what's maybe your biggest learning in the last year about like the future of communication or future of dialogue?
Speaker A: Uh, on the communication front, I think, um, the, the learning that we're having is basically how to build these systems that really can communicate with you in a super fluid way. Um, and that I think, you know, LLMs are the answer and I think it'll take us to essentially the human level dialogue system. And I think that piece of core technology will enable a lot of things. Like I have no doubt in my mind that um, an AI that can communicate with you but has access to all of the world's tools and knowledge, uh, can make a lot of different kind of products, uh, in, in the. Out in the world. Yeah. But I think on the, on the work side, uh, it's, it's very interesting. I think everything is changing right like in terms of our day to day. Um, it's really, it's really taken off in the last. Right like about the last year or so where like literally everything you do is changed in a significant way. And I think that it comes from like little things, but it comes from also big things. One thing I started doing recently was the way I write is completely different now. Um, because there's really great like AI, um, refine capabilities where you can like write badly and, and then it'll just like make it uh, grammatically correct, put it in bullets, like punctuation and everything. Like it just, it just makes it sound and much more coherent, better. So just like if I think about this as like a communication mechanism, I think writing is also a communication mechanism. Like what I do now is like I, instead of kind of being what I have to like write a proper doc or something, you always comes with a little bit of trepidation, right? Like, oh God, like it's like a, it's like a thing. Like what, how do I go about it and what I say? You're like typing, you're deleting and typing. And now I just like spew out everything and I just like just, just type just random stuff in my head. Like in bullets. Just, just put it out there. Uh, it's very cathartic. I highly recommend it. Um, you probably could do that with voice, too. And then.
Speaker B: And then just say, fix it.
Speaker A: Yeah, and just say fix it. That's it. Just like, you know, I wrote this huge thing and it looked really good. Uh, I sent it to something, somebody recently. But the way you do it is just that you. It's a way you communicate your thoughts. You just. Now you can just kind of just spew out, like, concepts and ideas and thoughts, and then you can essentially get the help of AI to draft it in a way that's understandable by people, that's a little more concise. And then. And then you can, like, you know, mess with it a little bit yourself to make sure that it is. It is the most appropriate format for what you're trying to communicate. But I think that's amazing, right? It's a way you, again, express, uh, yourself is just completely different in this world.
Speaker C: What an exciting future we're going into, and it's going to be an evolution in communication. So thank you so much for this conversation, people. How fascinating.
Speaker B: Yeah, yeah, thank you. You're amazing. Um, and I think this is so exciting just to hear about your journey and where you've been and where things are going. And I think we're all really excited and lucky that we have people like you who are working on solving these sorts of problems, because it's a fun time. Um, a little scary at points, but it's really fun.
Speaker A: Thank you. Thanks so much for having me. It's been a really fun conversation with both of you.
Speaker C: Thanks for joining us. For more information about our guests today, please check out the description.
Speaker B: And for more great conversations like this, hit that subscribe button. Until next time, thanks for listening.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.