
Denoised · 2026-05-28 · 40 min
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
Google's Gemini Omni represents a shift in video generation architecture - it's a world model, not a diffusion model, making it exceptionally good at video modification and inpainting tasks. Unlike Kling or Runway's cinematic generation, Omni excels when given an input video and a modification prompt: change the dog to a robot, restyle this miniature city, make a building take off like a spaceship. The avatar feature, hidden in the UI but remarkably refined, uses phone-based calibration (saying numbers, not phrases) to create likenesses that rival or exceed Sora's quality, with improving voice synthesis that captures inflection better than competing systems. The model demonstrates sophisticated world understanding - it can generate multi-object sequences spelling out words, maintain camera movement while completely restylizing scenes, and reason about spatial relationships. Guests Joey and Addy contextualize this within Google's broader I/O narrative: Demis Hassabis announced a three-year AGI timeline, Google's commitment to AI for Science (echoing AlphaFold's protein-mapping success), and the growing market divide where AI-specialized roles command million-dollar compensation while non-AI positions face talent oversupply. The episode also touches on Yann LeCun's work on continuous learning architectures and troubling real-world failures like Waymos driving into flooded Atlanta streets - exactly the edge cases Genie world models are supposedly training against.
Omni is a world model with physics understanding and spatial reasoning, not a diffusion model like Veo. It excels at video-to-video editing and inpainting (modifying existing videos) rather than text-to-video generation, making it fundamentally different in architecture and use case.
The avatar feature uses phone-based calibration where you record yourself saying numbers to create a personalized likeness tied to your account. It produces near-photorealistic results (estimated 98% accuracy) with improved voice synthesis, though it subtly enhances appearance and doesn't retain 100% likeness by design.
Yes, the model can generate multi-shot sequences where objects are selected to spell words (D-E-N-O-I-S-E-D) while maintaining natural composition, demonstrating sophisticated world understanding and reasoning capabilities.
Hassabis stated AGI could arrive in three years, though the hosts speculate the realistic timeline is five to six years out. He defined AGI as AI matching one human's learning capability and proposed testing it by giving models knowledge only up to 1910 to see if they could rediscover major scientific discoveries.
There's a sharp divide: AI-specialist roles face talent shortages and command million-dollar compensation packages, while non-AI positions have talent oversupply with thousands of applicants competing for roles, creating increased pressure for non-AI workers.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains moderate substantive content about video model capabilities and architecture (Omni vs. Aleph 2, world models vs. diffusion models, video-to-video editing), but is frequently diluted by casual banter, tangential discussions about swag and Waymo anecdotes, and repetitive example walkthroughs that don't add new information. The technical insights about how these models work are present but scattered.
This is the video version of Nano Banana Pro. It's a world model. It has world understanding, and it understands physics
this is like V1 of this new Omni model. And the things that it is really good at is editing videos and modifying videos
The episode mostly reports on Google's announcements without significant original analysis or contrarian takes. The hosts provide personal testing results and comparisons (Omni vs. Aleph 2), but largely validate marketing claims rather than offering fresh perspectives. The framing of video models as 'editing tools' is somewhat useful but not particularly novel.
don't look at it as, like, a video generation model- Right... but a video modification model
There is no, uh, grand master plan to tie everything underneath with, uh, like a master workflow
This is a co-hosted discussion between two podcast hosts reflecting on their Google I/O attendance. Neither are practitioners/operators with hands-on experience shipping video products at scale; they are content creators and commentators on AI tools. One brief mention of speaking to 'Justine Moore from Basics NZ' is insufficient to elevate guest caliber. The conversation lacks input from actual AI researchers, video engineers, or practitioners building with these models in production.
I was speaking with, uh, uh, Justine Moore from Basics NZ, and she was like, 'Oh, you got to try the avatar feature.'
they're, they're, they're, they are aware of what happens online [laughs] and what, what people talk about, so they're always looking, you know, even if it's just bug fixes
The episode includes concrete examples and video demos with specific outputs (robot dog transformation, LACMA miniature city, building-as-spaceship, avatar comparisons), but lacks quantitative metrics on performance, resolution specs, pricing, latency, or detailed technical specifications. Many claims about model behavior are illustrated visually but not backed by numbers, benchmarks, or rigorous testing methodology.
change the dog to a robot. Hmm... And, like, the dog turned into the robot. The movement looks good. Everything else looks pretty stable.
I thought this is way better than Sora's avatar was. It's way better than Sora's avatar. And even, like, I think you used avatar for the champagne shot here- It was the avatar, yeah... as well as the- These were all avatar shots... the Miami Vice gray T-shirt shot. Like, to me, that's 98%, Joey.
The hosts ask reasonable follow-up questions (e.g., about avatar sharing, Flow API specifics, character consistency) and occasionally push back on claims, but the conversation often meanders into tangents (Waymo honking, Meta layoffs, AGI timelines) unrelated to substantive video model analysis. There's insufficient probing of technical limitations, use-case constraints, or critical examination of marketing messaging. The tone is collegial rather than investigative.
And I asked, I'm like, 'Well, is there a plan where you could, you know, if you give me permission, like with Addy, can I pull up Addy's avatar if he gives me permission'
I was pegging them with questions of like, 'Is there something in the model?' Or when you- when they eventually roll out the API, 'cause there isn't an API yet, um, will there be something that like treats characters differently as other reference images?
Computed from the transcript - who did the talking, and the words that came up most.
Joey returns from Google I/O with hands-on tests of Omni, Google's new video world model, comparing it head-to-head with Runway Aleph 2 on the same shots. Plus: Demis Hassabis puts AGI three years out, Google Flow gets agentic workflows, and Google Pics turns AI-generated images into editable layers. - The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently
Transcribed and scored by The B2B Podcast Index.
I think the best way to think about it and the way that they were explaining it at the event was like, this is not Veo 4, and it's a completely different model. This is the video version of Nano Banana Pro. [static] All right. Welcome back to Denoised.
Addy, good to see you. Nice to see you, Joey. Welcome back to town. I am back in town.
I only just spent the last few days at Google I/O, which was a lot of fun. I am decked out in swag. I got a Gemini hoodie and vibe blank, vibe code hat. Maybe I'm just getting old, but when did tech gear free stuff start to look so good?
[laughs] It got- 'Cause remember- It got tender to it... you used to go to CES, like, 15 years ago, you'd get these free T-shirts. You're like, "Ugh, I can't wear that." [laughs] Yeah, the husky shirts- Yeah...
or whatever, [laughs] the cheapest ones you could mass print- Exactly... and give out. Uh, no, this stuff's very nice. Okay.
So yeah. Well, there's a lot of Google updates, obviously, and a lot of I/O updates, or updates from I/O. Uh, so let's talk about it. Yeah.
The biggest one, Google Omni- Mm-hmm... their new video world model. Yeah, big update. Um, I know we get video models that are coming out every couple of weeks, and Google also dropped...
You know, they're also in the fall. But what I noticed, I mean, please go through the Omni specs with me, and then I'll tell you what some of the things that I thought was really amazing. I've been sitting on this [laughs] for, since, for about a week, of like, where does this fit in or what is this good at? My initial...
When I first had access and messed around with it, obviously my first jump to, like, okay, how does this compare to Kling and C-Dance? And at, on the surface, I would say for, like, if I'm looking to do cinematic visuals, it still is not at the level of, or, and it's a completely different model. Um, this is the video version of Nano Banana Pro. It's a world model.
It has world understanding, and it understands physics, and it... Basically, Veo is more of a diffusion model, uh, that was just really good. And so this is completely different. This is like V1 of this new Omni model.
And the things that it is really good at is editing videos and modifying videos. I feel like we're gonna get stuck with calling things editing videos. I still think editing videos is like snip-snip, you're cutting videos together. But that lingo has- Yeah, I call them video-to-video models...
video to video or just video f- vi- video modification. Yeah. You give it an input video and you change something. Yeah.
Inpainting, yeah. Inpainting, yeah. Inpainting a video. Um, really good at that.
Really good at just giving it world understanding prompts, and really good or pretty good, um, avatar feature, which we can talk about in a second. I thought the, the, like, the biggest sort of, um, you know, the biggest gain in quality was definitely building a human from the likeness of like, I don't know, uh, like a one-minute or 90-second calibration- Not even... very similar to what you did with the Sora app. Yeah.
So the... Okay, so yeah, the avatar feature, let me say if you give it an input image, and I tested it with, like, a couple of, like, photos of me or my wife and, like, image-to-video, uh, and it has... And if you go on the Gemini app, there are a lot of, like, templates set up of just, like, fun things you could do with your stuff. Yeah.
It really, the likeness fell apart, and that was my first impression when I kind of gave it some images of me from, like, a trip, and it tried to make a video. It completely changed our faces. Um, and my initial impression was like, "What the hell is this model?" [laughs] What's the fuzzy?
[laughs] But yeah. And then, um, actually at the event itself, uh, I was speaking with, uh, uh, Justine Moore from Basics NZ, and she was like, "Oh, you got to try the avatar feature." And I was like- Right... "What avatar feature?"
[laughs] And it's, like, buried in the UI is this avatar feature that it's attuned to your account. You basically use your phone. It's very similar to Sora. Mm-hmm.
It has you, like, turn your head left, turn your head right, and then you say a bunch of numbers. Yeah. And that is your avatar. It's tied to your account, and then you can make videos of yourself with your avatar, and the quality there is way better.
Oh, so good. I thought this is way better than Sora's avatar was. It's way better than Sora's avatar. And even, like, I think you used avatar for the champagne shot here- It was the avatar, yeah...
as well as the- These were all avatar shots... the Miami Vice gray T-shirt shot. Like, to me, that's 98%, Joey. Like, I, I've seen you in person enough to know, like, like, maybe you're not- If you see this, it's like a little off...
you're not this much of a stud. No offense. [laughs] But, like, it's pretty fucking close. Yeah.
No, I like Omni, that it definitely, you know, kept giving, making me more jacked than I am. [laughs] Yeah. It, it just ga- gave you a sharper jawline and just fuller hair. Yeah.
Like, it just, like, amplifies... Like, the common denominator is, is, like, a really handsome man or a woman, and it just kind of pushes your avatar into that direction. It doesn't retain 100% of your likeness, but I think that's a creative choice on Google's part. I think if they really wanted to, they can dial that aesthetic thing down and just keep it true to who this man is.
Yeah, I, there's probably a big balance of just, like, what is it actually capable of and how much are they letting people do? There were times where I'd try and test out prompts and stuff, and it was just like, either it was overloaded or it couldn't do it or wouldn't do it. Yeah. Um, and sometimes it was like, "Hmm, what's the reason behind that?"
The other, the other really impressive thing about the avatar feature, there you go, yeah, you're pulling it up right now, is the, is the vocals. Like, the, the AI voices miss a lot of the inflections and some of the, like, the "hey," you know, and then the "ooh." Like, the range is very limited, but here I feel like they're, they're getting a little bit better with that range. It's still not 100% as natural as how we sound every day, but it's getting there, and it's, it...
I think Omni really made a big progress with that. Kinda. Like, the issue with the avatar, and so you can see, like, this was the source image. Like, I trained this at night in my hotel room on my phone.
So from the visual quality, uh, really good of what I was able to pull out from my phone. The audio of my voice, let me try the phone... microphone because it was. And there's no, like, enha- like, no Adobe podcast audio.
There's no, like, audio enhancement that this model's doing. Google just dropped Gemini Omni. It's Nano Banana for video. Yeah, that doesn't sound like you at all.
[laughs] It's funny because, uh, some, some of the other examples that I've heard is pretty spot on to the people that built the avatars. Oh, really? Yeah. Mm-hmm.
Maybe, like, uh, you know, if I plugged in a DJI mic, it, and trained it with a better microphone, then maybe it would give me better outputs. I don't know if it's the mic. Does it make you just say the same thing for everybody? No, it's in the prompts.
It's whatever... Oh, oh, oh, you mean- Like the calibration process. Yeah, it's just the calibration is literally you're saying numbers. Okay.
That's it. That's it. Okay. Yeah.
But I'm just, you know, I'm just holding up, like, you know, three- So you don't say, like, Mary sells- 24... she, she sells by the sea No, you're literally gonna say numbers. Yeah, okay. But you're, it's whatever, you know, your mic is here, so- Yeah...
yeah, if I had a better mic plugged in, then maybe I would get better outputs on this. Yeah, maybe. Yeah, but I, I was quite impressed by the avatar feature more than anything else on Omni. Al- although Omni does have really good reasoning in the same way that Sora does, uh, whereas, like, if you give it just an idea, it'll expand on that idea and really go to town on creating the entire editorial or the cut.
Well, yeah, let me show you some of the outputs. First up, in honor of you. Oh, yes. [laughs] The Grifft 1970s New York City walking down the street.
Mm. Although it is technically... I mean, these cuts are incredible, right? Yeah, and this is a very basic prompt.
Like, I didn't give it any info. It just kind of reminds me of, um, Saturday Night Live, um, like that Travolta scene where he gets the pizza. Two slices. Two, two.
And then he does the walk. [laughs] World understanding stuff. This was a video-to-video test, and so the original video was just, like, my wife walking the dog- Mm-hmm... uh, on Venice.
So it's like, okay, dog's standing, panning around. And then I said, "Okay, turn the dog into a robot." Hmm. And, like, the dog turned into the robot.
The movement looks good. Everything else looks pretty stable. Yeah, no, no distort- the only thing I saw was her left hand and the leash was a little funky when she grabs it again right there. But I'm being super picky.
Obviously, the, the sheen on the, and the finish on the metal seems very, uh, plasticky and not as reflective and, you know, responsive to the environment. Yeah, I think those robots are matte plasticky. Yeah. It even has the same kind of movement of it, but it also kept his arm placement on him.
I'm being super picky. This is incredible. Yeah. Yeah.
Kicking like robot dog zero. Yeah. [laughs] Yeah. Yeah.
[laughs] Like, wait a minute, this is- Yeah... the best. Well, 'cause also I will, once we get to our later updates of OLIF 2, I will- Oh, okay. Yes...
I may process the same shot and- Looking forward to that. [laughs] I haven't tested. I assume you could give it- Mm-hmm... I could give it an input image, like a restyled image, to give it more direction.
It'd be like, change the dog to this. This is my prompt is literally- Right... change the dog to a robot. It goes and picks the robot that it wants.
Like, I gave it the most basic prompt possible. Yeah, pick the robot, but also, like, you know, literally understood the dog, swapped the dog out, everything else, like the transparency of the background. Everything else looked pretty smooth. But the fact that it changed it to a robot dog- Pretty crazy...
and not, like, a robot the size of a dog that's a biped, that's reasoning. Like, it's going through some sort of intelligence, you know? Yeah, it kept the same- Right... it kept the same dog thing.
Yeah. I mean, I also did this one with same, same source video. I said change him to a chi- chimpanzee. Oh, that looks better.
See? Like, when you don't do metals and reflections, to me, that feels more real. Yeah. These look really good [laughs] out of the box.
It just kind of... These video-to-video, once it was like, okay, don't look at it as, like, a video generation model- Right... but a video modification model. This one, this was a shot from one of my favorite installations at LACMA.
I forgot the name of it. The, um, Metropolis City with the miniature cars. Oh, absolutely. Really?
Oh. It's like a permanent installation. It's, like, a massive permanent installation. Those cars are running full speed.
It's super cool. Yeah, it's, like, a bunch of, like, Matchbox cars in this, like, crazy futuristic city, and I said restyle it to a futuristic city but, like, keep the same geography layout. Oh, yeah. They kept that shot right.
Yeah. And kept the cars- Mm-hmm... kept the stuff. A little bit of warping there, but, like- Mm-hmm...
this was one of the best outputs where it still kept the- The shape of the tracks. Yeah... structure, the camera movement. Yeah.
And it just changed everything else. One more. This one was actually fun. This one, I just gave it an image of one of the Google buildings, 'cause they had a bunch of these, and they just looked really awesome.
And then I said make the building take off like a spaceship. Yes. [laughs] Yeah, it did. [laughs] It's like it kept the building.
Oh, my God. [laughs] It's like it- Yeah... added this extra shot I wouldn't have thought, but, like, it kept the building. Yeah.
It had it shake- Right... and bust out of the ground- That's awesome... and have rockets underneath. Crazy.
Um, okay, the world understanding this. I said create single shots of vintage objects. The first letter of each object spells out the word denoised. The letter should also be on the object- Mm-hmm...
in a natural way. So basically, I wanted the first letter, if it was, like, uh, you know, whatever, like, I said vintage objects, but if it's, like, dog, like the first letter's a dog. Yeah, it's a, it's a good, uh, reasoning test. Yeah.
Uh. E-N- Yeah, so does it understand objects?... O- Newspaper, oils. Yeah.
E. Uh. D? [laughs] Didn't get the last one.
[laughs] Don't know what the last one was, but it got basically, uh, dial phone- Yeah, the film and all that is so, so right... Edison bulb 'cause I asked them to, like, explain- Love that. Yeah. Yeah.
I said explain what you, what the list was. Uh, a- and it also, the list objects here are slightly different than what I put in the video, so the understanding between the two is a little mismatched, but N, newspaper, O, oil can, uh- The ink thing- Do it... with old-timey pens. Yeah.
Ink, ink holder. I, yeah. Um, S, stopwatch. E, I think that was Edison fan.
Eyeglasses? Okay. Uh, all right. Well, look.
I don't know. And then the key, I don't know. Did you generate this in one shot? That one went off the rails.
But- That's crazy... this is literally this one prompt. The fact that it i- is able to generate so many different things- This is a one prompt... and then cut it together into one clipYeah.
That's impressive. Yeah. That's where the world model part comes in, of like it knows objects, letters, uh, you know, much sharper at text rendering. Yeah.
It just, it just trained on the internet, man. Pretty crazy. [laughs] Like, it just knows every single object known to man- [laughs]... through the, through the years, through the decades.
Yeah. Yeah. Like it'll know all types of bulbs, not just this bulb, right? It'll know fluorescent bulb, LED bulb, and so on.
So Omni, probably the biggest announcement. Super cool. They also had Gemini Flash 3.5- Oh, the names-...
and 3.5 Lite... always confuse me. Tell me Nano Banana.
[laughs] Yeah. Names are confusing, and it's still basically they're like, it's not the pro model, but it's beaten benchmarks for them that their 3.1 Pro has, and that they're like kinda pushing it as the new- For video?... main model for now until there's like a pro model.
Oh, no, no, no. The actual LLM. This is just Gemini, like their text model for coding. Okay.
Yeah. The LLM, yeah. 'Cause, uh, correct me if I'm wrong, but the official name for Nano Banana is also like Gemini something, right? It's like Gemini Image three-point something.
Yes. Yes. Yeah, no, you, you said you saw Sundar, the CEO, as well as Demis Hassabis, uh, speak on stage. What, what, what did they say?
What happened? Oh, yeah. I mean, so yeah, they had the keynote. They, they, they, they did that, and then there was like a kind of fireside chat side things, uh, that you could go watch.
And so yeah, I know for Demis, it was AGI coming in- Oh, wow... that many years Three years. They pushed that timeline way back. That's his timeframe.
Yeah. That was his timeframe. Yeah. Well, okay.
That was his timeline. [laughs] Yeah. And there was definitely a bigger push. You know, his was interesting 'cause like it was a, you know, debate of like the AI doomy, doomers versus like optimism- Mm-hmm...
realistically optimistic or something. Um, but basically it kinda came down to like how you talk and frame, uh, when ta- uh, around AI and the risks it brings, but also the benefits that it could bring, which is also why they kinda did do a big push on like AI for science, and I think it's a whole separate, um, division now or focus on- Yeah... with Google- Yeah... AI for science.
I, I'm glad they're doing that 'cause I don't think, um, OpenAI and some of the other competitors are, are probably not investing as heavily in s-science. Yeah. Not, not talking about it enough. Yeah.
And that was like his original, uh, you know, a lot of it, the original DeepMind stuff with, um, AlphaFold, uh, you know, mapping- Yeah... um, proteins and stuff. So like, y-you know, that's something. It's like if you can get rid of diseases or cure diseases, like no one- Exactly...
no one is saying anything about that. Yeah, and if you can cure cancer, then you can go make a little bit more slop, and it kinda- Mm-hmm... just evens out. [laughs] So they understand it better than we do.
[laughs] Um, and yeah, I, I think they're ha- like initially, if you remember a couple years back when like the first ChatGPTs were coming out, they were like, "Yeah, at this point we're gonna be able to like figure out, uh, nuclear fusion, and cancer research," and all, all the stuff that needs heavy, heavy supercomputers, they're like, "Yeah, we can do it now." And then it kinda just went away, and then we had a bunch of [laughs] meme generation capability and, uh, you know, a lot of job displacement unfortunately.
Like, uh, a lot of companies bet on AI and had to, uh, fund the capital expenditure, so they have to fire people to fund that. So it was just all been negative since then. So I'm hoping like with Demis' sort of push back into goodness for humanity, maybe, maybe, just maybe we can solve something huge here that will really benefit mankind. Yeah.
I mean, yeah, I think, I think it's, there's like a big just what is the benefit to me thing with AI and all this stuff, and it's like kind of been hard to see, especially with like a lot of displacement, job loss, all of that stuff. And yeah, I think if you could just be like, "Oh, well, you know- Yeah... we can improve human health and wellbeing and cure terrible diseases," that is a clear benefit to society that is hard to dispute. So about 12 months ago when we started the podcast, they were saying that, "Hey, AGI's gonna be here very soon, but ASI will take a while."
So general intelligence is, uh, what I think is defined by an AI system that is as good as one human being, that can learn as well as we can, new skill sets and so on. ASI, artificial super intelligence is some, it's like a bunch of human intelligence put together into one system, so it is super intelligent than us. Mm-hmm. And if they're saying AGI is now three years out, I'm guessing it's more like five to six years out, [laughs] and ASI is probably decades away.
It's probably a good thing. It's probably a good thing. We don't need this stuff right away. There's enough disruption as it is.
As you saw just a couple days ago, Meta had th-I think the biggest layoff yet, right? 10% of the workforce. Something like that, and they were also like everyone that was staying- Yeah... they're like training models.
There's leaked audio from the Zuck that, yeah, we're, we're training on everybody 'cause you're, you're, y-you guys are really smart, yet we don't need you here. I haven't dug into the full post, but also the CEO of ClickUp, they did a layoff, and g-I, the ji... I didn't read the whole thing yet, but like the gist was like, you know, people that stay, you know, AI should make them 10 times, 100 times more effective. But like, I'm also looking at compensation tiers- Yeah...
that are like a million dollars, you know, if you are the person who like leverages AI and then could do like 10x, 100x more and then get compensated appropriately for that. Yeah. There, there's a really interesting divide in the talent f-pool, like in the job market now, and I'm, I'm kinda starting to notice it a little bit more, and it's becoming more and more pronounced, is that there is an oversupply of talent on the... Like there's a hard fence in the industry on AI jobs and not AI jobs.
The side of non-AI jobs, there's an oversupply, and so you're fighting and wading through like thousands of applicants and so on. On the AI side, there's not enough people. Like they literally can't find people, so they have to overpay and get those million-dollar contracts out to hire whoever. And the other about AGII think he was asked, and I'm gonna obviously paraphrase and try to remember the best I can, but he was asked like, "How would you even know, like, what, you know, if it does achieve that?"
And Demis's, one of his tests was like if you had the AI model and you just gave it world knowledge up to like early 1900s, like 1910 or something, would it be able to figure out, you know- Right. Right... relatively- Exactly... like things that Einstein discovered and other- Mm...
things that humans discovered in that timefr- uh, you know, after that time period. Uh, that was like one of his benchmarks for like determining- Yeah, it, it all goes back to what Yann LeCun is doing now, right? Which is, which is that when we train a model, during the training it, it absorbs and learns everything, but then after training it's ha- the clay has hardened, right? So you can't retrain it unless you train a new ChatGPT version or so on.
Like it has to be an entirely different model. So how do you build in mechanisms for a model to keep learning obsessively over and over? The same way you and I, like, hey, do you remember a time, Joey, before cameras and before editing and before color science? Like you learned all this stuff, man.
So like how do you make an AI do that? Mm. And that's what Yann LeCun is working on is, um, he's abstracting away the notion of tokens essentially. So instead of like data captioning and images and videos being the tokens, how do you have a more abstract form of neural network that just relies on vectors and like there is...
I'm forgetting here, but ba- e- essentially he is making a, a more abstract version of a neural network where it's multimodal by default, and it's much more than that. It is also able to absorb inputs much more easily into the network. Yeah, what I could do is I'll do a little bit more research, and then you can ask me about it on one of the episodes we shoot later. [laughs] Yeah.
How do you use NotebookLM? I don't use NotebookLM. It, it just hasn't served me, um, high utility yet. You will when you're researching this.
Just watch Yann LeCun's... Okay, the best thing to do is to watch some- [laughs]... of his YouTube stuff and go to the Gemini summary. Yeah, that's also helpful.
Yeah. Last one about the Demis thing that it came through that I thought was interesting and that also has turned out to be really, um, ironic in the last 24 hours is he talked about how, um, they were using, uh, Genie, the, uh, generative- Yeah... world model, uh, you know, that we talked about where you can spin up an, uh, uh, a world and navigate through it. Um, they're using that to create worlds to test the Waymos in to, with like super fringe case studies to see how they would react.
Um, one example was, you know, if the Waymo's driving in a forest and like a forest fire breaks out and it's surrounded by flames, like what would it do? How would it behave? And he was talking about these like one in a billion like, uh, fringe scenarios that just like they're not gonna like, um, you know, ha- that's not gonna usually happen on regular training data with- in the real world. Uh, I was like, "Huh, okay, interesting."
[laughs] Fast, fast-forward to- No... yesterday. Have you seen this? The Waymos in Atlanta have [laughs] been going full send into flooded streets where they have now killed, or they have, they have now stopped, uh, Waymos in Atlanta in flooded streets.
Okay. For a second I thought that was an AI-generated image. No, that's terrifying if it's real. No, this is real.
The Waymos have been driving into flooded streets in Atlanta, [laughs] so they need to spin up- Right... some more Genie models- 100%... to test out [laughs] the fringe cases- Yeah... 'cause it happened.
Wow. And they drove through it. Sometimes I feel bad for the Waymos, but then they're machines. Who cares?
[laughs] I did take my first, uh, highway Waymo in, uh, at the event. Yeah, I, I texted you about my first- Oh, fun. Fun... uh, Waymo fight.
So I was on, I was in Santa Monica, Joey's neighborhood, and, uh, if you don't know, Santa Monica is like the unofficial capital of Waymo in LA. And, um, there was a Waymo next to me, and I was like, "What happens if I mess with it?" And just so you know, this was for educational purposes only. I don't recommend that you actually mess with Waymos.
Disclaimer. Okay. So I took my car, and I just lightly started to go into its lane. And at first it kinda just slowed down, and then I just was like, "Huh, it's pretty polite."
And then it ca- caught up, and then this time I was a little bit more aggressive and I was just like, like just did a real quick jerky move, and it honked at me. [laughs] And I was offended. I was like, "Hey, don't honk at me." I've never- [laughs]...
seen a Waymo honk. But I totally deserved that honk. J- uh, but I did wanna check how responsive their, uh, driving system was, and it was quite responsive. Somebody brought this up, and I felt the same way, where like if I [laughs] am like crossing the street, I feel way f- more safe, like, or don't think as much if I'm jumping in front of a Waymo- Absolutely...
crossing the street than a human driver. The Waymo will most likely stop, whereas a human driver might be on the phone or whatever. Yeah. Yeah, distracted.
Yes. The Waymo's got a bunch of sensors. I was at a group... Yeah, I guess I should also probably disclose, like i- i- the, the trip was paid for by Google.
Uh, and it was, you know, a lot of the, it was like the builders group, so it was also a lot of kind of content creators and people that are messing around with the Google products. But I will say they were genuinely interested- Mm-hmm... in like getting- Right... user feedback and like making these products better.
Yeah, and they, they're, they're, they're, they're, they are aware of what happens online [laughs] and what, what people talk about, so they're always looking, you know, even if it's just bug fixes or things, but also just like what use cases there are and how can the models get better, and also explaining why they throw a bunch of stuff under Google Labs and then kill products sometimes because they're like, "Yeah, sometimes the products suck or they just don't take off," and, you know, just wanna see what works, which sometimes is good and sometimes makes-Figuring out which Google product to use for your [laughs] problem- Yeah, I mean, end of the day-...
uh, challenge... they have to run a business and the product has to have some type of, uh, revenue promise, right? So they can't just keep making R&D things happen all the time. Yeah.
And then also try to figure out new things and new tools and uses. Um, so yeah, other two quick things I wanna touch on that were interesting and relevant to our audience. Uh, one is Flow. So Google Flow, which is their web app- It's like an editing platform-ish...
that you can use to, uh, for... Yeah, for video creation. Yeah. It's like kind of their- Mm-hmm...
narrow version of like Free Pick, uh, for like video creation. The... A couple updates there. One was, uh, obviously the new Omni model's built into it.
They added a character system where you can create characters and then call them up. I was pegging them with questions of like, "Is there something in the model?" Or when you- when they eventually roll out the API, 'cause there isn't an API yet, um, will there be something that like treats characters differently as other reference images? And basically the answer I got was like, no, basically what Flow is doing is just sort of structuring the data a little bit differently under the hood, but there's no- Hmm...
special Omni model that it's using. Um, so it's basically like anyone could kinda... They just built a clever interface and have some- Right... stuff helping with character consistency under the hood.
Yeah. There is no, uh, grand master plan to tie everything underneath with, uh, like a master workflow. Uh, basically there's no special model that Flow is using- Right... that anyone else- Right...
wouldn't have access to via the API. It's just a regular model just with like, uh, some extra stuff built on top to, to direct it. The master plan thing though that you mentioned reminded me we never finished our Avatar talk. The thing with Avatar and comparing it to Sora and Sora characters- Mm-hmm...
you can only make one avatar of yourself. Uh, you can't make avatars of other people. And I asked, I'm like, "Well, is there a plan where you could, you know, if you give me permission, like with Addy, can I pull up Addy's avatar if he gives me permission, and like make videos like we did with Sora?" Where it, it was a permission process and if you're okay that people could use your avatar in their own creations.
Um- Yeah. And they said, you know, maybe, but it wasn't on the roadmap right now. So I don't really know what... I'm like, if that wasn't on the roadmap, then what are you gonna do with avatars if it's just yourself and you can't...
I can't share it with anyone- I don't know... I can't like get access to anyone else's... I just think it has huge YouTube implications. Like, um- Yeah.
No, that's a good point... a lot of the faceless YouTube channels- Yeah... that make a ton of money, if they added a, a synthetic face, they would make more money. Stuff like that, you know?
Yeah. Uh, you know, actually, Rob, you're... That's, that's actually not bad. That's probably a, a direction to- And, and with the quality level where it's at, like I don't think they really care that it's not cinematic and it's not meant for our world.
I think if it's good enough for like a talking head YouTube video, it's more than good enough for them. Yeah. I mean, I'm sure they'll improve the quality on it. Um, but yeah, I mean, the use cases of, of us- Yeah...
of like needing- We need [laughs] we need aging-... more cinematic... de-aging, we need costume, hair, makeup- 10 bits to get 4K... Yeah.
Like our, our needs are up here, man. Like no- nobody's doing that anytime soon. Yeah. Okay, so back to Flow.
Uh, cool thing they added, which, you know, uh, obviously seems to be a trend with a lot of these tools, is, uh, an agentic workflow. So there's sort of like a Gemini sidebar and so you can just chat with it and be like, "Hey, generate 10 shots of like a man walking in the forest." Or if you have a bunch of shots, the example they gave here of like a daytime scene and you could say, "Okay, change all these shots to nighttime," and it'll just batch process and generate these new shots just through this- That's pretty impressive...
uh, basic chat interaction. I love stuff like that. Um, gosh, I'd hate to call it like, uh, agentic. It's probably not, but it's more like assisting you in the creativity process, you know?
It's agentic in the sense that like it is trying to figure out, like if I just ask something vague like, "I want a..." Um, I want... I, I think I tested it. I said, "I want like 10 different shots- Yeah...
of the same person walking in the same field," and I just gave it the text. It spun up and generated an image of the guy, an image of the, the, the field- Right... and then started making the shots using those images as inputs for consistency, and I just told it the one thing. So it had that agentic understanding of like, okay, to make something consistent for videos, I need to have consistent inputs.
Yeah. I gotta make the inputs first 'cause they don't exist. Also, like, like a real creative- Spin that up first... would never just go make that s-storyboard, right, all those shots, and then just go right into production.
Like, that's their iterative cycle, right? They'll start there, they'll modify shot three, modify shot seven, keep going, and I think that's- Yeah... still so much faster and gives them so much more options than like every little storyboard from scratch. Oh, yeah, exactly.
Yeah. If you're just like- Yeah... "I'm gonna copy and paste this prompt and redo it, copy and paste this prompt, redo it." Exactly.
Uh, it's like, oh, no, if you just wanna batch change something in a shots you already have, just tell it and, and then have it do it. Obviously, it's still charging you credits for every th- single thing you do, so- Right... [laughs] you will eat through your credits faster. But, um, you know, I think, I think this sort of agentic workflow where you're talking to stuff and it's doing and setting up things for you- Yeah...
is where a lot of these tools keep going. And then the other thing in Flow, and I'm curious to see how people use this, is... And this was sort of like a trend with Google I/O in general, was building more, uh, vibe coding and tool at... tool building capabilities into more tools.
So Flow has this creative tool builder inside Flow that is basically like a very white version of a vibe coding tool set where you just kinda describe what tool you want- Whoa... and it sort of spins up these- That's amazing... applets. Yeah Some of the examples were- That's powerful, man...
like a motion tracking one. Yeah. And can you download other people's tools that have been previously created? That's great.
You can remix. Yeah. There's like a... There's already like a, like a directory of like, um, publicly created ones that you can- Like-...
copy and remix and make your own... this was my idea for iOS. Like Apple should have rolled out something like this where it's like vibe coding for idiots.In iOS, so you can build little apps that only- Yeah...
exist on your phone. Yeah. Something, yeah, spins up a widget really fast. Yeah.
Um, yeah, for, yeah, this is like that idea, but it's like lives on your computer, in flow. Um, I'm just curious because it's like, you know, it's a- [laughs] Would this turn into the gateway drug- Absolutely... of like you mess with something here- Oh, yeah... and then you start going into- Yeah, you're, you're spending $3,000-...
tools along Cloud Code... on Cloud Code the next day. Yes. Yeah.
Yes. Anti-gravity, uh, uh, Google's version of Cloud Code are, are creatives thinking, uh, that way, where it's like, "I want a tool that does something that solves a specific problem?" Yeah. It, it's, it's a, it's a obvious idea now that we say it, but for them to actually go through with it, I'm sure there's a ton of engineering under the hood that's happening for user-generated app- application on the fly.
So that, that's a, that's a first, um, especially on the creative side of things. Yeah. Uh, okay. And then this one's about, uh, be my last one, but I...
This one I think was the most impressive or coolest surprise thing that I saw, and it's called Google Pix, which is a confusing name because there is also Google Photos. But, um, [laughs] Google Pix, it's basically, it's turned Nana Banana into like a Canva-like editor. Uh, so you can generate an image, you can generate a flyer, you know, with Nana Banana that has like text, and images, and design elements. And normally if you did that and then you wanted to change something, you'd have to like redo the whole thing.
And then it served us as like, again, this is one of those things which is like they're doing something under the hood. It's not a special model, but I'm not quite sure how they figured out how to like in paint in Nana Banana- Yeah... to blend the edges. So it turns something flat into many, many layers, then, then you can individually edit.
Yeah. This was like a, not the best recorded demo, but at the booth, basically this, from scratch, it generated this rooted future flyer thing. Mm-hmm. And then you can see the mouse hovering over every element, and it turns every element into something I can click, and then I can re-prompt to say what I want to change about it, and then it'll just edit that one section and not touch anything else.
But the outputs are like- Yeah. That's wild... so cohesive. It was kind of wild.
Um, this was another test one I did where I, I had already prompted it. I think I said, I clicked on the text and I said, "Make the text fancier," and I clicked on the person on the left and I said, "Put them in like a, a spaceship, uh, in an astronaut outfit, and the person on the right in a, um, [tsking] I don't even know, safari outfit," and then here's the output. And so it changed, the text is the same, it changed the text. It kept the person's likeness, but it changed their outfit.
Didn't change anything about the layout. Yeah, it's still there. Uh, even that- Yeah... tape thing on the top, the tape's still there, it's still blended in, but you know, it changed the, uh, underlay- Yeah...
of the wording underneath it. Bro, dude, graphics design is getting so easy. I mean, like if you compare or if you complement this with GPT-2, which is really good at gaph- graphic design, so you take, you know, your first pass at GPT-2, you generate the thing that you want and you're like, "Now I wanna kind of change it," you bring it into Google Pix and you're doing layer by layer and element by element adjustment- Mm-hmm... without- I don't know [laughs] Oh.
But now I'm curious, uh, if they'll let you- Yeah. I hope so... import an image to edit. Yeah.
I guess they would [laughs] if they're like, "You gotta- Also-... you gotta make Nana Banana first"... calling it Pix is like the complete opposite of what it's doing. [laughs] Has nothing to do with real photos per se.
Yeah. Uh, yeah. [sighs] Look, man, they're good, they're- Yeah... they're good at making stuff.
Yeah. Sometimes they're not the best at naming stuff. [laughs] Like, I don't know. Yeah.
No, that's, that's amazing. Look, we've covered IO in the past, but this feels like a really eventful one and, uh, a slightly less evil one if we're pivoting into science. Yeah. I also didn't touch, I know people are upset 'cause they kind of also announced that they're basically shifting main Google search to like AI-centric search first.
Separate topic, but- I mean, how are they gonna make their money? Is SEO dead? Uh, maybe. All that comes from AdSense, right?
They probably already thought about all this. True. Moving on. Yeah.
Runway Aleph 2. Runway Aleph was- Videos... Runway's model that was like- Yeah... one of the original edit video, video to video, give it a video, just tell it what you wanna change.
It was okay. How was your experience with that? It was okay. I thought the, um, I thought the quality lacked, but the promise of Aleph wa- And Aleph was one of the first, if not the first video to video model, predated Kling and some of the newer ones, and I thought, yeah, this is absolutely the direction we should go.
We should not have to generate everything from scratch. We could just generate something that's roughly there and then iterate, and iterate, and iterate on it. Um, some of the generations that I ran had, um, like high frequency noise and like blobs and things like that, which I thought, "Hey, it's just the first generation, it's gonna get better." But the promise and the vision for Runway, what I thought was really, really strong.
Uh, yeah, similar boat. I've like, uh, I've always had a hard time getting good outputs out of Runway. It's just never really like vibed for me. Aleph, everything I tested, it would just either change too much about the video, warp too much stuff, too much would be soft or fuzzy.
It was just never really usable. Yeah. Aleph 2 definitely feels like a huge bump in quality. I'm not sure what the resolution is, but just even from like lighting, composition, you know, the, the tone map looks, a lot of that is subdued and just looks more and more natural.
One of the improvements is you can get more specific about what you wanna change. So it has a slider where before you kind of had to describe your change or you can kind of change the first frame, but this slider lets you pick any frame in the video and then modify that frame. So you can give it a better starting point, uh, and guidance of like what you wanna change. And then you can upload images, or you could also, um, change this frame and describe it with what you wanna change.
Oh, so it's like a- Um, and then create the first frame with Nana Banana. Like an image reference insertion-On the frame that you'd want Yeah. And obviously you could just download this frame and then like- Yeah, but that, that's like round tripping it... use Photoshop or whatever to- You don't want to be doing that.
Yeah... really dial in whatever you wanna change. Yeah. But yeah, at least you have it built in here.
That's cool, man. You can modify it here- Yeah... and then change it. So that's the new workflow.
Uh, I did the same test- The LACMA thing... where I gave it that, the video of the, um- Yeah... Metropolis sculpture. First image it generated from the first frame kept too much- Yeah, yeah, yeah [laughs]...
of the LACMA railing and balcony, so I didn't use this frame. [laughs] Uh, I ended up just- Okay... getting one where the camera tilted down, and used this bottom frame as the guidance. Oof.
And this was the output. Uh, yeah. It's so much softer- It's really noisy... and lacking so much detail than, uh, what we saw with Omni.
A lot. Yeah. And then this was the same. Let me change the- Yep...
shot to a dog, or change the dog to a robot. This was the robot that it came up with, and I was like, "Sure, cool. Works." No.
And then this is the output. Oh, the leg disappeared for a frame or two. And you can see the dog with the legs warping. It's still moving like your dog- The face is, like, distorted and-...
but not a robot dog, whereas the Omni model was really the mechanics and the rigging was like a robot dog. Yeah, th- this, this is like they're feeding it through OpenPose and just extracting the, the anchor points, and then just attaching it to the new dog, whereas the Omni model was really doing something different. It fundamentally translated that animation into a completely robotic animation. Yeah, so you see here, like, those are backward bending legs in the front.
Like, that's hard to do anyway. And yeah, like, moving like that, that's how a Boston Dynamic robot spins- Mm... like, on all fours, but not a real dog probably won't do that. So from an animation quality standpoint, Omni nailed it.
Aleph, not so much here. Having said that, like, it is still really successful in painting, 'cause you could see, like, her hand on the leash, and just overall crowd work and all that. Like, none, none of that changes. No, I mean, at least it's not messing with anything else.
Yeah, th- the shadow swims a little bit. You know, the, the left paw, the right paw just kinda disappears. It's, it's tough, man, 'cause Runway's probably been working on this for, I don't know, six months to a year. They release at the same time as Omni.
Obviously Google is way more funded and has a bigger team, and they're gonna release a more superior product. Unfortunate thing is they are coming out at the same time, and we're gonna compare the two against each other. I mean, it felt like this is probably a thing where they w- you know, it's like, when do you release it? And it's like, okay, well, like, Omni's coming out and, like, Taotian doing- Yeah...
the same thing. So, like, [laughs] we better drop this update. You know, if you have success with Aleph, let me know. I'm always curious, 'cause I, I know some people like it and have used it, and I've just never really- I'll just say one shout-out to Runway, that the f- the fact that being, like, a small company, they're not a Google or an OpenAI or, you know, they're still hanging with the big boys, right?
Yeah, yeah. So Kristof is doing something right. All right. Links to everything we talked about at denoizpodcast.
com. If you'd like to meet us in person next week, we'll be at AI in the Lot. Joey's gonna be a little bit busier than I am, so I'll be walking the show floor. Uh, come say hi if you see me.
Yeah, I'll be buried in a back room. But let me know if you're gonna be there. I'll be around at the, uh, after party stuff. Thanks, everyone.
We'll catch you in the next episode.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.