The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/HR/Hidden Layers
Hidden Layers artwork

GPT-5 Release Fallout, AGI Timeline, Google's Genie 3 and Meta's DINO V3 | EP. 45

Hidden Layers · 2025-09-03 · 25 min

0:00--:--

Key moments - from our scoring

Substance score

64 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality11 / 20
Guest Caliber14 / 20
Specificity & Evidence14 / 20
Conversational Craft12 / 20

The episode unpacks OpenAI's GPT-5 release, which despite solid benchmark improvements - MMLU jumping from 56% to 84%, AIME from 7% to 87% - disappointed many due to a botched rollout, router failures, and removal of older models like GPT-4o. The hosts debate whether GPT-5 represents meaningful progress or signals stalling returns without major algorithmic breakthroughs. This connects to a broader AGI timeline discussion: while Ron argues we've already crossed the AGI threshold (meeting criteria from a decade ago), Dr. CZSI emphasizes embodiment and physical world interaction remain unsolved, and Michael notes the wall around unsupervised agent autonomy. The conversation then pivots to more optimistic territory - Google DeepMind's Genie 3, an action-conditioned video generation model generating 24fps at 720p with emergent spatial consistency, and Tencent's open-source Hunyuan Gamecraft. Michael closes with Meta's DINO V3, a self-supervised image encoder solving a known limitation in dense feature extraction through gram anchoring, maintaining fine-grained patch consistency across training iterations.

Key takeaways

  • →GPT-5 shows meaningful benchmark jumps (MMLU 56→84%, AIME 7→87%) but the release was operationally botched with broken routers and model confusion, reducing the perceived impact.
  • →Genie 3's action-conditioned video generation achieves real-time 24fps at 720p with emergent spatial and temporal consistency without explicit 3D geometry - a major step toward embodied AI through world models.
  • →DINO V3 solves the dense feature degradation problem in self-supervised encoders by maintaining Gram matrix consistency across training, enabling semantic segmentation without fine-tuning.
  • →The AGI timeline debate hinges on definition: Ron contends we've met pre-2015 AGI criteria, while Dr. CZSI and Michael tie remaining gaps to embodiment, physics understanding, and unsupervised agent reliability.
  • →Inference cost reduction and model routing optimization appear to be OpenAI's actual focus with GPT-5, traded against headline capability improvements.

In this episode

  1. 1GPT-5 Release Analysis and Performance Metrics
  2. 2AGI Timeline and Definition Debate
  3. 3Google's Genie 3 World Model Breakthrough
  4. 4Tencent's Hunyuan Gamecraft Action-Conditioned Video Generation
  5. 5Meta's DINO V3 Self-Supervised Image Encoder Innovation

Mentioned

OpenAIGoogle DeepMindMetaAnthropicTencentKung Fu AIGPT-5ClaudeGenie 3DINO V3Hunyuan GamecraftRon Green

Guests

Michael WardenDr. CZSI

Topics in this episode

World modelsGPT-5Genie 3DINO V3Hunyuan Gamecraftaction-conditioned video generationself-supervised image encodersgram anchoringMMLU benchmarkAIME benchmark

Questions this episode answers

How much better is GPT-5 compared to GPT-4 on benchmark tasks?

GPT-5 shows significant jumps: MMLU (general knowledge) improved from 56% to 84%, software engineering from 1% to 65%, and AIME (high school math) from 7% to 87%. On the meter challenge (long-horizon tasks), GPT-5 handles up to 2 hours 15 minutes compared to GPT-4's 10 minutes.

What went wrong with OpenAI's GPT-5 release?

The release suffered from router failures that routed complex questions to fast thinking models instead of the thinking model, removal of older models like GPT-4o that some users relied on, and general confusion around model selection - operations issues rather than core capability problems.

What is Genie 3 and why is it significant?

Genie 3 is Google DeepMind's action-conditioned video generation model that generates interactive 3D worlds in real-time (24fps at 720p) based on text prompts and user controls, with emergent spatial and temporal consistency achieved purely from pixel-level learning without explicit geometry encoding.

How does DINO V3 solve the dense feature problem in image encoders?

DINO V3 uses gram anchoring - maintaining cosine similarity consistency of Gram matrices (pairwise feature comparisons) between a teacher model from earlier training and the student model - preventing fine-grained features from degrading as training progresses.

Have we already achieved AGI?

Ron argues yes, claiming current LLMs meet AGI definitions from 10 years ago (broad general intelligence, expert-level conversation), but Dr. CZSI and Michael contend AGI requires embodiment and reliable unsupervised autonomy in physical worlds, which remain unsolved.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode packs several substantive technical details - GPT-5 performance metrics (MMLU 56→84%, AIME 7→87%, 2hr 15min task horizon), Genie 3's 24fps 720p video generation, and DINO V3's gram anchoring technique - but is diluted by meandering AGI definitional debate that circles without resolution. Useful for someone unfamiliar with these model releases, but lacks deep operational insight for B2B operators.

On the multiple choice general knowledge, MMLU, it went from 56 to 84%. This is GPT four versus five.
GPT five is now the top model there. And this is really impressive. It's now, it has a 50% time horizon on tasks up to two hours and 15 minutes.

Originality

11 / 20

The take that OpenAI deliberately optimized for inference cost reduction over raw capability is a useful counternarrative to pure disappointment narratives, and the embodiment/physical intelligence framing of AGI is well-articulated. However, the core analysis largely synthesizes existing public discourse (Gary Marcus review, scaling debates, world models as AGI prerequisite) without genuinely novel frameworks or contrarian positions beyond the 'we've already achieved AGI' claim, which is asserted rather than rigorously defended.

I think that OpenAI has not come out and explicitly said this, but it seems pretty clear to me that they were focused a lot on reducing inference costs.
I genuinely feel like we've already achieved AGI. And I believe that based upon what I would have defined AGI as 10 years ago, four years ago. That's it. I think we're just moving the goalposts.

Guest Caliber

14 / 20

Guests are credentialed practitioners (Dr. CZSI, Michael Warden as VP of Engineering at Kung Fu AI, host Ron Green as co-founder), speaking from direct hands-on experience with cutting-edge models. However, they are not household names or widely-recognized industry leaders with massive scale exits or dominant market positions, and their contributions, while informed, remain largely reactive commentary on others' published work rather than original operational insights.

I'm joined by my two amazing Kung Fu AI colleagues, my co-founder and distinguished engineer, Dr. CZSI, and VP of engineering, Michael Warden.
So I've used the GPT basically every day. And between Claude and the GPT five, I kind of use some half and half.

Specificity & Evidence

14 / 20

Strong on model metrics and technical details (MMLU, AIME, DINO V3's gram anchoring, Genie 3's demo specifics) but weak on business impact, customer outcomes, or deployment scale. Concrete examples of model behavior (zebra legs, fruit market segmentation) are present, but the episode lacks data on adoption, revenue, competitive wins, or operator-level ROI that would ground these capabilities in real business context.

on the AIME, the high school level math, it went from 7% to 87%.
they trained out about one million gameplay videos from about 100 games.

Conversational Craft

12 / 20

The host asks reasonable follow-ups and pushes back on AGI claims (e.g., 'did this push out your timeline?'), and guests do challenge each other's definitions. However, many questions are surface-level or rhetorical ('are you using Claude for your non-coding tasks as well?'), and when the AGI debate heats up, the hosts allow circular reasoning and goalpost-moving to persist without sharp interrogation of the contradictions or empirical standards underlying the claims.

Did that push out your timeline on AGI or not affected?
Do we review human code? Yeah. Yeah. That's a good point.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

model24world16feel15intelligence12five12tasks12features11video10genie9fine8consistency8meta7general7models7real7performance7

Episode notes

In this episode of Hidden Layers, we dive into the most important AI developments of the month. We cover OpenAI’s highly anticipated and controversial GPT-5 release, debate where we really are on the AGI timeline, explore groundbreaking new world models like Google’s Genie 3 and Tencent’s Huanyuan Gamecraft, and unpack Meta’s DINO V3 image encoder breakthrough.

Full transcript

25 min

Transcribed and scored by The B2B Podcast Index.

Welcome to Hidden Layers. We're exploring the people and technology behind artificial intelligence. I'm your host Ron Green. Today we're covering big news out of meta, open AI, and Google.

And we're also going to share our thoughts on AGI and exactly where we each think we are on the timeline to achieving artificial general intelligence. Tell me make sense of it all. I'm joined by my two amazing Kung Fu AI colleagues, my co-founder and distinguished engineer, Dr. CZSI, and VP of engineering, Michael Warden.

All right, great to see you guys. You ready for now that I've said? So great. All right, let's just jump right in.

I'm going to talk about what I think is probably the most anticipated model of the year, which is GPT-5. It's been almost two and a half years since GPT-4 came out. This was definitely... I think this is one of those situations where open AI put themselves in a little bit of a corner because there was so much hype almost anything they came out with was going to be a little bit of a letdown.

A little bit of a letdown. I was shown as a big whale compared to GPT-4. Exactly. That's exactly right.

And so I've spent, you know, it came out a couple of weeks ago. I think the timing on this episode's great because it allowed me to do a bunch of research. A few things. I think there's a general consensus even two weeks out that the release was a bit botched.

It wasn't so much that the model was a huge disappointment, but there's an element to that. It was just a lot of the confusion. You know, they removed some of the pre-existing models. Did you guys see that there were people, like especially in some, you know, user groups that were really upset about 4-0 being a 4-0.

And you know, it's a very almost like sycophantic model. You know, everything's a great idea that you have. And I saw some people feeling saying, you know, they felt like they lost a friend. You know, there were issues with the router.

One of the things I think opening out was trying to accomplish with this releases moving away from having users select the model. That's not a great choice. You know, like, you know, O3 and Turbo and many and things like that. So the router, which was supposed to take in your question to decide whether to do a fast response model or a thinking model, that was broken.

So a lot of the questions that needed more thought went to the fast thinking model. There's some other things too. It was pretty underwhelming, I think, from a capability perspective compared to what people expected five to be given all the time. And that's led to a bunch of what I kind of feel like real sort of hype backlash.

And it's gone beyond just GPT five in my opinion. And I actually want to ask you guys about this in a second. There's this idea that open AI has lost the lead or they failed or the model wasn't. Everything was expected to be maybe AGI is now the timeline on that's been pushed away.

Maybe scaling doesn't work anymore or the AI bubble is bursting all these sorts of really negative trends. So before I go any further, I want to ask you, have you guys had a chance to play with GPT five much yet? And if so, what are your thoughts? I personally am a big anthropic guy.

I mean, just because I'm a developer, I typically is that for most tasks. So I've looked at a lot of examples. I've obviously read Gary Marcus's scathing review. But I haven't really played with it much other than just a few little test prompts.

Are you using Claude for your non-coding tasks as well? Yeah, I just have the desktop app and use it for everything. Yeah, how about you easy? So I've used the GPT basically every day.

And between Claude and the GPT five, I kind of use some half and half. Claude is really good at coding. I use Claude code. And the Opus model is really good.

And Sonnet is not bad either. For GPT five, I totally agree. I feel like for common simple tasks, I don't feel there's a big difference compared to 4.0.

Yeah, I agree. And Opus 3 is really, really good. Right. Yeah, I know the naming can be really confusing.

I know. Then GPT five pro, if you have the patience, it is really, really good on some of the more deep thinking tasks. So if you're trying to brainstorm about some mathematical theories, some proofs, and trying to do some survey, and I do find it to be giving me some good insights. And it's a great learning tool for me now.

Yeah, that mirrors my experience as well. I feel like when the day-to-day stuff, not a big change if I'm asking you to summarize something or like that, but when it kicks into thinking mode, there has also been great. And I'm finding that they advertise that there were less hallucinations. I'm feeling like there are less hallucinations on a personal basis.

All right. So I went and did an analysis of its performance on different metrics. And it's pretty interesting. Surprisingly, the jump in performance between April 2024, which was GPT four turbo, and April 2025 to GPT three, that was bigger than the jump between three and four, which is kind of surprising.

I think we forget how bad three was relative to, and hindsight really saw how weak four was. A few metrics. On the multiple choice general knowledge, MMLU, it went from 56 to 84%. This is GPT four versus five.

So a big jump on the sweet software engineering went from basically 1% to 65%. And then on the AIME, the high school level math, it went from 7% to 87%. So these are pretty meaningful jumps. Probably more interestingly, the meter challenge essentially that measures the ability for AI to carry out long tasks.

GPT five is now the top model there. And this is really impressive. It's now, it has a 50% time horizon on tasks up to two hours and 15 minutes, which is really kind of astounding. And by contrast, the GPT three could only handle tasks of a few minutes, of a few seconds to a minute, and four could go up to 10 minutes.

So I mean, this is a significant jump in the last couple of years. So here's my takeaway on this. I think the reality is, it's O3 was so strong that I think it was just necessarily going to be let down with five. And I also think that OpenAI has not come out and explicitly said this, but it seems pretty clear to me that they were focused a lot on reducing inference costs.

This was about getting, getting, trying to at least keep parity for most average users, but dramatically reducing the inference costs. And there are other things like getting rid of model selection and having the router and stuff like that. And so I think that, I think net net, it was sort of a botched release, but it might, my guess is it's going to get the monkey off OpenAI's back. They've now figured out the naming, and now they are sort of clear to maybe take a big swing on the scale front again.

All right, so I've got a question for you guys. How does this affect your timeline on AGI, meaning the fact that GPT five, I think it's fair to say it's more evolutionary than revolutionary. Does that pull back, push forward your timeline on AGI? And I know there are some people that maybe even think we've really already crossed that threshold, and we should qualify these models as AGI.

I'd love to hear what each of you think. That's a great question. Yeah, I think different people have very different definitions of what AGI is. And for me, I always thought AGI is not near, because to me, an AGI needs to be able to complete a simple task for a human like, hey, make a coffee and bring it to the table.

Oh, so you're thinking like embodiment, you will be a physical embodiment, okay? Yeah, it needs to be able to interact with the real world. The language intelligence is great. The mass coding capability is great, but it's everything's in a digital.

It's confined in a digital world. It doesn't really deal with the messiness of the real world. And I view the language, mass coding ability as a tip of iceberg, kind of in an IV tower. But the thing that's underneath, the physical intelligence, the spatial intelligence, like grabbing a cup reliably, not breaking it, you know, you operate in the coffee machine and maybe it's dirty.

Maybe you need to wash it. So many messiness in the real world that I feel like. Okay, so you don't think we have AGI right now. Do you feel specifically with GPT-5?

Do you feel that that pushed out when you think we'll get to AGI or didn't really affect your timeline? I don't think it affected my time. I would have more interest in following the world model, the video generation type of work stream. I feel like that can unlock a lot of the things that's under the tip of iceberg.

Okay. Yeah, that's not true. Honestly, it's going to be pretty similar. In these AGI conversations, always tend to lean toward it.

Okay, let's define it. What are we talking about? I kind of suspected that there would be some sort of stalling performance, at least without some major algorithmic sort of upgrade. Yeah.

So, you know, it's interesting to watch it play out. I mean, there's the biases that these models have. Like, the intrinsic features of the way that Ellen's work seem to be consistent regarding, excuse me, regardless of scale so far. Like, if he's really strong priors and I know Gary Marcus is the go-to naysayer for all this stuff.

Yeah. And, you know, he shows these photos in his latest substack post that's like, oh, look, how many legs does this zebra have? And it's like a photo of the zebra. It's clearly got five legs.

But there's such a strong prior around how many legs a zebra should have that you just can't get it to admit that they're five legs. Right. There are other things that play there. But, you know, I really do think that the spirit behind the AGI question is, you know, can this do what humans do?

And I think that we've come a long way, for sure. And there's still a ton of distance that we need to go. And it's just kind of like a, I don't think there's an AGI winter ahead of us. Okay.

But you did it. All right. Same question for you. Did this push out your timeline for AGI or not affected?

I don't think affected because it feels kind of like in the noise in terms of how far we still have to go. Okay. Okay. All right.

This is great because I'm going to take a contrary in perspective. We should love it. I love to hear it. I genuinely feel like we've already achieved AGI.

And I believe that based upon what I would have defined AGI as 10 years ago, four years ago. That's it. I think we're just moving the goalposts. Yeah.

You know, if you go all the way back to touring. Yeah, touring test. You know, in the mid 20th century, you know, the whole idea was you would interact through text prompts because you didn't want the embodiment problem to be challenged. And, you know, again, seriously, 10 years ago, if you would have said, give me your criteria for AGI.

I think we'd have checked all those boxes, right? That's totally with you on that. And it's because these models now have broad general intelligence. You can interact with them and they can converse increasingly at an expert level on pretty much any topic.

Now, they'll make mistakes. But I actually think that's a bit of an element of the nature of intelligence. That sort of fractal boundary of intelligence. And if you think about humans, I mean, I like to think I'm a smart person.

I'll do the dumbest things occasionally. And I can't believe it. I will misspeak or I'll walk into the room and forget why I went in there. And I don't think that those lapses just qualify me as having general intelligence.

I know this is a little bit of a fringe take, but you know that. I think it's mostly we've moved the goals. And I think it's fine. This is a reason I'm pretty excited, which is, it takes easy.

You're like, well, you want AGI to have full embodiment and an ability to move around a 3D world. The fact that we're even moving the goalpost to that point now, show how much progress we've made. Totally agree. What about the, I mean, I guess this is get to broader conversations.

I'll try to keep this little brief. But I feel like with agents, except for the very, except for some very narrow tasks that are wrote where the domain is unchanging, you can't let these things just go run off in the wild and do their own thing. There's still always got to be some sort of human supervision. And to me, it feels like there's some key call it AGI threshold or something beyond which humans don't have to be involved in things that they normally are.

Yeah. And it feels like we really have kind of like bumped up into a little bit of a wall there. Yeah. I don't think there's an actual wall there.

We'll bust past it. Right. But there's something about still having to review every line of code that gets generated in like my IDE that feels like, you know, if we had AGI, I feel like I wouldn't need to do that. Right.

Right. Right. Right. But see, all right.

The reason the reason I feel comfortable saying this is we're talking about general intelligence. Yeah. Do we review human code? Yeah.

Yeah. That's a good point. That's a good point. We have to.

Yeah. No, maybe ASI. Right. Super intelligence.

No. That's a new goalpost. That's the new goalpost. She's just handling it.

Okay. Zizi, I know you've got some news on the world of model front for us. Yeah. So I got two quick updates that I am really excited about in the last month.

Both are about world models. Well, really action-conditioned video generation models. You know, the Genie 3 that came from Google DeepMind. Super impressive.

Yeah. I watched their demos. That's really great. I don't think they have APIs accessible yet, but the demos definitely look super impressive.

So it's an action-conditioned video generation model. Basically, you can think about it as, you know, you can write a prompt and it will generate a role where you can just use your arrow keys or mouse to move around and then explore the world and in real time because they, I believe they optimize their model to be, you know, generating at 24 frames per second now at 720p kind of resolution. So the video, and we'll put a link in the podcast, but the video where they showed the perspective of the man painting the wall.

You know, I'm talking about it and it has memory. I've shown that video to, I don't know, half dozen people. And every single person that saw it asked me about halfway through. Hold on, is this, is this how I generated?

They thought it was real. Yeah. And yeah, there's another demo that I really like, which is, you know, you can control the player kind of, you know, in the first person view and go towards a water puddle. And then you move forward and then you look down and then you can see the footstep into the water puddle.

And then the ripple and everything that looks like so real. Yeah. Another crazy example is where it's like a world in a world. So I think someone prompted Genie 3 to imagine Genie 3.

So in that video, you know, there's a person looking at a computer screen. And on that computer screen is Genie 3, you know, that's incredible. Generating there. So yeah, that is super interesting.

And another adjacent update is called Hunyuan Gamecraft. It's also a world model. It's an action-conditioned world model. It's an open source one from Tencent.

So what's interesting there is that, you know, they made some architectural innovation. Where they basically map the mouse movement or the arrow movement into kind of camera and movement space. And then they also did something with the kind of the hybrid history so that they can keep generating a long video without drift, which has been a challenge before. Yeah.

Because I believe the Genie 3 video, is there a lot of minute, two minutes max, something like that? Yeah, yeah. So between Genie 2 and Genie 3, I think that lends up the video without drifting, without, you know, having unfeasible kind of conditions has improved a lot. And I think they trend, well, for Hunyuan Gamecraft, they trend out about one million gameplay videos from about 100 games.

Yeah, that's a super interesting model. So there are no tricks under the hood to enforce consistency or anything like that. It's truly just kind of like an end-to-end, data-end model out kind of situation. Yeah, I believe they don't have explicit geometry, 3D geometry.

Like everything is just learned from the pixels. And I remember the same thing on the Google model, the Genie 3, that they said that the consistency was emergent, which to me is just amazing. I mean, you know, that is, that's when things like that happen, it makes me feel like we're on the right track. That's a bit or less than, right?

That's a bit or less than, baby. And related to the AJA topic, you know, if we want the, yeah, to have embodiment, to really understand the world, well, you need to simulate the world first. Well, that's not my worst ass Elon Musk words. That's great.

That's awesome. Michael, what do you got? Yeah, I was going to talk a little bit about dyno V3. Good, good.

So yeah, it's from meta. And I feel like with, with meta, it's finally across that threshold where it's more comfortable to say meta than Facebook. I feel like the same thing with like the Twitter X thing. Yeah, they're meta now.

Good job. But yeah, so it's a self-supervised image encoder, you know, meant to be very, very general purpose, much like Clip, which we use for image embeddings for tons and tons of things. But they had a clever little innovation that I think was kind of like the secret to unlocking a little bit of performance with dense feature performance. So for things like really fine grain image segmentation or, you know, things where you need really dense features.

Apparently, it's this really big deal with self-supervised imaging coders where if you train them for a really long time, you'll notice that if you pluck out a set of weights in the middle of training, and then you fine tune it for either a classification task or something that requires dense features. Classification performance just improves monotonically. But eventually the semantic segmentation or those like dense performance sorts of tasks will peak out and then drop off and start degrading.

I want to make sure I understand this will happen when you're fine tuning it. Period or if you remove some weight. It's basically like, like let's say you train for 100,000 iterations, but you've still got some way to go on the self-supervised task. Yeah.

If you use that as a basis for a fine tuned task, it's going to do great with classification, great with segmentation. But the more you train those features, the more they're useful for global tasks and not local tasks. Interesting. Okay.

So like are the features that are being extracted useful for everything? Yeah. It sounds like they are. They're losing the very fine crane features somehow.

That's exactly it. Through. Yeah. So apparently this is like in the self-supervised image and coder world, this is like a very known limitation.

And what they did to get around that was this thing called gram anchoring. No, I don't know. It's okay. So the idea is that this gram matrix is what they're calling.

You take the feature map from one model's output and then you do pairwise comparisons of all the features. So it's an in-squared sort of problem. Like if you can imagine a vision encoder and you have the 16 by 16 feature map or for the patches or whatever your architecture is, if you were to do pairwise comparisons between all those features, that is very... If you enforce it to keep that consistency as your training, then those dense features don't degrade.

Oh, okay. Sounds... I know it's kind of convoluted. And when you say you're comparing those patches for consistency, I mean, is it like a dot product, you know, specifically what they're doing?

Literally just like cosine similarity. Okay, yeah. Okay, okay. Yep.

And that's all it takes for it to not... I sort of average out that those find great features. Yeah, they'll literally pluck out, like as they're training, they'll just take one of the earlier models and that becomes the teacher. And then as they continue training, those downstream become the student.

And then they just enforce consistency relative to a previous teacher. Oh, okay. I miss the progress. Okay.

I was thinking within the same model, this is over time, you keep the consistency over time. Yeah. Oh, interesting. Yeah, it almost feels like, you know, when you project an image into like 16 by 16 kind of latent vectors and in a latent space, they have certain geometry, certain relative position.

Yeah. Maybe over time or during the progress of training, you don't want them to be just randomly shuffle around. You want to keep it kind of invariant, like relative, their relative position in a latent space to be consistent. Yeah, I can show that.

Yeah, you want the structure. It's almost like that platonic representation hypothesis paper that we've all been talking about. Yeah. Like there's some structural consistency of that embedding space that you want to maintain roughly at the point that it starts.

It starts doing really well on those semantic segmentation or dense tasks. Yeah. And the results are kind of mind blowing. You know, they tested a bunch against a bunch of different things because it's an encoder that's universal.

So yeah, I'm not going to get too into the weeds at the evaluations. But the takeaway is that anytime dense features are important, it blows the competition out of the water. So things like semantic segmentation. You know, there's no supervised training with this whatsoever.

And there's a figure. I think it's figure three. And it shows if you look at the feature map on this, it's like a really, really crazy photo of like a fruit market. Like a fruit standard market.

It's got apples, bananas, oranges, cantaloupe, whatever. And if you look at the, you just pluck one pixel out. And then you take similarity with all the rest of them. It's basically got a semantic segmentation model built in already without any sort of fine tuning.

Because it's that good at that structural consistency. Oh, wow. Okay. I did not see that coming.

Yeah. You can use it for really amazing sort of segmentation without any sort of fine tuning. Wow. And this is open source.

And it's, well, yeah, yeah. With an asterisk. Yeah. I mean, you have to give them your name and your data birth.

Yeah. It's kind of like a scandal. Yeah. It's meta.

Yeah. Yeah. That's great. All right.

Well, I think we'll call it there. Because this was, this was a really, for me, I think a pretty exciting month. And I think we covered the three most important things that have happened. GPT-5, the Genie-3 model, which I think is actually just groundbreaking in an amazing way.

And then the stuff coming out of meta with done at three. This is great. All right, guys. I appreciate your time.

Thank you so much. See you next time. Thank you for listening to Hidden Layers. This series is hosted by Kung Fu AI, a management consulting and engineering firm focused exclusively on artificial intelligence.

If you have any questions or thoughts about today's episode, or if you know someone we should feature, please visit us at kungfu.ai.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Open Source Self-Driving with Comma AIPractical AI · on World models91 / 100
  • Kiwi co-founded world‑builder hits $2.5 billion valuationThe Business of Tech · on World models87 / 100
  • Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI TodayUnsupervised Learning with Jacob Effron · on World models85 / 100
  • The Future of Tech - Key Themes for 2026 and BeyondThe B2B Podcast · on World models82 / 100
  • Where AI Delivers Real Value in Pharma Today with Emily LewisPharma Sessions · on World models80 / 100
  • Tech Shift Survival Guide with Aparna Sinha, SVP of Product at VercelNext Gen Builders · on GPT-570 / 100

More from Hidden Layers

All episodes →
  • AI Is Designing the Next Cancer Fighter | EP.5383 / 100
  • Anthropic Code Leak: A Rare Look Inside Frontier AI | EP.5282 / 100
  • The "AI Bubble" Bubble | EP.5174 / 100
  • Did AI Kill Programming? | EP. 5072 / 100
  • Your AI Is Too Big, Too Expensive, and Probably Wrong | EP. 4978 / 100
Explore the best B2B HR podcasts →
All Hidden Layers episodes →