The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The Gradient: Perspectives on AI
The Gradient: Perspectives on AI artwork

2024 in AI, with Nathan Benaich

The Gradient: Perspectives on AI · 2024-12-26 · 1h 49m

0:00--:--

Nathan Benaich, founder and general partner of Airstreet Capital, joins to reflect on 2024's pivotal moments in AI. The conversation opens with O3's surprising performance on complex reasoning benchmarks (ARC-AGI, math problems, PhD-level research questions), though Benaich cautions that limited access and high inference costs - up to $1,000 per problem - make evaluation difficult. The hosts explore the economic implications of expensive frontier models and the likely multi-year gap between accessible systems like Llama and DeepSeek R1 versus prohibitively costly cutting-edge models. Benaich frames this through infrastructure analogies: just as you don't use a Mac Pro for everyday tasks, frontier models may eventually become cheap enough to displace cheaper alternatives entirely, but we're currently in "blitz scaling wars" prioritizing performance over margins. The conversation shifts to Airstreet's successful 2024 - marked by increased interest in AI applications, major biology-focused Nobel Prizes (protein folding, deep learning), and surprising trend reversals in defense/security and robotics. Benaich emphasizes evaluating Gen AI founders on execution, design taste, and customer affinity rather than ML credibility, since frontier models now allow rapid iteration without massive compute budgets. Examples like Synthesia and ElevenLabs demonstrate how startups can outmaneuver incumbents through superior UX and market timing, contrasted against Apple Intelligence's underwhelming integration. The episode concludes examining the "bitter lesson" in biology - AlphaFold3's diffusion-based approach outperforming engineered inductive biases - while noting ESM3's success with multimodal constraints, suggesting different scaling strategies may suit different biological challenges.

Key takeaways

  • →O3 demonstrates reasoning progress but remains prohibitively expensive ($1,000+ inference per problem), and it's unclear whether scaling inference compute is the right long-term approach for commoditized AI applications.
  • →A significant capability gap will persist for years between frontier models (O3, advanced proprietary systems) and open-source alternatives (Llama, DeepSeek R1), creating distinct competitive tiers in the market.
  • →Gen AI startup success depends more on product execution, design taste, and domain expertise than ML engineering credentials, since founders can now iterate rapidly using frontier models without massive pre-training budgets.
  • →Startups like Synthesia and ElevenLabs have beaten incumbents by crossing the quality threshold first and building superior UX, while big tech companies (Meta, Apple) often hesitate to enter emotionally sensitive domains like digital avatars and voices.
  • →Biology research is experiencing transformative results by importing language model insights (scaling, DPO, alignment) and multi-modal learning approaches, with AlphaFold3 and ESM3 showing different benefits from diffusion-based versus constraint-based architectures.

Guests

Nathan Benaich

Topics in this episode

GitHub CopilotLlamaElevenLabsOpenAI O3SynthesiaAirstreet CapitalDeepseek R1ARC-AGI benchmarkAlphaFold3ESM3

Questions this episode answers

How expensive is it to run OpenAI's O3 model, and why does cost matter?

OpenAI spent approximately $1.6 million on total inference compute to solve ARC-AGI benchmark problems, with individual problems costing upwards of $1,000 in inference compute - far too expensive for routine use until costs drop dramatically through optimization and competition.

What's the difference between how AlphaFold3 and ESM3 approach biological modeling with AI?

AlphaFold3 uses a pure diffusion model approach without engineered inductive biases, while ESM3 retains inductive biases (separate encoders/tokenizers) to integrate multiple modalities like DNA, RNA, and protein structure - suggesting different scaling strategies suit different biological problems.

Why do startups like Synthesia and ElevenLabs beat big tech companies at AI products?

Startups achieve faster market fit by crossing the quality threshold first with superior UX and having fewer organizational constraints, while big tech companies often hesitate to enter domains perceived as risky (digital avatars, synthetic voices) and must retrofit AI into legacy platforms.

How should investors evaluate Gen AI startups if the underlying models keep improving?

Focus on founding team execution ability, design taste, product sense, and customer affinity rather than ML credentials, since frontier models now enable rapid iteration without massive compute budgets - the best ideas often come from domain experts, not AI researchers.

Will cheaper frontier models eventually replace cheaper open-source AI models like Llama?

If inference costs for frontier models drop far enough, they could displace cheaper alternatives across most use cases, but we're currently in a "blitz scaling wars" phase prioritizing performance over margins, so this transition may take several years.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B65%
  • Speaker A35%

Most-used words

seems39different35feel32question32systems30back30system29models28hard26money23model22interesting22openai21super21seeing21point20

Episode notes

Episode 142 Happy holidays! This is one of my favorite episodes of the year - for the third time, Nathan Benaich and I did our yearly roundup of all the AI news and advancements you need to know. This includes selections from this year’s State of AI Report, some early takes on o3, a few minutes LARPing as China Guys……… If you’ve stuck around and continue to listen, I’m really thankful you’re here. I love hearing from you. You can find Nathan and Air Street Press here on Substack and on Twitter , LinkedIn , and his personal site . Check out his writing at press.airstreet.com . Find me on Twitter (or LinkedIn if you want…) for updates on new episodes, and reach me at editor@thegradient.pub for feedback, ideas, guest suggestions.

Full transcript

1h 49m

Transcribed and scored by The B2B Podcast Index.

Speaker A: Merry Christmas or Happy Holidays. Whatever you're celebrating, I hope you're enjoying time with your family or, um, grinding through Christmas and New Year's, if that's your thing. As a holiday gift, this is one of my favorite episodes of the year. For the third time, I sat down with Nathan Benesh, who is founder and general partner of airstreet Capital and who I'm very lucky to call a friend. As always, we got into selections from this year's State of AI report, as well as some other news that we've both been following. You might hear that we've both got a little bit of the winter sniffles. It's very cold outside these days. As a guy who grew up in Arizona, winter is definitely not my terrain. But with that, I hope you enjoy the episode. I hope you have a wonderful holiday season, a wonderful new year, that 2025 is everything you're hoping it will be. And here is everything you need to know about AI in 2024 with Nathan Benesh. This is our 2024 edition of the end of year. What's everything important that happened in AI this year? I'm sure we'll be able to fit it all into this podcast, no doubt. We're talking a couple of days after the OpenAI03 announcement. And, I mean, I have a bunch of questions for you about what you've been up to, but I feel like we kind of have to address the elephant in the room first. The question that everybody listening to this is probably asking is, uh, am I cooked? How do you. How do you feel about things?

Speaker B: Um, suppose the honest answer is hard to know, uh, because, yeah, for the avoidance of doubt, no one actually has access to this other than, um, I think, uh, safety testers, red, uh, teamers. And, yeah, the benchmarks were on, uh, really difficult math problems and the RKGI prize, um, and then some, uh, PhD level, um, sort of research question. So I don't know if I'm cooked yet in biology. I haven't seen, like, the evidence of it yet, but, uh, I would say, like, at least it pours some cold water on the recent narrative that, um, yeah, it's hitting a wall. Um, I think this is to some degree exacerbated by media who kind of grip onto, like, mentions of, like, hey, scaling is hard. I think what's, What's. What's really exciting here, um, is that reasoning seems to be making really good progress. So the last couple of years, there's this debate of language models or stochastic parrots that they reproduce statistics and they're just next token generators. But if they can generate really complex reasoning traces that are A effective and B sometimes better or distinctly different than what humans would do, I think that's a pretty exciting result. Now like with other AI systems and certainly, um, kind of like these chatbot style interfaces is like you really, you really need to know how to poke them and how to coax them into a behavior that you want. And I still think that that's, there's a really big like learning slope that's going to have to come to bear. It's a bit like we need an Apple Genius bar for like how to, how to different AI systems for our everyday jobs. Um, and only at that point will we know whether we're cooked or not.

Speaker A: Yeah, it's uh, it's hard to say. I mean it seems like we're still in the regime where this stuff. So I guess kind of backtracking a bit, this is falling off of zero, uh one where we first saw this case of like whoa, scaling up inference time compute a bunch yields really, really great results. And now we have all of a sudden another dimension along which we can start scaling. And so everybody's thinking about this. People were kind of already focused on mitigating inference time costs and now it seems like that's going to be even more of a priority going forward if 0103 style models are to be the kinds of things people want to use. But maybe some details for us to add here about O3 was that for ARC AGI they spent upwards of $1,000 or something on inference compute per problem to the point where they'd spent a total of $1.6 million on just doing inference for this to solve the problem, which is hilariously expensive. So on the one hand you might think okay, we can use this stuff. And um, you said people will have to prompt them correctly and things like this. So maybe it gets reserved for only the most super important problems that people have. But inevitably this stuff is going to get cheaper and cheaper. And so maybe you at some point in the future get something like O3 that is cheap enough that you can apply it with less judgment about what exactly you're going to have to do. But what do you think about all that?

Speaker B: Yeah, I think this is a killer question. Um, and we, we wrote like a trilogy of essays on, on this question of like frontier model economics and does it all make sense? And uh, you know, I tried to uh, on the one hand like draw a simple analogy of, of like if you work on a laptop every day. You don't. And um, use Mac. You don't really go and buy like this super expensive cheese grater station. Like it looks really cool but it's completely unnecessary for the vast majority of us. In the same way you wouldn't really drive a sports car every day to go pick up your coffee in New York because it's unnecessary and not fit for the job. And so even if you had a system that was super powerful, why would you use the God mode for simple tasks? I think all that kind of goes out the window if God mode is as cheap, um, as entry level mode. But um, I'm not sure to your point whether that's yet true. And it might not be true for a while as we're still in this kind um, of blitz scaling wars, um, where we care less about margins and care more about performance and market share, which is probably fine in the history of technology. But again, I think it's a bit hard to reason about O3 and OpenAI systems because we don't know really what's behind the hood. It could be, um, a smorgasm of models that are working together based on the complexity of the task that's instructed by the user. Um, and so perhaps there's a efficient way of handling these queries that we don't yet see.

Speaker A: Yeah, there could be. I think this also opens up another question and something that I've been seeing people talk or kind of worry about since O3 came out and everybody noticed these costs, which is there is going to be a period of time. We don't exactly know how long it is going to be yet, but there will be a period of time where the best of the best models like O3 are going to be prohibitively expensive for most people. So a lot of people will be able to use Llama and things like this. And I don't know how cheap DeepSeq R1 is exactly. I know that was sort of competitive with O1, but there's going to be for some period of time it seems a gap between the ability of the models at the very, very frontier that people are able to have access to on those other models. And so you do kind of see this gap start to emerge and maybe eventually it gets closed, but it seems like it's going to exist for a while. And so the kinds of questions people might be asking about that are then what does that mean for economic inequalities or just like the sorts of problems and economic value generation that are able to occur between those who are able to cough up the money for O3 and its ilk versus everybody else. How do you think through that?

Speaker B: Well, uh, I was just trying to think on the fly reason on the fly about analogies in the real world. And I mean there are for sure society and the economy is built in a way that bunch of services are available at different sort of scales of your willingness to pay, um, and your. And your. And the necessity that you have to procure that service. Also, I wonder if, uh, you know, these really complicated queries tend to be ones that are asked by people who anyway, like have budget, like for example, complex scientific questions. You know, the average person doesn't wake up in the morning thinking about that or really spending their time working on that. But like scientific organizations that have significant grant money are certainly spending their life working on this. So, you know, while like one of the big promises of technology is to democratize or to like reduce the barriers of access to information, computation and just like doing work, et cetera, I do wonder whether still reflecting some of these like asymmetries and what people have access to in the digital world is just normal. Like that's just kind of how the, how the world works. And there will always be individuals who have certain proclivities that are more expensive than others and they'd be willing to pay for. And that's how capitalism work.

Speaker A: Yeah, that's totally right. And maybe another take on this is something like when I was pointing to economic value creation and the generation of a business or something like this earlier. It's not necessary that the complexity of the question you're trying to solve or how smart you are correlates to how much economic value you end up creating. Probably there's some amount of that, but certainly you see a lot of businesses out there that just generate insane value that are, you know, the idea itself doesn't require that deep of like a scientific background. It's just executed really, really well.

Speaker B: Yeah, and I suppose as people get more used to using these systems, like, we'll probably be able to like trade off what they're good at and what they're not very good at. And you know, maybe we don't need like, uh, a system to autonomously do absolutely everything or, you know, we only want it for the parts that we kind of don't enjoy or want to supercharge or go faster with. And so, I don't know. I think this is all to play for and all to figure out. It's a bit hard to take a view, let's take a stern view about how the future is going to pan out.

Speaker A: Well, before we get into a little bit more about the future or maybe a little bit more about the past year, airstreet's had a very interesting year and I want to talk about that. So I've been following a bunch of what you've been putting out on airstreet Press, which I know has really put out a lot of great essays this year. You had a great chat with ISO Kant on um, scaling laws, I guess. How would you describe this year for airstreet?

Speaker B: Well, I think for me just being interested, I think like you in machine learning for a long time it feels like product market fit has been uh, sort of achieved or at least like product interest fit has been achieved. Um, so you know, many of the ideas and um, sort of like excitement for opportunities that I've had over the last like 10 years have really come to bear now. And it just feels like go time. So we're trying to be present, be in the arena, as people call it, and try uh, to make some good decisions along the way and work with um, the best founders we can find. And um, I think it's been for me also cool because um, the science and tech project side, which is really why I do this job, big year for science and biology. There was obviously a Nobel Prize for deep learning, Nobel, uh, prize for protein folding, and then also Nobel Prize for micrornas, which was like the class of molecules I worked on in my PhD. That was like really cool to see, uh, kind of like some vindication on that. And I think it's also been uh, a year of huge vibe shifts. I mean some themes that were totally rogue or not popular for a long time now became super popular, like defense and security being one of them. But more recently robotics, um, you know like 12 months ago or something was still uh, in the category of like hardware is hard, don't touch it, it's expensive, it's not going to work, it's brittle, customers don't want it to now like, you know, please take my money basically. Uh, and um, that's been very exciting. And I think this idea of like software automation and like agents has also come to the fore a bit more. Um, you know, for a long time in the industry we're like, oh, when is GitHub going to release some AI coding tool? And, and then it turns out probably like the biggest hit out of all of them was not like a model company or not like a, a big Tech company but you know, a startup team that just created an amazing user experience around these systems. So I think the, the other part that I really enjoy in this job and like this year kind of shows it too is whenever you think that the chess game is sort of like added checkmate state and uh, you couldn't envision any other way to get out of it and rejig the power, it happens. And so yeah, just trying to look at the present and uh, not take what's happening today as the final state of the board. It's still early innings.

Speaker A: I mean everything you said I think also just points to. This is maybe a question for anybody listening who's also an investor or a vc. It feels for all of those reasons like a really hard time to understand where do I make bets? How do I invest well, how do I make good decisions? How do you think about doing that?

Speaker B: Yeah, one of the changes I've made is to not be too fixated or too judgmental about the starting idea in Genai, particularly because this is really like uh, you know, V01 experiment land and things will change really fast and the best teams will change really fast. Um, um. And so uh, I think I've, I've made some like painful mistakes by overjudging what um, you know, first demo I was shown. Um, you know the best founders obviously iterate around that they find adjacencies around what they're working on that are, that make you know, the opportunity more attractive. Um, so it's really not being too judgmental. And then the second thing is uh, frankly like not like kind of de emphasizing the like machine learning credibility uh and gravitas that affects founding team has and emphasize more um, their capabilities in execution and taste around design and product and like their affinity for who they're building for because um, you know, with these like frontier systems like the AI space has been gifted artifacts that allow it to iterate at the speed of SaaS without having to burn their entire seed budget on train data collection like hyperparameter tuning, blah blah blah, before they get to test whether their product hypothesis even works. And so I think it's really going to be very exciting to see what products will come out in the next few years as they're no longer just designed by AI people and AI research people, but they're going to be designed by people who have those problems and are really gifted uh, in building consumer apps like uh, consumer Internet services, et cetera, et cetera. That AI folks are not like waking up every day thinking about, and I

Speaker A: guess the theory of the case here too, about there's a differentiation between somebody outside comes and builds something that's leveraging whatever OpenAI anthropic is building. And they have this really good sense of design taste of the consumers in their particular domain. And I think a lot of people also wonder, what If Google or OpenAI builds a team internally to do that? I think the kind of classic back and forth you have about this is small team, greater focus, maybe even greater sense of what the customer is looking for in this domain. And so that's naturally an advantage. But you've also written about this a bit. So one of your pieces I really liked was called Being first is a mode. And I do think that there's something to be said about just getting what, like getting mind share, getting, you know, what it is that you're building into the head of the consumer and it being like the first thing they go to with a very nice user experience. So how do you think about what that competitive landscape looks like when you do also these big labs which probably have a greater concentration of ML talent, but could probably also hire some great designers.

Speaker B: Yeah. So to take one angle and then the other. So on the one hand a startup that's done this very well is Synthesia in the avatar, um, domain and then 11, uh, labs in the audio domain. Both of these were, you know, to some degree capabilities that big tech companies were working on. You know, meta had Facebook reality labs that I don't know how many bajillion dollars they spend, like scanning humans to try to create 3D representations of them. And then they came out with this like hyper gimmicky VR product that no one liked. And then on the other hand, like for 11, um, Amazon and other big tech companies had text to speech systems, but they sounded super robotic. Um, and there was a bunch of open source, um, tools that were out there. But I think in their case, like we hadn't crossed the uncanny valley yet in a way that was super easy to use. And potentially in both cases there was like some, some kind of trepidation, like uncertainty amongst big tech companies to kind of go into that world of like digital humans or digital voices. Like, I don't know, maybe it's a bit sketchy, uh, in their, in their view, which was like the wrong view to take. And that creates a bit of open air for startups to show that hey, it is actually really cool and customers really want this. But then on the other hand, um, you have examples where like the Kind of bigger labs are there first and it's hard to dislodge them. And having recently upgraded iPhone to try the Apple intelligence, I am so unimpressed with what they've produced. Um, even if it's just kind of integrating intelligence into a camera where they're trying to feed ChatGPT into your iPhone camera and the, the UI like looks terrible, it's super slow. And then you kind of realize like, you know, the camera is not the uh, interface through which I'm going to ask follow up questions or something else mind and I'm going to go ask, ask it this just like I would use ChatGPT for. So I'm like in that case, like, you know, ChatGPT is basically with normal people synonymous with AI. And so it's really, really hard to say like, oh yeah, but if you want to understand the world around you, like use your camera. Even though I can just take a picture with ChatGPT and ask it the same query and then move on to my other stuff and have all the other conversations in there.

Speaker A: Yeah, it's um, I mean in part there's this question of like where, I mean for Apple, anything they do with intelligence is going to have to integrate into like the stuff they've already built and they already have a sense of what their UI is. And that was originally built in a time when ChatGPT wasn't exactly a thing. And so I guess to some extent it makes it uh, harder for them to reimagine that. I mean even somebody else who is externally trying to develop how an AI system is going to fit on your phone or something like that, they do have to play within Apple's bounds. But maybe to the extent that Apple has to natively integrate this into their software, their iOS, that imposes certain limitations for them. That may be somebody thinking all this in you doesn't quite have.

Speaker B: Yeah, it's, it's, it's possible. Um, at least on my side, what I would have been excited with, with Apple's intelligence is just its ability to orchestrate instructions with other apps. So you know, you're like uh, dreaming in this phone, like, oh, I need to go to this like meeting, can you please like book an Uber for it? And then the only thing it does is it opens Uber. Well, I could have done that with fewer characters. Some things are actually easier to just tap on your phone, so maybe we will get there. Um, and as you mentioned, apps have to play with Apple so that's to their advantage as well. But it just unfortunately feels really half baked and not in a great way.

Speaker A: You were talking a little bit about some of the stuff that was really exciting for you in the research space, mostly with bio earlier. And I think that's a theme that we've also talked about the other two times we've done this. And so I want to get into this a little bit. Maybe part one of this that I wanted to get into that I thought was kind of interesting was you noted in this year's State of AI report the way that AlphaFold3 gave us another case of the bitter lesson appearing in this kind of World of Biomolecules 2, where instead of the equivariance constraints, this inductive bias that AlphaFold2 had in terms of modeling proteins, AlphaFold3 is like, let's just use this diffusion model and scale up uh, even more. And that seemed to perform better. But it was also kind of interesting because there was another model where you kind of had the opposite story, I think, where you still got a good amount of performance with this inductive bias built in. And so the story seemed slightly different. And I'm wondering if you have a sense of reconciling those two cases in your head.

Speaker B: Yeah, well in the esm case, like ESM3 if I recall correctly, they had inductive biases in order to integrate different modalities into the same model. And uh, so this is like trying to like integrate uh, you know, DNA, rna, like expression structure, et cetera. And so to some degree you have to have like an encoder, um, for each of those, um, for each of those modalities, at least a tokenizer for those modalities that represents them appropriately. And then in ah, AlphaFold's case it's, it's purely just dealing with um, with 3D structure. And so they're, to that extent you need probably less modalities to feed into it. But I think it's still very experimental. Um, my bet for next year is we'll see some even more exciting results from uh, scaling different aspects of model development in bio. And I think a lot of the lessons that we've learned in language are actually applicable, which to me is quite astounding because in language you can perfectly perceive the input and the output and then you can reason and say, yeah, this is good or this is bad and you're not really integrating across, I guess, lots of hidden or unknown states. Apart from of course, if you're like, I don't know, asking a model to reason about the universe and Somehow making a bunch of, um, intellectual leaps there. But with biology, you know, the amino acid sequence and you're trying to predict a 3D structure and there's all sorts of quantum stuff that's happening, uh, inside a cell that dictates how that protein will take shape and um, whether that shape will be static or whether it'll wiggle around and to what extent it denatures and forms complexes with other proteins. And it's just kind of wild. The model can learn that. Um, and it's not immediately clear to me how you actually properly integrate across all sorts of modalities, from like, genomics to imaging to like cell behaviors into a single system. And it'll be exciting to see the results because people are trying.

Speaker A: Do you want to speak to. You gave the example with ESM and how it's using different modalities. Do you have any sense of what you think the most promising directions are?

Speaker B: I think for starters, using what we've learned in language, um, models would be really helpful. So direct preference learning, DPO is useful. Um, scaling up pre training, scaling up alignment, um, figuring out what that even means for your use case. I think in some areas of biology it might be useful to have different modalities because as an example, in medical imaging there was evidence that deep learning classifiers could detect disease in images really well. But then when you did feature interpretability, you would find some weird artifact in these images to explain how it did its classification. But these were not biologically coherent. So you're like, okay, it performs well, but it learns the wrong thing. So is it actually like learning anything useful? Whereas potentially, like, integrating different modalities into that system might give it, you know, more help, uh, like a better scaffold, um, to make like, more scientifically sound predictions. But one of the other things I remember was, was, uh, that neurips, when, um, when Ilya said like pre training is over, et cetera, um, a bunch of the bioml crowd on Twitter was, were like, yeah, maybe in language, but definitely not the case in biology. We've got loads of tokens we're not taking advantage of.

Speaker A: Yeah. When you hear Ilya say something like that, I feel like some of the theories that float around are like, this is motivated reasoning because of the new thing he's doing. What goes on in your head when you hear somebody saying something like that?

Speaker B: Yeah. My main question is, uh, to people who uh, take it at face value is do you understand the incentives of the, the person that's saying these things, um, either way and I think it's really hard to take people at face value unless you really know that. But you know, like, I mean the guy is obviously brilliant so it's, it's probably not entirely but um, it might not also be be ah, gospel either.

Speaker A: Maybe a slightly stretched version of the truth or something like that.

Speaker B: Self serving stretched version of potential truth. Yeah,

Speaker A: uh, that seems like a good way to take it while we're still in bio ML, I feel like one of the important things that you noted in the report that I feel also makes it harder to evaluate progress in the field is exactly that. Evals and benchmarking in bioml remains poor and partially it seems like there's uh, a skills gap there. There just aren't enough people with skills to both do the frontier model training but then also give it a rigorous biological appraisal. Uh, those are two pretty distinct skills. I think you've seen a lot in academia now. A bunch of people with serious linguistics backgrounds, for example, who have the ability to do stuff with these models and then also subjected to a rigorous appraisal along lines of linguistics. Um, so you see people like Tallinnson doing all this great, great work. How do you think about, I mean I know people are still trying to work this out and it's been discovered that there were some kind of physical violations in AlphaFold 3 for example, and recursion is doing stuff on this. Where do you think we need to go to kind of mitigate some of what we're seeing here?

Speaker B: I think the answer is cross disciplinary teams. Uh, I think it comes back to education. In a way it is wild that engineers and research, um, contributors are building general purpose systems and then have to be responsible for the behavior of these systems in any one of the possible ways that a general purpose system can be used. And that's like an inhumane ask or a superhuman ask rather. Uh, so it's certainly easier in areas like math, coding and language because that's what people who are good at AI are also good at makes sense. But in biology it's still not um, baked into the curriculum that um, individuals who do a PhD have to know computer science or have to at least know scripting or have to know how software works, et cetera or bioinformatic tools work. So I think efforts that try to improve this cross disciplinary collaboration will actually do a lot better. And that's probably some competitive advantage that some teams will have because they'll know what useful alignment means, um, in whatever scientific domain they work in, uh, compared to the average Contributor who doesn't know the end use case.

Speaker A: Are there successful examples of this right now that you're seeing like ICM recursion and some of the folks that you're in contact with seem to be doing a pretty good job of putting bio people, tech people together. What does it look like to do that?

Speaker B: Well yeah, so I think Recursion does a fantastic job of that because they have very strong biologists, they have the experimental infrastructure from their company and then they acquired Valence, one of our portfolio companies that's really strong in generative modeling. And everybody works together in a cross disciplinary way in that setting. It helps to have both dry lab as they call it, and the wet lab, uh, under the same roof so you can troubleshoot things. Know, in the earlier stage startup side, like I think the folks are profluent, which is another team that we work with on AI and protein engineering, like do a fantastic job there as well. Like they're really at the frontier, um, on the modeling side and at the frontier on the biological side too. Um, and like it shows, you know, like in the board meetings, like we spend a good amount of time like peering into like the experiments that they've run and like thinking, um, hey, there's actually not that many groups that are really pushing the frontier here in AI and bio like we are. And we can do that because we have kind of both teams that really sit in the same office and take each other's contributions seriously and troubleshoot things together. So I think the best companies really understand that there's data and then there's data that's actually natively useful for what you want to train a model to do. And certainly like in experimental biology there's data that's fit for, fit for purpose for machine learning and then there's just data that's useful for people to look at. And um, without wanting to like ramble too much on this, but like most of the biological data we've been creating has been built and represented to be interpreted by humans. And I think one of the interesting examples here, um, at Recursion and they've been public about this is like the company started by doing um, uh, what's called cell painting, which is basically trying to like kind of do image recognition, detect things in cells and like label them, uh, fluorescently or project a uh, label on it so humans can look at it and like identify things. And to do that you need to like fluorescently, uh, label proteins, label cells so people can make that interpretation. But it turns out like after Many years now that uh, they have AI systems that can perform equally well without the cells having been stained fluorescently at all. So like now they just use brightfield, which as a human looking at a brightfield image, you just see like structure, you see like the boundaries of cells, you maybe see the nucleus, but like, I can't see anything more than that. But like a machine can. And so like, who cares if a human can't see it? Like if a machine can see it. And that's useful. Like, great.

Speaker A: Yeah, it's almost, um, taking what we were talking about earlier, you know, there's always been a sense, we've often seen the case of spurious artifacts that these things are kind of latching onto in order to make classifications. So that's also the sort of thing that you can use if you're really careful about it, towards a more positive purpose.

Speaker B: Mhm. So we yet to really see what fruits that brings because it's still early days, but I think it's so far been a pretty exciting trajectory.

Speaker A: Yeah, I guess we've gone over a couple of things about this. We talked briefly about genome design. Is there anything that you feel in the bioml space that isn't on people's radar so much that they should be paying more attention to?

Speaker B: Well, I'd say still gene editing is pretty amazing. We do actually have two approved medicines now in the US and in uh, the UK for sickle cell anemia. And there are like, there is evidence you can do gene editing like in vivo, so in a human patient, like by just swallowing medicine. Um, and I think that's, that's pretty incredible. The other area is like cancer vaccines or like personalized cancer vaccines. So this is the idea that basically a cancer like has a bunch of mutations and as a result of its mutations, it uh, expresses certain proteins on it on the surface of its, um, cells that are not normal. And a personalized cancer vaccine basically, uh, rests on sequencing a patient's tumor, figuring out the mutations that are in the tumor, and then predicting what weird antigens or antibodies rather are on the cell surface and then designing MRNA's that produce antigens against those antibodies in a cocktail format. So it's like you predict all the weird, uh, antigens and then you stack rank the most popular ones basically. And then you make, make MRNA's for all of them. You put it in a cocktail and you jab that into a patient and that's, that's personalized. And you have Moderna and Biontech that are working on this and like early results seem pretty interesting and uh, given also like the success of MRNA for Covid. Like we know that for tumors that are, or at least diseases that are easy to deliver to, this stuff could really work. I feel like people are not really talking about that enough and it seems pretty magical that it works. And this also rests on machine learning, uh, to make those predictions from, uh, genetic mutations.

Speaker A: That's a really exciting direction. Yeah. I want to get into a couple of more themes on research and I feel like both of these are kind of wrapped into a bit of the AGI esque discourse. And I guess that's kind of come back to us with 03 and the two of them are general architectural innovations along the lines of what Deepseek has been doing. And they put out some really, really interesting stuff this year as well as open endedness as kind of a broad theme. Before we get into these though, uh, with the theme of AGI being a thing again, and this also came up in your discussion with Isocant about models for coding and things like this and what that looks like. What's your reaction when you hear people talking about AGI? What is your personal orientation towards that as a term and a concept?

Speaker B: It just feels like an ever changing yardstick. But what I think about is a system that can autonomously, uh, conduct and solve tasks that are economically useful at a level that's the same or better than humans. I'd say for certain things. It does seem like we have AGI in a narrow domain. Like a system can definitely do web search and web summarization better than I can, you know, um, and some of the complex reasoning seems pretty nuts. Like, I mean the math problems that it can solve is like, was way better than what I can do. Um, so like from my vantage point, that's like AGI and math for me, you know. But I think, um, the other aspect of this is, uh, just how quickly like technology diffuses, um, among society. And I've always felt like the discourse of AGI kind of presumes that once we have it, like it's like a light bulb switch that like you switch it on and then now the world's like totally different because you can see in the dark, um, and you can't go back. But I feel like in reality technology diffuses way slower than that, um, because we have tools that are available to solve all sorts of problems. But you know, a lot of the time like industries that have those problems don't want to adopt them because of like human inertia. Or like refusal to change or I don't know what, like regulatory compliance or just like the how. How things work or culture. Seems like it's only recently that Sam Waldman, like now acknowledges and says it publicly that, like, yeah, actually, like, inertia is quite big amongst humans. And so we might have AGI build, take a while to, to diffuse. And so in that sense, I think it'll take a while for us to all appreciate, like, yes, we have achieved AGI.

Speaker A: I feel like the quotation I keep hearing from Altman is something like, AGI is going to arrive a lot sooner than people think, and it will matter a lot less than people think. And I think that also wraps into one of the really interesting. I know we're jumping ahead a bit, but just because it seems related. There was a paper from economist Zari Nakamoglu this year and that argued that the impact of AGI on general economic productivity, and I think this was. I forget what the measure is called, but the measure of increase in economic productivity that is not due to just capital inputs like labor and things of that nature would be less than 0.55%. And I remember that paper got a lot of criticism. So I guess the question here is, I see on the one end, I don't know exactly where Altman is, but it sounds like there's analogies between what he's saying and what Akamoglu is saying. On the other hand, you do have a lot of people, I mean, this is terminally online Twitter people kind of thinking about the fast takeoff scenario and everything being exponential. And now that we have zero, uh, three people are taking this as further evidence to that point. And so there's this additional rush of kind of proclaimed urgency about needing to answer lots of questions about AI safety and things like this. And I mean, it's hard for us to just kind of sit here and aren't sure, philosophize about what's really happening here. We know better than everybody else. But where do you feel skeptical? What feels more resonant to you when you read things like this?

Speaker B: Yeah, I think people's views are definitely shaped by the environment that they work in or that they evolve in. And I can, I can certainly believe that if you're at the epicenter of all this action and this is all you live and breathe and read about, um, that your worldview is shaped by that and your timelines for acceleration are way, way faster than somebody who, I don't know, lives in a completely different part of the world that can't even transfer money online because of some bs. So it's certainly a gradation of problems, um, that um, alter your worldview. I think mine's like somewhat in the middle. Like I, I do think that a lot of tasks that are done on a computer that like a human can say are good or bad and there's enough evidence for how people like do do that task. Like eventually that'll get solved pretty fast. Uh, once you get into the behavior of swapping your workflow from doing it manually to uh, to getting a machine to do it, like you're not really going back. Um, as an example, uh, if you wanted to narrate an article, like you have to sit there in front of your computer and like when the siren outside went off, like, ah, crap, I gotta start again. Like when your neighbors like kid is crying. I mean it takes you like hours to do all this and now when you have a voice clone, I mean it's 95% good. So like why would I ever read a piece of text again? And uh, you know, I was watching like House of Cards again last night, like episode one, season one or something and I remember like, you know, in the newsroom of like the Washington whatever paper and like all these people bashing keyboards to write articles. I'm like, oh my God, this looks like back in like touring days, you know. Yeah. Like, why are people doing this?

Speaker A: Yeah, yeah, I don't feel that way.

Speaker B: So I think uh, I think this stuff like creeps up on you and it creeps up on some people way faster than others, but eventually it does. And, and then you kind of look back at the last 10 years and you're like, Jesus Christ, how far have we come? And a lot of the tools that we take for granted today, I think you showed to somebody 10 years ago and they'd say it's magic for sure.

Speaker A: Yeah. And I think some of what you said points to again that case of differential economic impact. Right. I mean, early adopters and regions where people are way more tech forward. The San Franciscos of the world are probably going to experience like visions of what this change looks like a lot more quickly than places that are not like able to put up, you know, a gigawatt nuclear power plant.

Speaker B: Yeah, yeah, yeah, yeah. Or of the volition to change. There's plenty of, plenty of like organizations that I think could operate way better by just, I mean, using better software and they still don't want to do it. And I don't understand why. Yeah, like a mentality Thing. So I think some industries are just still like in dinosaur land and eventually like the meteor will strike and then things will change.

Speaker A: Yeah, I mean it's still the question of like tech diffusion is hard. Right? It's like yeah, the meteor strikes, but then where does all the residue go? Do people actually like pick up on things as fast as we think they will or hope they will?

Speaker B: I think it has to be either has to be like out of curiosity, uh, and willingness and desire to improve and be at the frontier or like existential dread.

Speaker A: Yeah, that seems right. And I mean, I guess I wonder how much it's a question of. Everybody says this industrial revolution is going to be different but in some ways it does feel like the potential economic gap that could arise from some people embracing this technology and then other people not embracing this sort of technology. It does feel wider than before.

Speaker B: Yeah, but it probably has to be relatable like if Google is like going to make a ton more profits, like does HSBC care? Or you know, does like Wells Fargo care? Like I don't know, probably not. Has to be somebody relatable like in that, in that industry that really shows like hey, you know, customers are getting rinsed because like the service is bad and it's too expensive and it could be a lot better. And then you see wholesale, like hopefully wholesale movement of customers with their feet. Right. And I think that that's certainly. I use the banking one because I think that's one that's certainly happened over the last like 10 years. I mean you're like starting to work in London like where it's a big financial center. Like so many people work in banking and the service sucks. And then you have like Revolut and Monzo and Wise. They're like, hey, we can do this better. Like in part because there was some good like government regulation to create sandboxes for like new actors to, to go and try to start financial institutions and. And now you're like, oh, uh, I can't use Revolut or I can't use one of these other neobanks. It's really painful.

Speaker A: I want to step back into the how we get here question to the two topics I mentioned earlier. I'm thinking maybe we should start with Deepseek and what they're doing since that's kind of closer to the status quo right now. And then kind of get into open endedness, which I think is like very both technically but also kind of philosophically interesting for lots and lots of reasons. But Deepseek, I mean super interesting. They came out with Deepseek Moe this year. There was R1 competitive 01 their coder model. And we've both read I think the translation of the interview with Lian Wanfeng, their CEO, which had a ton of really interesting insights. One of the things I think he pointed out just to give people background on this and I'll uh, post a link to the translation and the show notes for this. But a couple of things kind of came out with the CEO. He's himself a really interesting guy. Both somebody who knows how to mobilize lots of capital and resources but also dives in and also does the research with all the other researchers. And um, has this philosophy of just organizing these really young people and giving them the money and resources to do their things and then they come out with multi adlaten detention and things of this nature. And one of his points in that interview was that uh, what we lack in innovation is not capital but a lack of confidence and knowledge of how to organize high density talent for effective innovation. And this is an interesting differential. And also take on the dynamic of where is the main innovation coming from. Because DeepSeq has been making architectural innovations that Americans aren't used to seeing out of China. Like MSR Asia did come out with the RESNET model way back when that was huge for computer vision, but we haven't seen as much since. I'm wondering for you, how do you see this dynamic here? Do you have thoughts on. We were talking earlier about inference optimization and things like this and DeepSeq has been doing a ton towards that end as well. Where do you think this goes?

Speaker B: Mhm. Well for one finance was one of the first domains where machine learning really had um, its ah, Hollywood moment basically. I mean there's been you know, so many storied quant firms that are chock full of people who would otherwise be working in like big AI labs. And uh, and to some degree there it's a bit similar to big AI labs where they try to find a metric and hill climb on it. And here in quant firms you know, you have the metric and it's you know, successful trades and you hill climb on that. But uh, so I'm not entirely surprised to see cool work coming out of quant firms like on that basis. I do think that the west has overwhelmingly been pushing a narrative in the last couple of years that China copies and Europe regulates and the US innovates been the broad thing and I think it's just probably overshadowed a lot of uh, stuff that happens in China. And also the other point in previous reports that we found was a lot of Chinese AI research is published in local language that's not captured on Arxiv. So we just don't see it. I think the other snippet that I really loved, uh, from this interview and I'm just looking at it, the interviewer goes, I think he's talking about how expensive like AI research is and it's like kind of eye watering costs of large models. The interviewer says, but this process is also a money burning endeavor. And uh, Liang, the CEO goes, an exciting endeavor perhaps cannot be measured purely in monetary terms. It's like somebody buying a piano for a home. First, they can afford it and second, such a group of people are eager to play beautiful music on it. Such classic, beautiful way of phrasing things from the Chinese. I really like that, um, there is definitely an art to all this, uh, and then also put in the report. But it's pretty amazing how uh, there was very little if anything open source contribution from Chinese groups in the last uh, you know, 24 months outside of the, the previous 12. So it's like really just kicked off. It's a bit perplexing to know exactly why. Um, I think probably the best bet is like these individuals would also be happy to come to the US if they could and you know, they might, might be able to and to retain them locally. Like this is the way, this is the way you have to operate. But it's been, it's been cool to see, like it certainly shook up the status quo. So like back to what we were talking about all the way at the start. Like, you know, when you think the status quo is, is fixed, like it's probably up for uh, a rejig.

Speaker A: Yeah, I mean the nationalism question that you're kind of poking at there is interesting and Lian Wongfeng himself was entirely educated in China. And it sounds like there is that sense of wanting to cultivate homegrown talent. And this kind of ties into something that we're probably going to get into later with Jensen Huang's comment of every country having a national LLM. And like Ian Hogarth, who used to work with you on the State of AI report, had that really great article about AI nationalism a couple of years back. And so it feels like we're seeing this kind of thing like intensified.

Speaker B: Yeah, and that's one also where like, if you don't know the sort of incentive behind it, it's hard to like read as to whether it's True or not, I mean I think it's a really good uh, kind of growth vector for the company uh, to be pushing this narrative. I'm not entirely sure like I buy why countries need to have like their own, their own language models, uh m. It harkens back a little bit to some of these um, European uh, projects, uh, whose name I'm now blanking on. But uh, you know one was like hey, we need to have our own search engine because Google is an American company and it's not going to retrieve and produce uh, artifacts that represent uh, the European cultural heritage uh, appropriately. So we're gonna like invest hundreds of billions of euros or whatever or into trying to do that and like obviously end up failing because most government like technology efforts like end up failing. Um, so I'm not like entirely enthusiastic about this idea of like domestic language models outside of maybe like low resource languages or maybe like maybe some countries that have a differential view on uh, on copyright as I understand in Japan, like they're quite permissive over uh, over copyright is one of the reasons why you have some lab space there and the culture is like super different. And so maybe that, maybe that's like a good justification but uh, at the limit. I think it, I think it uh, gets this question of like what is the incremental amount of data, new uh, data I can add to, to an already ginormous pre training data set that makes the model that much different or that much better. So it might be less about that and be more about again this preference optimization and fine tuning.

Speaker A: I want to get into open endedness in a second but maybe a last question on this Pete is there's kind of two conflicting things I can see going on here at the super high level where we're talking about lots of these architectural innovations, lots of these post training techniques that can really improve models and who's going to get ahead on the basis of what model is the best. And you and I for a while saw in the Twitterverse this whole lots of people were very loud about how great Sonnet 3.5 was and things like this and how it had overtaken OpenAI's models. And Leon Wanfeng from his part is also pretty optimistic about how deepseek can compete against the likes of OpenAI. He literally said OpenAI isn't a God and they won't necessarily be at the forefront. And that seems like very reasonable. At the same time though, to your point of being first As a moat, ChatGPT does have at least the American mindshare right now to the point where ChatGPT is starting to kind of become a verb more and more. And so when people think of this technology, the first person they're going to think of is OpenAI just because they were there first. And you know, you might have quibbles over whose model is exactly the best on evals or for your particular use case. And at least in my case I'll kind of move around between models depending on what seems best at the time. And so I would say that there isn't a lot of stickiness in my personal case except if you're subscribed to Claude or subscribed to ChatGPT Pro Enterprise. Totally different question, but how do you think about that question of there's going to be somebody with the best model, there's also going to be somebody with the most mindshare and those models might or may not be super, super close. But there is a difference in those two questions in terms of who is winning from a monetary level.

Speaker B: Yeah, I do increasingly think that if you're not at the frontier and you're not at the top like 1, 2, 3, it is basically irrelevant. And I think one analogy there, um, could be in the early days of smartphones there was iOS and then Android and Android had a huge plethora of, or an increasingly large like plethora of devices and uh, that were running it and you could fork it and do all sorts of wacky stuff with it because you own the operating system and it's open source et cetera. And then it feels like in the last few years I don't really meet anybody who is Android anymore. Um, and so it feels like at the limit people want good performance at a good price that's reliable with a nice user experience. At the limit again that's probably at this point OpenAI and um, anthropic on the consumer side. Uh, I mean Google seems to have great systems but like their UI is still like unusable, um, or it's just like unnecessarily unfriendly and complicated and then maybe perplexity on the search side and a few other services for different modalities. Uh, but once you're there and the service works and it's affordable, what's the point messing around with something else? And I think if the open source uh, angle was actually so compelling then you would probably see um, some kind of slowing in the revenue growth rate of the big um, Apple style companies in AI and you don't see that like do normal people Even know how to interface with like a llama model?

Speaker A: Definitely.

Speaker B: I m mean they say that they have what, 600 million like monthly users or something, something like this and honestly don't know where they come from. It might not, might be from different markets in the US because everybody I talk to in the US is like confused as to why I need to have it in WhatsApp, Instagram, Facebook Messenger. And in Facebook like you're searching for m your friend or something in there and it thinks you're writing a query. You're like this is so confused.

Speaker A: Yeah, I guess it speaks to the, it's like a more general version of that question of what does the kind of design look like when you're serving the customers. And in this case you have a different sense of OpenAI and Anthropic were kind of purpose built for these general purpose models and so they're building their UI around the existence of that and that alone. Whereas Facebook, Apple again have that, there's that incumbency, there's a lot of money. But then also they have the platforms that they like have to integrate these things into.

Speaker B: Yeah, yeah. Or they have to launch a separate service and then somehow like push traffic to there which, which the, you know, meta has done pretty well with threads. Uh, but then there it like from what I understand like caters to a different kind of person or like a sort of sub community within Instagram if I understand correctly. It was like designers and et cetera that really liked it. Um, whereas ML people are still on Twitter. So yeah, it has to be like a compelling reason to move. Uh, it's just like good old fashioned product 101 I think.

Speaker A: Yeah, the stickiness question is hard and it's always that the improvement you get from moving from your old product or interface to the new product or interface has to overcome whatever the friction is for you making that switch. And the gap had better be pretty big for you to even notice that.

Speaker B: Yeah, yeah, yeah. And that's, that's the point with uh, you know being early as a mode is we're in this sort of like generational for the moment transition of companies onboarding AI service providers for various things and I think once they make their decision they're gonna stick with those ones because the schlep it's gonna be to switch is not worth the like improved benchmark performance which is not even relatable to anybody outside of like neurips.

Speaker A: Exactly. Okay. I guess last kind of set of things I wanted to talk about AGR related we should also talk about robotics. But first there's a lot going on with open endedness. I feel like I've been seeing this come more and more into the limelight recently and I think that in general it's just a super interesting take on what does it look like to build an intelligent system. Um, and we've seen lots of papers since then. There was like a position paper out of Google DeepMind even before that. Joel Lehman, who I've talked to on this podcast did a really interesting work at uh, OpenAI called Evolution through Large Models. And there's this idea of continuously generating artifacts that are novel and learnable to an observer and so kind of abstracting that to the way that we humans experience the world. We leave artifacts in the world and we learn how to do new stuff. And the environments that we're in are continuously changing so you're never overfitting to a particular environment and you have to figure out what is that common substrate or whatever that allows you to continuously adapt to exactly whatever it is that you're facing. And it seems like the open endedness research is the area that has this set of problems, the best kind of technical articulation of this set of problems in my view. But I was kind of interested to see you also put out an open endedness all we'll need essay. So I'm wondering what about it has caught your attention and then how you're thinking about its role in the grand scheme of things.

Speaker B: Yeah, yeah, I think it's because I, I just vibe with this idea that um, that like language models trained on static data sets, like obviously it created a kind of shortcut or made it very easy to consume and access and work with information. But without some form of, of like ability to recursively self improve or plan how to do that, it's hard for me to see how a system can generate something that's like genuinely new to like me as a user new to it and then it can do some useful work with it. Doesn't mean like the former is not useful because faster Google is definitely useful but um, but uh, you know we're all in like well that's probably not true but like some of us are in the business of like trying to advance like knowledge and do new things. Um, and yeah, so I think as a new frontier or as like an emerging frontier for how like software systems can take agency over tasks that we want to accomplish. Like it seems like um, being able to explore, do search recursively, self improve is like pretty critical. Um, and it could be even you know, back to like the O1 and O3 stuff like their you know, progress like really worked because we had very good reasoning traces for, for how uh, experts would try to solve mathematical problems, et cetera. Um, but having a system that could somehow figure out those traces on its own without having expert demonstrations would be pretty sick.

Speaker A: Yeah, the self generated feedback loop does seem like the most important version of this and I guess we've seen proto versions of this so you mentioned AlphaGo earlier as a very good example of that. There was the more recent strategist out of Meta as well for multi agent games. That was also a self improvement kind of deal. And so I guess the question is a lot of these are pretty, there's stuff that's coming to be useful at least in neuro domains and I think maybe a question then coming out of that is like where does that training process, that improvement process meet what we're thinking about in terms of agents and things like this that are doing what we call economically valuable tasks and stuff of this nature?

Speaker B: I'm not sure. Uh, so in biological science it would be useful to have these open ended systems because you can't potentially some data points are just really expensive to capture in the real world. Um, some logical jumps from um, one modality to another are difficult to make. Jumping in and out of different fields that we're not experts experts in is hard to do as well. So it'd be pretty cool to have a system that could um, that could take like a starting hypothesis and then figure out kind of roll through what experiments you might need to do to validate or invalidate the hypothesis and then maybe to your point like it invokes like a human needs to do an experiment uh, to actually test if this theory works before you could continue further. And those are probably where like in the short term I'm more like bullish on this. Um, yeah, this idea of having like these kind of proto open ended systems who can help like direct our own research to produce artifacts that would be useful for their own reasoning because otherwise it just feels like a system is going to magically like explore the unknown and do all the hard work for us and somehow like figure out the theory of everything on its own. And that seems a bit unrealistic if it can't like sample from the real world to generate new, new um, data points and is otherwise operating on everything that we've created to date like in a way, I don't know, it's like um, in the early Days of telescopes, like they put simply, they had, they were like very pixelated, the resolution was super crap. And so you point that into the sky and you probably can't count the number of stars and planets because you just don't have a good enough telescope. And now we've got like super HD stuff. And so you perform the exact same experiment and you get a different result. And so if an open ended model is purely trained on data from like the crappy telescope, how is it possibly gonna figure out that there are actually like way more stars out there without somehow sampling again with like better tools? I don't know how without more information about the unknown, it can somehow figure out the unknown.

Speaker A: Yeah, I guess uh, at some point we turn into real world contact and more data for the system.

Speaker B: Yeah, hopefully not in a doomsday way.

Speaker A: Hopefully not. I think a last beat on research that was really interesting. One of the vibe shifts you'd mentioned earlier is robotics becoming en vogue again. In 2021 we saw OpenAI disbanding its robotics team and now this year it's been rebooted. We've seen Google DeepMind doing a bunch of stuff in robotics and I think a lot of that impetus has been from just seeing language models and vision language models, being able to do things like planning and kind of working with robot arms. We thought of various versions of the robotics transformer. Really, really great work. And it seems like diffusion models kind of helping improve policy, action generation, things like this. So a lot of it seems like there was of course direct robotics work happening, but then we saw these advancements in other areas and then LLMs and um, the concept of them developing world models which there's lots of different definitions of and opinions on. But that really kind of fed back into the robotics world and seemed like it opened up a new set of avenues for things to work a lot better than it did before. But that's kind of my take. But how do you understand the arc of robotics kind of going out of the limelight and then coming back again?

Speaker B: Yeah, I think it's um, driven by a, like on the, on the tech side, the fusion of vision plus language, the ability for like systems to perceive like a visual scene, reason about it and then pass that information to a robotic system that can actuate and perform the action. So like for example, like we have this company Cereact that provides like um, industrial pick and place and other warehouse automation systems and they have a product which scans every single bin that's in one of these warehouses and tells you what's in it, which is good for analytics. But you can also then instruct uh, the robot system to go like grab things from that bin just in natural language, um, which would have been entirely impossible to do like 10 years ago. Which is great for debugging and programming. But um, what we found is actually like you now no longer have to have a vision system that's trained to recognize all the objects that one customer has and then train an entirely new system for another customer because the systems are not general enough. Now you can just have like one system that performs really well across like all customers, um, which for customers that have been sold robotic solutions for the past 10 years that have kind of been brittle and like hard to go live with and hard to maintain, uh, against the backdrop of like labor crisis post Covid is still a thing, um, is now kind of creating a perfect storm of like we actually want this stuff and we want it like now. And we haven't seen that in like a decade I would say. And then there's like you know, a fringe of groups that are trying to apply the same scaling principles through demonstration mostly by, by saying uh, like a human wearing like a cyborg suit but like, but basically like commanding robot arms and like demonstrating actions of how you might FL fold T shirts uh, or like pick and place things and then using that demonstration data uh, to uh, to train a robot system to copy that kind of ah, motion. And it, it seems to be from all the Twitter demos like pretty impressive compared to what we've seen before. Uh, and then moving into like humanoid space, uh, where again like demonstrations look pretty impressive. But we don't know if this, these robots are teleoperated or truly autonomous. Um, uh, but yeah, like across those three domains there's been a real resurgence of interest. Um, not even mentioning the fact that consumers can use self driving cars in San Francisco pretty reliably and in Arizona. So that's a pretty amazing feat that was poo pooed for I don't know how many years before it became an overnight success. Success.

Speaker A: Yeah, I mean I guess in that domain too there was a lot of over promising with robotics. There's like similar stories to be told. A lot of potential promise out there. Although I don't know, I don't know if I've got the same sense of, I mean besides like Tesla's you know, humanoid bot. But I feel like I'm not seeing quite as much of this. You know this is totally going to change the world like immediately kind of discourse around Robotics right now when I compare it to the way the self driving discourse looked a decade ago, maybe slightly less than a decade ago, but that seemed a lot louder than what I'm hearing from people doing robotic stuff right now. I don't know if you feel the same way.

Speaker B: Well, I certainly hear and see quite a lot of companies that are training robots to fold T shirts and do your dishes, cut tomatoes in the kitchen and stuff. Way more than what was happening in the past. And generally they're quite convinced this will get solved. So maybe it will happen. But then back to your point as to like who can afford this and how broadly, um, is it distributed? Um, but um, the other cohort I'd say is uh, we talked about this a little bit earlier at the start of uh, kind of like the bitter lesson and end to end learning works and self driving that being true. Or um, you see performance and scaling of wave system like the company in London that sort of took a very early bet on end to end learning. Uh, and then testing kind of gravitating towards their approach and, and Waymo with their GEMMA model also going back to their approach. Um, so yeah, there have been people that have been pretty vocal about certain aspects of robotics for a long time. They're now like being proven correct. Um, um, and I'd say the humanoid crowd like seems pretty loud, but maybe it's my, my microcosm in Twitter. Uh, yeah. And then that's like not even talking about all the kind of autonomous um, defense and warfare systems uh, that are built by Anduril and like Delian and a bunch of other companies. Um, and I think those are all pretty exciting and we'll probably scale even more now in the new uh, administration coming into the U.S. yeah, uh, I

Speaker A: think we'll come back to some of those things in politics. I feel like we should do a little bit of time on the compute question and this is more of kind of getting into the industry section of things. So like we'd mentioned actually I guess we didn't mention in this. I was listening to another, I think I was listening to your conversation with ISO content. I'm like conflating that with some of the things we've talked about. But I remember you two were talking about just like the GPU problem and the question of distributed training is really, really hard. And we all know this is because of having to manage you know like just the size of data centers, the networking that's going on between GPU racks and things like this. And when you're Able to grow something or create a system that's got more chips on it. And you've got Nvidia working on GB200 and it's oh, that does like mitigate the problem to some extent. And so we're seeing Nvidia's ambitions growing. We're also seeing some of the competitors and I feel like I've seen more from the competitors this year. I mean we saw like the graph core stunt. I feel like Cerebras is getting more mind shared than I've seen them get before. And like even Salmonova I feel has done like a couple of things that have gotten them into the limelight, which is really interesting still. Like nobody is coming anywhere close to Nvidia, but I am seeing more attention being paid to these folks. And your compute index and what we're seeing on the State of AI report still does bear out still massive, massive differential. People are not moving away from Nvidia that much. But there's more people who are trying to let's say, launch stones at the giant. How are you thinking about what we've seen this year?

Speaker B: Yeah, I mean uh, it's entirely unsurprising given the ginormous revenue and software margins that Nvidia is making. So uh, with a quasi monopoly of that size, why would you not try to um, you know, get some of that pie? Um, so it seems like the more credible efforts are Google CPU and then potentially like Amazon Systems. Like um, but then outside of those companies, um, I'd say like it's still pretty, pretty um, little impact. Um, we saw the gap at least in uh, open source AI research papers that you know, make mention of using a specific vendor's uh, AI accelerator. Um, in that experiment like we saw the gap for Nvidia, all Nvidia papers versus like everybody else, uh, grow smaller this year than what it was last year, largely driven by like a 500% increase uh, in number of papers that use TPUs. Uh, and then on the startup side like there was an inflection in number of papers that made use of Cerebras systems. But um, but it's still like tiny. It's like orders of magnitude smaller than, than what Nvidia has. Uh, I think Nvidia annually is like roughly 30,000 papers or something like that and cerebras was like 200. Um, so uh, yeah, and then um, and then I think the other way to look at it is uh, like these Neo clouds like Core Weave, uh, or uh, Cruzo or Lambda, um, or Nebius, like None of them buy anything else than Nvidia. So if there were demand for non Nvidia stuff, they would probably stock it. Yeah. And then just for fun, because we got asked this question, a lot of, uh, why should I not put some money in an Nvidia competitor? Now seems like a good time, etc. And you're sort of like, oh, I've seen this TV show before, it doesn't end well. We looked at like the $6 billion billion that was invested across like six competitors over the last like few years. Um, and like looked at the portfolio that that 6 billion would be worth today. Um, and it's roughly like 32 billion as of when we published the report. And half of that value is Cambercon, which is like a Chinese entry company is listed there. But if you bought Nvidia stock, uh, 6 billion worth of Nvidia stock on like the same day that the press release came out that, uh, you know, investor X invested Y in competitor z, uh, that 6 billion would be worth 120 billion, um, back when we published the report, and now it'd probably be higher. So it's like a 20x and I think it's since 2016 or so in like eight years, which is obviously like, you know, Candyland is a VC and this is a public company, you know. And so I feel like. I don't know what more to say basically, you know, but the amazing thing with all this is like, um, Like I've never, I haven't really found like an Nvidia person that'll like, kind of gleefully tell you that like, they're. At least the people that I meet are still like, acutely aware that this is insanely competitive. Like, it's a. It's like an amazing reward from all their work and it's flattering to see that. But like, they know that, you know, if we don't keep innovating, like, we could lose the race. Um, and they seem to be willing to like, you know, make, make bets or build systems that might even eat into their revenue, as long as that's what customers want. And I think that's a pretty killer mentality, um, to have at that scale.

Speaker A: Oh yeah, it does seem like they're not slowing down that much. There does seem to be something fundamentally quixotic in lots of ways. But we'll have to keep watching, I guess. We've been seeing a lot of collaborations and relationships between generative AI players. We've been seeing various acquisitions. So there is, of course, the huge databricks Mosaic ML acquisition. And we're also seeing antitrust come to the fore again in more ways than one. Not just about AI, but in the AI world. You're seeing model builders and their partnerships with big tech companies, of course, most famously OpenAI and Microsoft. And so antitrust regulators are now kind of worried about where that goes, the Microsoft OpenAI thing right now, externally. And from what people are saying, it feels a little bit shaky at the moment. It's kind of unclear where that's going to go. But the regulators do seem to be also trying to figure this out. They're feeling nervous about Nvidia. Huh. Do you have, like, a read on where some of these regulation questions go?

Speaker B: I mean, it feels to me like it's. It's kind of. It's go time, you know, I think I, um, mean, if you want to, like, kind of consume the edge opinion here, you know, you listen to Alec at Alex Carp from Palantir, uh, who basically, like you unabashedly says, like, you know, there are nations and countries that will adopt AI and they will be part of the future and everybody else is going to die. I feel like, I feel like those who are pro growth kind of get that, or that's, that's their narrative. Um, whereas countries that are more along the, like, oh, we should regulate, are just going to miss the boat. Um, and, um, I mean, this is like, I mean, maybe not directly relevant, but given the sort of narrative in the US right now, because we're all reading the tea leaves, is like, what the next four years will look like and kind of imagining our own features there. But if you kind of believe the, like, uh, that this is going to be great for technology, great for, like, regulation and stuff narrative, then probably this holds, which is like the, the market cap appreciation in the S and P, like from the date of, uh, November 4th or November 5th, like a week or ten days later, plus, like, the appreciation of like, crypto assets like Bitcoin. Uh, that amount was like, worth single digigit percentage points of, like, the entire European annual GDP. Um, it's like whatever, 1, 2, 3% based on, like, my back of the envelope calculation, um, which I feel like is pretty significant. Um, and, and then like, contrasted with like, you know, Thierry who was like, pushing a lot of these, like, regulatory narratives in the EU Commission, then, like, leaves and, like, change his Twitter profile to, like, entrepreneur. It's just like, next level, like, vibes. Uh, and, and I hear like, on the ground, like, people who are involved in, uh, the EU AI act are like feeling, oh, we might have gone a bit too far. Um, and, and uh, probably the EO in the US will get like torn up. Anyway, it was non binding, but, but given how close a lot of tech executives are, um, with the new administration, like uh, you know, these kind of AI and crypto czars that are getting appointed that are clearly like, you know, tech and VC people, um, it's hard to see how like we'll enter into like a super regulated, like AI is dangerous time. But a lot of things can change because, uh, I think there's some um, kind of opinions that don't really align. You know, on the one hand, like Elon wants to have his own AI company, but he's also like thinking that AI is going to kill us. And um, you know, he's rather like kind of ah, dovish with China. And then the Palantir Teal crew is very like anti China. So how does this all get sorted out? I don't know. Then Trump seems to be anti automation, it seems.

Speaker A: Yeah. I mean the vibe shift is interesting and complicated and to what you were just saying, different sections of the world and different political factions seem to have different reads on the landscape. The way you kind of described the overall vibe shift was from basically last year, worries about AI safety and this stuff taking over the world to now Please buy my consumer app. You see Claude ads literally everywhere. I saw one in D.C. you know, they're just like all around. And so we had this. And then like just as we mentioned earlier, O3, there's like more AI safety discourse kind of coming back, but this is more of. Everybody has their own super extreme take on what O3 means. So it's like hard to fully characterize, I guess the landscape of vibes. Here you can see where people are coming from and why they land where they land, but there's just like a lot of different voices there right now. It feels.

Speaker B: Yeah, for sure. I think that two of the areas that companies have been pretty consistent about with regards to safety is A, like biorisk, bioterrorism and B, kind uh, of military systems. Um, and on the military side it's been a crazy vibe shift to see uh, OpenAI change their terms of service to then say like, hey, our systems can be used by uh, you know, for like defense use cases and then meta, like also making that change after, you know, they had this whole hoo ha in China of like, oh, we fine tuned a llama system to do some sort of military planning exercise. We don't even know if that's like, useful or not useful and anthropic also defense partnership. So I'm not saying that's bad, like, it's probably a good idea actually, but it's still wild that those same individuals are now like, pro, uh, pro defense. Um, and on the bio side, uh, we wrote an article about this, uh, quite recently about, uh, hey, is some of this debate warranted? Is it overhyped? Um, can AI systems design bioweapons? And we already have pretty specialized biological design tools which are quite useful for designing, I don't know, new sequences of anthropology racks or things like that. But it turns out that whenever you're working with some, uh, sort of biological substance that can kill people, it probably kills you first. Um, and so the sort of Venn diagram of people that want to kill people and have enough knowledge to use AI tools to go way faster and have lab automation and can buy the reagents and manufacture all this and distribute it is pretty small. And at the end of the day you probably end up in this. I can't remember the. An acronym of this bell curve you see on Twitter every time, where it's like, yeah, use some fancy sexy AI tool to design this crazy biovirus that no one's ever seen before. To impart a bunch of damage is in the median distribution and on either end of the fold. This is like build an explosive device. Uh, I don't know. I feel like it's kind of overstated for now. And we certainly shouldn't be putting arbitrary compute, um, thresholds on biological models above which they're suddenly dangerous because that's going to be pretty problematic for a lot of really exciting drug discovery research. And, um, for AI and autonomous systems, I think the reality is you need it because, um, there's no way you're going to defend against huge swarms of drones if you don't have something autonomous. It's full zone.

Speaker A: That seems totally right. I feel like, um, one articulation of how people, some see the sort of regulation or light regulation scenario playing out is what you mentioned earlier. Like different countries have these very different takes on regulation and we can get into that more. But the US being pretty light on it or sort of reluctant EU going forward in lots of regulation. Although there was some kind of walking back with the UAI act, and I know that they took off facial recognition, they walked back a bit to allow this for law enforcement. With China, they've really been doing a lot on this, imposing pretty strict standards and wanting the AI models coming out of their top labs to be to kind of not look like there's censorship going on, but sort of toeing the party line. And one way I hear the way some think it will play out in the States is something like, you know, the models keep developing and then like something really, really bad happens. Somebody dies on account of something that happened with a model. And then people are like all of a sudden, oh, let's actually begin to regulate. This is maybe getting into like predictions for the future kind of territory. But I'm curious how you think about the different regulatory landscapes playing out in terms of where things are going to go for startups, how the impacts are going to look as they interact with the different regulatory environments.

Speaker B: Um, well, a base case. I feel like countries that are uh, you know, host to companies and organizations that create and uh, and monetize these tools will probably be softer uh, on them. Like it would be, it would be so crazy if in the US there would be a bunch of legislation that rate limits how AI ah, can be used by companies and uh, and how big like models can be trained. I mean this would like basically pop the bubble. And like if you pop the AI bubble I really don't know what's left. So for that to happen I mean you got to be like, yeah, like a death wish basically. And um, so that's super dangerous. I think it's way easier for countries that don't have an endemic um, in a strong like AI technology base to be like exerting regulation because some degree that's like one of the ways they could put brakes, one of the perceived ways potentially they could put brakes on other countries lapping them. Um, it's like yeah, we'll make it a bit harder for you, but it turns out we'll make it even harder for ourselves as well because we have a bunch of other issues that there. Yeah. And then the other like big question probably is um, around like uh, just large scale M and A in the US because um, many of these organizations are so expensive to run on the AI side. You know, at some point they're going to need to have a home somewhere or they go public, but most of them will probably need a home. And if you can't create like a natural cycle of like money flowing through the system, that's also bad. And it's bad for everybody because like, you know, if like the highly valued AI, uh, company or, or company that raised a lot of money from like good investors doesn't find a home and doesn't see that money come out then like that pension system and that university endowment's not getting the money to go like uh, provide to its constituents or its members or its students. And that's bad for the country as well. So everybody's got to eat basically.

Speaker A: Yeah. I feel like on the opposite side of a lot of this regulation, some countries are very enthusiastic about what's going on. So I think that the main example of a sovereign wealth fund getting into this has been the government of Abu Dhabi. And they have their hands in a lot of pies like Cerebras. Uh, I mean Sam Altman was reportedly talking with them. Um, how do you think about what's going on there?

Speaker B: Yeah, well, um, I think it's smart because it's an incredibly rich country. Um, and um, they're generally quite good at finding the new, new thing and being involved in a new thing. But um, what seems interesting is they're tackling it from two angles. They have one which is like let's import foreign technology into the country which is through this G42 initiative, this cloud computing behemoth through which they have investments and bought a bunch of systems from Cerebras and Microsoft and Nvidia, um, to build like DCs etc in the region. Uh, and then they have like the let's Try and Home Grow Something initiative which is their AI71 company, um, basically um, which kind of spun out from the uh, Falcon initiative which was like a open source, uh, Arabic language model. Um, uh yeah, open source project. And they're putting like serious like money and talent behind that and they've recruited people um, from you know, large AI labs, um, and universities, like less so on the American side but more like from various European countries, MENA and, and Asia. Um, and then trying to see like which which one of those two basically wins.

Speaker A: Um,

Speaker B: but, but yeah, they gotta like they got to transmit money that comes out of the ground into useful capital and this is a great way to do it. Um, and then when you see um, some of these uh, cloud CSPs like Crusoe, that have a pretty cool um, value proposition for countries that are energy producers and that want to align both kind uh, of clean energy with like clean uh, energy powered uh, computing infrastructure, like this kind of a perfect, perfect fit. So I think it'll be, it'd be pretty cool to see like what will come out. Um, and my guess is that more people will be attracted towards going to some of these regions because you know, other countries in Europe just don't seem that exciting anymore. Like, you know, I think most people want to be somewhere where like people have ambition or taking risk or putting real resources into the ground to get things done rather than meandering around legislation and oh, this can't be done. Oh, there's no solution to this. It gets boring after a while. You either go to the US and if you can't go to the US you go somewhere else where the similar mentality is uh, present.

Speaker A: There's also the question of what does it mean with US companies getting like really, really significant investment from these sovereign wealth funds. And you could see analogs of the China hawks having had trouble with US companies having investments from China. Do you see much of an analog there? How does that look to you?

Speaker B: Yeah, it could, it could be, I think any, an investment of any above a certain scale is going to be like questionable from national security perspectives. Like, you know, why, why is like a sovereign wealth fund investing in OpenAI when like the US government could invest in OpenAI or the DOE could invest. It's not like America lacks the money or can't print the money. So why are we like seeking it abroad? Um, then there is, there's going to be like some crescendo of like tastefulness where like China's probably on like the uh, wrong end of that spectrum. And then um, and then various countries in the UAE are probably more in the middle somewhere. And in the day like it looks like um, Abu Dhabi and like the UAE is like a very close like strategic partner of the us. Um, and um, and so on that basis like there might not be an issue because so many like American venture funds are anyway supported by these groups and if I'm not mistaken they buy like all the advanced like US military equipment. So until that stop, until you can't buy the F35 anymore, I think you're probably in okay land it seems.

Speaker A: Another kind of direction I wanted to go with this was just the international environment regarding. We were talking about the EU AI act and regulation earlier and one of your 2023 predictions was that there would be limited progress on AI governance beyond high level voluntary commitments. And that's pretty much what we saw borne out. The UK had an AI Safety summit in November. There was a summit in Seoul in May. France wanted to move in a different direction and called theirs the AI Action Summit. So lots of high level non binding stuff. And I think you see to what you said about Abu Dhabi for example, countries sort of aligning themselves with respect to what is their stance on AI, what is their stance on technological innovation, like, who do we want to attract here? And so then this AI safety question too, and whatever summits you have around it seem to become one of those platforms for, hey, you're a startup, you're really ambitious, come to our country, come here instead. Because we're aligned with your values of moving fast and being low regulation and giving you the sort of space you need to do your work really well. And like France is in some sense there's a lot of AI talent there. So you're kind of seeing this work out in different ways. How do you see this kind of continuing to play out?

Speaker B: Yes and no. I mean France has always had a lot of technical talent because of their university system. Um, so many major contributors in the US are French and probably will never come back. Uh, but the total number of contributors in AI that are operating in Paris is probably still very small. Um, and so I think there's been a good amount of positive spin and marketing around that. But on the ground I wouldn't say there's that many people. And then when you see how politically volatile the environment is, um, where now I don't think there's that much certainty over what direction the country is going to take. I, uh, mean it looked like Rosie a couple of months ago and then, um, and then now with like a big reshuffle and like a mixed government, seems like complete mess basically. Um, so I think, I think it's still, um, it's still like, it's still hard to be operating um, in Europe I would say, um, it's just like easier when you're in a bigger country with like a cohesive market. Everybody speaks the same language. You got one regulator to care about. You don't have like two dozen. Um, um, and still like the playbook is like European companies launch offices for particularly for go to market the very least like in the us And I don't really see that changing too much. So um, I would just be careful like calling, calling winners in other countries like too early with by sort of uh, underestimating the seriousness of politicians, uh, uh, own seriousness when it comes to AI.

Speaker A: So I guess kind of what I'm hearing there then right at this moment, very hard to work in the eu. Of course, eu, AI act and gdpr. It's also made it hard for US labs to launch their products there. And the UK is definitely going further into this. They launched the AI Safety Institute. There is also aria, so they've got their kind of own lab for AI safety research and it does seem like the UK seems to be, wants to do a lot in terms of putting themselves at the forefront of thinking about AI safety and stuff like this. So do you see that difficulty, I don't see that difficulty of operating in Europe changing for the foreseeable future. Do you see any hints of like a vibe shift or anything like that there? Or do you think this is just kind of how it's going to be?

Speaker B: I mean, there's like micro vibe shifts in certain communities that want things to change, but I think some of them are kind of tone deaf to the reality of politics and then others are connected to politics. But like, there are other priorities that um, that unfortunately, um, take precedence, whether that's, you know, education or health care, immigration. Um, you know, while like the current French president like, likes AI and technology, I don't think any of his like, competitors care. I mean the existing UK government like, sort of cares about technology, but I wouldn't say is like, well known to be very well versed in it. Um, so, uh, so yeah, and you know, on the, on the AI safety point in the uk, like, I would agree with you, it looks like that institute is doing like, really good work and got um, you know, pretty high praise from Anthropic for like testing their like Sonnet 3.5 number two. But, but at the same time, like, does having an AI safety institute attract like for profit companies that are going to create like market cap and jobs and like tax taxable revenue in your country? Like, I don't think so. Like, what would be cool is maybe, um, drawing a lesson from what we were talking about earlier around, uh, creating like regulatory sandboxes for new companies to have a shot. And so, you know, one of the topical things now in the west is like, hey, we need to have defense, because turns out that if we get invaded, we're like beeped, basically. Uh, and so why doesn't the UK for example, create, create some kind of procurement sandbox for advanced autonomous AI systems for defense, uh, and put a bunch of money behind it, do advanced procurements. You don't have to deal with any of this favoritism with archaic primes, um, and weird political hoops, uh, that you have to jump through and put real money to bear here and run actual competitions and use things on the field. That would be sick. Like companies would really move for that. Um, same thing in like, you know, in, I don't know, gene editing, like, you know, we did the approval and that was probably one of the good things with, with brexit or at least one of the good things that was implemented post Brexit is having like a separate regulatory agency for medical uh, approvals, um, and so managed to push one through. Like it would be cool if like they got a bunch of like AI design companies to do more around gene editing and created like a faster path to clinical trials and approvals. So I think that that's the kind of, that's the kind of uh, policy architected like tech advancement that I think would be cool without being like a politician myself. That sort of first principles seems a good idea and could actually create like economic value.

Speaker A: Yeah. I guess the question is if the politicians feel that way and how likely we are to get there in the future.

Speaker B: Yeah, exactly. I just think the average voter just doesn't care about that. The average voter cares about, do they get their pension? Do they get health insurance? Is their energy bill high or low? Can they afford school for their children? Um, are immigrants taking their jobs? That's unfortunately what people care about. It's not, is AI going to kill us? They probably have no idea what that means.

Speaker A: Yeah. And I think that definitely adds on to what we were saying earlier. Not quite the same as diffusion lags in terms of stuff being used, but just the gap between the average San Franciscan or somebody in London thinking about AI versus just somebody who rightfully is worried about a lot of these questions which are far, far more pertinent to their everyday lives.

Speaker B: Yeah. And you got to pitch to the level of the audience you're selling to and make it relevant for them.

Speaker A: I think another kind of aspect of this is just how people are articulating the harms of AI and there's this differential. So one of the things you pointed out in The State of AI report was, um, Google DeepMind had done some interesting research on whether we're actually focused on the wrong harms here. And there's lots of sophisticated, really smart, interesting exploits of AI systems and defenses against those. But it sounds like most cases we actually see of misuse are from easily available tools. We're seeing many different cases of deepfakes and things of this nature and that is costing people money. You are seeing deepfake porn, which is extremely unfortunate and that's getting easier and easier to generate. And so that does seem where the bulk of this is going to happen because why would you invest that time when deepfakes are a pretty effective tool to scam people out of money? So that seems like how things are going to continue to look going forward. And I remember when there was back and forth about deepfakes and sells, where on the one hand you had this stuff is going to continue to get better and better, but then you also had researchers on the other side who are like, wait, people are actually pretty discerning about what is and isn't real. But it seems like you flood the space enough then you can get cases where if deepfakes are super cheap, you just keep doing enough of this and eventually something works and you do get money through nefarious means. Somehow it seems like the incentives are there for this to get worse.

Speaker B: It could be. But I think we also um, slept walk through one of the most major political events with a lot of fears that deepfakes and fake voices and all sorts of stuff would have a big impact. And it seems like nothing happened or like I haven't read anything. There was no like New York Times headline about anything happening, which was pretty, pretty like positive news I would say. I think people are just like quite good at adapting. Like people are not dumb. Um, but, but like at the limit. Yeah, you could flood the market with all sorts of opinions and things of that nature, particularly around like bots. And if it's in the written format, I think that's probably harder to, to really discern the image. And video is maybe a bit easier because we use multiple senses and you're sort of triangulating across those senses and it's harder to like, to fool. But just reading text and unless it says like, uh, uh, what is it like, ah, uh, and then certainly, uh, you want to know if an AI wrote it.

Speaker A: Yeah, yeah, certainly.

Speaker B: Let me tell you why this candidate sucks.

Speaker A: Uh, we've got some tells now with some of the systems, but that seems right. Fair worry to have. I think we should talk through some predictions. Uh, this is always the really fun part of the report. So one that we kind of touched on earlier already was the question of investments from a sovereign state into US large AI lab. And you pinned the number at uh, a 10 billion plus investment of that nature would invoke a national security review. I guess we'd already been building up the reasons to think that this sort of thing might happen. Given what we're seeing from the Saudis. What's your sense of if that happens, where it's most likely to come from?

Speaker B: Yeah, I think based on who has the money, it'd probably be Saudi, uh, Abu Dhabi or maybe Qatar, not um, Dubai because they don't have the money anymore. Uh, or at least the oil's gone. Um, and Then could be, I mean other countries that have a ton of cash would be like Canada, um, or Norway, uh, or Singapore frankly, even like Australia. Like their pension system is enormous. Um, but those ones seem like pretty, pretty like vanilla. So I'd say it's probably more something, something like Middle east or Asia. But then again like OpenAI has that NSA person on the board. Uh, I don't know if Anthropic has anybody like from the security complex or government on their board. Maybe this was like a bit of a quid pro quo thing, I don't know. Um, but it could be like a newcomer like the, like uh, the one that some OpenAI alums are allegedly working on. But also the other thing is probably with these big AI labs that raised massive rounds through existing venture funds. Who did SPVs, like these sort of companies that are raised specific to invest in, uh, companies. Those ones potentially have already some sovereign states in them. So at what point uh, does the concentration get too high and then that invokes some kind of like, hey, what's up here? And then realistically, even if a sovereign state owns equity in a company, if they don't have board control or uh, information rights or anything else aside from not funding it anymore or actually getting some kind of right that would block future financings or block M and as and therefore be a way to exert hard power over the company, they're going to have nothing more than self power, which may be enough.

Speaker A: I see like a lot of scrutiny coming up. I'm curious with like the new administration how all this is going to look. I know that there was the platform of them walking back the executive order from, from Biden, unclear what their perspective is going to be on the labs. I mean, I know you'd said Trump was kind of like anti automation, but

Speaker B: then he has potentially like uh, Elon parroting in his ear who's very like anti OpenAI.

Speaker A: So, yeah, so, so I could see More scrutiny on OpenAI as a result of that. I think another one that kind of ties to some of our themes was one of your predictions was the early EU AI act implementation. And we still have yet to see how implementation questions play out. Ends, uh, up softer than anticipated because lawmakers worry they've overreached and I guess we've kind of said a lot about this already. But do you have theories about parts of the AI act or things that you feel are most likely to be stepped back and just what that looks like?

Speaker B: I don't like a good enough recollection of the specific terms. But I think, um, what seems very frustrating is just like you can power up the exact same device in North America and power it up in Europe and you get different services. And so in one sense you're like living in the future and the other one, you go back in time and that seems disadvantageous for consumers who are paying the same price for this thing. I don't see how that really protects them in any way. And like, if we've learned anything from ah, GDPR with this, like cookie pollution on websites that make them basically unreadable. Unless you like, X out like five different boxes on a, on a website like. And I don't see how that protects the consumer either. I would hope that somehow people can like express this and be like, can we be done with this nonsense? So it might be more along the lines of what you said of like, some companies are really getting forward or like, kind of going a lot faster because they have access to capabilities that their competitors don't. And then we sort of like call it quits on hamstringing. You know, like European companies and contributors because of legislation that's supposed to protect them but actually just makes them worse off.

Speaker A: It's like a lot of echoes of gdpr.

Speaker B: Yeah, I mean, if you ask a normal person what is gdpr? They'll just say it's cookies. They don't even know what it is.

Speaker A: I had initially some other predictions I want to discuss, but I'm m actually maybe more curious just for you. On the side of things that you think have a greater than 5% or greater than 10% chance of happening in the next year, are there any particular things that you think are likely to happen in the next year that are very high on your list of things that it would be great if this happened? And then also on the other side of that list that you'd be kind of terrified to see happen, um, I

Speaker B: think would be great would be, um, this prediction we had around an app or a product that's built by somebody with no coding ability goes viral and Flappy Bird ends up in the top 100. Because I feel like that's so liberating and cool. And I think we're already seeing snippets of this and I sort of can't wait to see that happen. Just see way more cool things. And then I think it'll also kind of unshackle the fact that consumers of, uh, big tech company products are victims of the average user phenomenon where Google has to sell to the average person who Uses Google Maps, not like the niche communities that use it in specific ways. And those people will never be catered for because. Because they're not average and they can't edit it because they have no capabilities to do so. And now it's like, oh, you don't like that? Well, just build your own. And then, um. And like obviously what would be terrible is like the flip side of the things that I think are overhyped which is like uh, yeah, like some. We do actually invent some like insane like biological system that's like super dangerous and like doesn't kill the person that makes it and is spread very easy and, and can go undetected. Like that would be pretty scary. We had one prediction around a research paper written by an agent or an AI scientist would be accepted at a conference. And um, I think that would be not scary. But I think it does call into question not our ability to detect fake stuff because it doesn't mean a machine did it, that it's fake or not of good quality. But it's sort of like, oh, maybe I'm not the best at this anymore for a. That's like used to being the ones that are architecting these things.

Speaker A: Uh, the existential dread.

Speaker B: Yeah, exactly.

Speaker A: Maybe as a last question, what's on the horizon for airstreet for the next year? What are you excited about? What do you want to do this next year? Whereas you mentioned a lot of stuff that kind of come to fruition this year in terms of things you've been thinking about. So what does that look like over the next few months? The next year?

Speaker B: Yeah, well, I think a few things, um, excited about a lot of the in person like community building. We've been doing like, we've been running like um, AI meetups in lots of cities like sf, New York, like Munich, Berlin, Paris, London. And that's been awesome as everybody wants to meet like peers and learn from peers and it's. It's a great way for people to like learn best practices. So excited to kind of get back into the swing of that tour. I think there's some, some companies that are like reaching interesting inflection points like you know, Profluent and gene editing and Delian and Defense and um, a few others. I try to constantly think about what is the non consensus thing today that eventually will be proven right and I don't have a super good answer to that yet. Um, so still kind of pondering and I think uh, go back to another point. I'm just excited to try to meet Problem owners, like product designers, designers who are inspired to build things with AI and, and can kind of supercharge this whole AI startup space into a new direction that's no longer kind, um, of rate limited by what AI research people think models can be useful for. Um, and then I'm kind of looking forward to whatever ideas I'll come up

Speaker A: with

Speaker B: during the year. I guess what's kind of humbling, a lot of things are humbling with investing, but one of the, the ones that continuously is, is that like, there's no playbook for how to do this job right? Because if there were one, everybody would use it and there'd be no value in it anymore. So, uh, you really do have to like, think creatively about, you know, how do I, how do I do something today that could like, team me up for finding something new in the future that would be exciting and, and being the person for, you know, whoever is the entrepreneur that's building against that opportunity. So in that sense, we always try to do stuff that really shows that we care about the domain more than anything else and want to be contributors to it. So I'm sure I'll come up with other ideas of how to contribute beyond trying to reform university spin out trying to reform government procurement of AI systems. Uh, do State of AI do Air Street Press. Do merch.

Speaker A: Yeah, the merch looks really solid. I'm excited to see more of what's coming out of Air street this next year. I've been very closely following this year, so it's been good to see that grow. Well, I think this is a good place to close in. As always, Nathan, thank you for doing this. It's very fun to end the year with these. This is like, always one of my favorite episodes to do.

Speaker B: Yeah, same.

Speaker A: And I appreciate the State of AI report. I appreciate Air Street. Thank you for doing what you're doing.

Speaker B: Yeah, for sure. I mean, it's great to have, um, friends that I can share this with. So thank you for that.

Speaker A: As always. Thank you for listening. Happy holidays. I do appreciate your feedback, comments, suggestions, compliments, if you have anything to say. I'm pretty easily reachable by Twitter email substack and I'd love to hear from you.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Utilizing AI internally to iterate faster and empower smaller teams to upskill w/ Vivek Raghunathan #263The Engineering Leadership Podcast · on GitHub Copilot96 / 100
  • Why Developers Are AI's Canary in the Coal Mine | AI For The C-Suite EP 76AI For the C Suite with Chad Harvey™ · on GitHub Copilot85 / 100
  • Small Models, Massive Wins: The New Shopify AI FormulaBeyond The Pilot: Enterprise AI in Action · on GitHub Copilot85 / 100
  • ClearGrid's Mohammad Al-Khalili on Building AI for a Trillion-Dollar Debt MarketFWDstart · on ElevenLabs85 / 100
  • EP 115: The Moment Claude Co-Work Replaced 4 Days of Work in 30 MinutesEmbracing Marketing Mistakes · on ElevenLabs85 / 100
  • 044 - Synthesia: Data Director - Why Data Teams Should Stop Trying to Be in Every RoomThe Stacked Data Podcast · on Synthesia79 / 100

More from The Gradient: Perspectives on AI

All episodes →
  • 2025 in AI, with Nathan Benaich77 / 100
  • Iason Gabriel: Value Alignment and the Ethics of Advanced AI Systems
  • Philip Goff: Panpsychism as a Theory of Consciousness
  • Some Changes at The Gradient
  • Jacob Andreas: Language, Grounding, and World Models
Explore the best B2B AI & Data podcasts →
All The Gradient: Perspectives on AI episodes →