The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The Gradient: Perspectives on AI
The Gradient: Perspectives on AI artwork

2025 in AI, with Nathan Benaich

The Gradient: Perspectives on AI · 2026-01-22 · 1h 1m

0:00--:--

Key moments - from our scoring

Substance score

57 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality11 / 20
Guest Caliber14 / 20
Specificity & Evidence12 / 20
Conversational Craft9 / 20

Nathan Benaich, founder of the solo-GP venture firm Airstreet Capital, returns for his fourth appearance to assess the AI landscape in early 2026. With a PhD in cancer research and bioinformatics, Benaich has built Airstreet to back AI-native startups across North America and Europe since 2018, investing in companies like 11 Labs, Synthesia, and Crusoe Energy. He articulates his thesis: that general-purpose AI represents an evergreen investment opportunity spanning industries and geographies. Rather than base solely in San Francisco, Benaich maintains geographic flexibility across Europe and North America, arguing this provides perspective unavailable in the SF consensus bubble and access to diverse opportunities in defense, biotech, and enterprise AI that others miss. He's notably positive on 2026, citing vastly improved capabilities in writing, research, web agents, and computer use, though emphasizing the industry is in an "instruction manual phase" of learning how to deploy these tools effectively. The episode covers his skepticism of export controls' efficacy against Chinese AI progress, the nuance of capability gaps (citing Epoch AI's finding of a ~7-month Chinese lag), and the emerging importance of reasoning models and inference-time compute.

Key takeaways

  • →Airstreet's solo-GP model and geographic diversity enable contrarian investments (defense, embodied AI, biotech) that SF momentum-driven firms miss, prioritizing founders with no alternative vehicle for their vision over those opportunistically starting companies.
  • →The 7-month gap between Chinese and US frontier models is concerning not because it's large, but because it's small - equivalent to two quarters - making export controls appear insufficient as competitive policy.
  • →Capability benchmarks mask consumer preference dynamics where users default to familiar tools ("Coca-Cola vs Pepsi") and quickly reset expectations, making it unclear which model advances will generate genuine market differentiation.
  • →Investment risk concentrates less on whether AI works than on valuation outrunning actual company progress and financial engineering in compute infrastructure buildout, suggesting downside is primarily a pricing problem rather than technology failure.
  • →Europe's macro geopolitical shift post-Munich Security Conference, with ~$1 trillion mobilized for defense, creates upstream venture opportunities in defense tech that would be invisible from San Francisco but visible from Europe.

In this episode

  1. 1Introduction and Air Street Capital Investment Thesis
  2. 2Geographic Strategy and European AI Competitiveness
  3. 3Comparison of Investment Philosophy Between Regions
  4. 4Progress and Sentiment on AI Development in 2025
  5. 5Range of Outcomes and Valuation Concerns
  6. 6China's Capability Gap and Export Controls
  7. 7Reasoning Models and Inference Time Scaling

Mentioned

Nathan BenaichAir Street CapitalDeepMindOpenAI11 LabsSynthesiaCrusoeAndurilHelsingDeep SeekNvidiaEpoch AI

Guests

Nathan Benaich

Topics in this episode

Embodied AI11 Labsreasoning modelsCrusoe EnergySynthesiaInference-time computeTest-time computeAirstreet CapitalDeep Seek R1export controls on China

Questions this episode answers

What is the current capabilities gap between Chinese and US AI models?

Chinese models lag the US frontier by approximately 7 months on average since 2023, as analyzed by Epoch AI, with Deep Seek's R1 matching OpenAI's O1 on benchmarks like AIME and MATH-500, but no Chinese model yet surpassing the top American frontier models.

How does Nathan Benaich's investment approach differ from typical San Francisco venture firms?

Benaich's solo-GP structure means investment decisions reflect his taste alone; he invests across diverse industries (defense, embodied AI, protein engineering) rather than following momentum consensus, and he's willing to work with teams that communicate poorly if they can deliver, unlike SF firms that weight exceptional communication highly.

What does Benaich think about export controls on China for AI chips?

He believes export controls will have limited effectiveness because selective pressure on capable systems drives innovation around constraints, and a 7-month gap is too small for controls to prevent convergence; he suggests collaboration for better visibility as a potentially more effective approach.

Why does Benaich maintain geographic flexibility between Europe and North America rather than basing solely in San Francisco?

San Francisco's strong tractor beam toward consensus makes original decision-making harder; other cities like New York offer greater industry diversity, and Europe's macro geopolitical shifts (especially post-Munich Security Conference) create upstream venture opportunities invisible from the US.

What does Nathan Benaich see as the main downside risk if AI progress slows significantly?

The primary risk is valuation exceeding what companies will actually achieve (a pricing problem), plus financial engineering in compute infrastructure buildout driven by rates and geopolitics, rather than AI fundamentally failing to work.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode surfaces a handful of genuinely useful observations - the 'instruction manual phase' framing, the 2/20 venture bifurcation, and geographic diversity of deal flow - but large stretches are conversational and high-level rather than densely packed with novel per-minute insights. Many beats are standard AI-year-in-review talking points.

it's no longer 2 and 20. It's either you're in the business of the 2 or you're in the business of the 20
I found that while one could be in San Francisco and sort of drink the Kool Aid every day, I found it like not particularly conducive to making original decisions

Originality

11 / 20

A few notably fresh framings stand out - the evolutionary pressure argument against export controls, the SF 'tractor beam' dynamic crowding out incremental ideas, and the sharp 2/20 venture distinction - but much of the conversation revisits well-circulated takes on China's capability gap, benchmark skepticism, and regulatory overreach without materially advancing them.

anytime you put a selective pressure on a system that's capable of evolving, it will. And so China is a country that's incredibly ingenious and has talent
the diversity of companies and industries that are represented in New York is far greater than San Francisco. You know, you'll have like defense media, finance, wealth management, consumer media

Guest Caliber

14 / 20

Nathan Benaich is a genuine, long-tenured practitioner - solo GP with a PhD in bioinformatics, running Air Street since 2018, with named portfolio companies at scale (Eleven Labs, Synthesia, Crusoe, Poolside) and author of the State of AI report. He is not an operator who built a product, but his technical background and investor track record give his opinions real grounding.

I managed to work with like 11 labs, Synthesia, Crusoe, a lot of different industries and great people
everything that is invested in at Airsuite is a result of like my taste and my taste alone

Specificity & Evidence

12 / 20

The episode uses real survey data (1,366 respondents, 92% productivity gains, 10% paying >$200/month), cited acquisition prices ($20B Grok, $1.6B Sambanova markdown), return multiples (6x Chinese vs 26x Nvidia), and a concrete operational claim (3-4x faster shipping). These are solid, though several macro claims are left uninterrogated and methodology is thin.

44% of US businesses are paying for AI. This was up from I think 5% in early 2023. The average contract values are up as well to 530k from 39k
a 6x return versus a 26x if people had invested on Nvidia at that day's stock price. And for US challengers it was 2x versus a 12x

Conversational Craft

9 / 20

The host arrives with genuine preparation - citing specific essays, reports, and numbers before each question - and occasionally draws Nathan into useful elaborations (the 2/20 observation, the data center NIMBYism prediction). However, there is almost no pushback, contradictions go unchallenged, and many transitions are mechanical ('maybe a next beat'), making this a structured briefing rather than a probing interview.

You've had various predictions over the years. I feel like you all grade yourselves quite harshly on prior predictions
the act was born in a world before GPT4 and will take effect in a world shaped by GPT7

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B69%
  • Speaker A31%

Most-used words

feel35model23safety23different22data21seems17general16last16models16europe15trying13feels13china13part12hard12chinese12

Episode notes

Episode 144 Happy New Year! This is one of my favorite episodes of the year - for the fourth time, Nathan Benaich and I did our yearly roundup of AI news and advancements, including selections from this year’s State of AI Report. If you’ve stuck around and continue to listen, I’m really thankful you’re here. I love hearing from you. You can find Nathan and Air Street Press here on Substack and on Twitter , LinkedIn , and his personal site . Check out his writing at press.airstreet.com . Find me on Twitter (or LinkedIn if you want…) for updates on new episodes, and reach me at editor@thegradient.pub for feedback, ideas, guest suggestions. Outline * (00:00) Intro * (00:44) Air Street Capital and Nathan world * Nathan’s path from cancer research and bioinformatics to AI investing * The “evergreen thesis” of AI from niche to ubiquitous * Portfolio highlights: Eleven Labs, Synthesia, Crusoe * (03:44) Geographic flexibility: Europe vs. the US * Why SF isn’t always the best place for original decisions * Industry diversity in New York vs.

Full transcript

1h 1m

Transcribed and scored by The B2B Podcast Index.

Speaker A: Hello. If you've been listening for a while, this is the fourth episode that I've done with Nathan Benesh. Rounding up what we've seen in the year in AI. Uh, as usual, we go through a few of the themes from the State of AI report. Also what's been happening for airstreet Capital, and just Nathan's general takes on issues over time. I hope you enjoy. I think you're like the one and only person I've brought on this podcast four times by now, which is really great. And I think this is easily my favorite, one of my favorite episodes of the year. So welcome back, Nathan Benesh. I feel like we will probably have a good number of returned listeners, but possibly also new listeners who are trying to figure out what the hell happened last year. So towards that purpose, I thought we'd begin the conversation by talking about Air Street Capital, what's been going on for you and I think also in particular some of the ideas around your core investment theses. So maybe we can start there.

Speaker B: Sure. Always a pleasure to do this and I guess it serves as a good landmark for where we've been and where we might be going. For me, my formal training is in cancer research, bioinformatics. Uh, I've always been interested in science and technology. I finished my PhD in 2013 and that was roughly around the time that modern day kind uh, of deep learning started to work. Maybe one or two years after that with ImageNet and deep learning, et cetera. As somebody who kind of grew up more through the scientific path rather than the finance, consulting, professional services path I've always been, or having more affinity for the technical crowd and like trying to find problems that you can't solve without new approaches, as opposed to this software product exists today. I'm going to make one that looks prettier and therefore like the customer will shift over or marketplaces or E commerce that were very popular at the time. And having started my career in London, where, you know, DeepMind was born, and having friends there and you know, being very close to like the academic community, I had the opportunity to see like this kind of nascent AI ecosystem, like grow firsthand. And I took the view like many others who decided to spend their life on AI. While this technology looked niche and maybe a bit toy at the time, by definition, everything that exists in the world is naturally produced as the result of intelligence. And so if you can have a system that can generally learn any task and probably over time start with simple ones and gravitate to other ones and go into new industries, et cetera. Then it's this evergreen investment thesis which you can ride from being niche to being ubiquitous. So that's like, let's say why I'm interested in this area. And I uh, had the opportunity in 2018 to start something new. And still at that time most venture firms were not particularly interested in AI. Like there was no real evidence outside of like bytedance or Google or Facebook, et cetera, kinds of monster companies producing lots of money in AI. And so I set out as a solo gp, which means I run everything myself and make the investment decisions myself to build airstreet as the venture firm that uh, you know, could support entrepreneurs building or making use of AI in their startup and do that across North America, Europe, and go where industry opportunities look like they'll be ripe, you know, in maybe two, three years time, but don't look super consensus today. And I'm now in my third fund and it's been an exciting time. I managed to work with like 11 labs, Synthesia, Crusoe, a lot of different industries and great people. And yeah, other than investing, I spend my time like reading, research and trying to uh, stay close to the community and speak to people like you and others. And so whether everything goes pear shaped for a bit or it's even more exuberant like I, I'll still be around, still be investing. Yeah.

Speaker A: One of the things I think I appreciate most about your work in public advocacy, I think is just how sticky you are, both with the sticking with your thesis, but also you've been doing a lot and we'll get more into this in detail later, but the high level of wanting Europe to be more competitive in the AI space and we'll talk about things like regulation and all of this later, but a question that makes me think about is obviously you're able to invest in American companies as well and you spend time in the States and things like this, but you're largely based in Europe. I think you generally identify as a European venture capital firm. And one option you could have or think about is, well, Europe was clearly not as competitive as I might want it to be. Why don't I go and base myself in the States or something like this and not deal with it. But clearly you're taking what to me feels like a higher friction route, at least in the. I'm going to base myself here and really try to figure out how do I get Europe to be a little bit better on these fronts that I care about, what animates you there.

Speaker B: I've taken the approach of trying to be geographically flexible and split the year in various places in Europe and in North America. And that just means you have to be more able to just jump on a plane when you need to be. But I think because the nature of this industry has expanded so much, it's more diffuse than perhaps it was before where there was a clear center of gravity in London. And then um, post OpenAI clearly one in San Francisco. And then you could argue there's a good one in New York. And so you do have to be in many places at the same time. The other thing is there are different ways of doing this job and who knows what the right one is. The result is the result. So I sort of learned to try to do things in the style that leads me to make better decisions and to be around the right places. And I found that ah, while one could be in San Francisco and sort of drink the Kool Aid every day, I found it like not particularly conducive to making original decisions and not being able to have a bit of sort of perspective, like global perspective on things. For example, I found San Francisco's fantastic place to discover the new new thing and harvest all energy that is not in the new new thing towards the new new thing. And so the incremental odd idea feels like it's less commonly coming out of SF than it is coming out of, you know, other cities where there isn't maybe so much of a strong tractor beam. For example, when I do the meetups that you've been to in SF or New York and I sort of look through the attendee list, the diversity of companies and industries that are represented in New York is far greater than San Francisco. You know, you'll have like defense media, finance, wealth management, consumer media, whatever and sf. It's like all the obvious names that we're all aware of. So I think for that reason it's good to have like geographic perspective and diversity. The other part around being long Europe. I think Europe has always had like a lot of strengths that probably don't need to be necessarily repeated here because they are repeated ad libitum. But like generally speaking, intelligent people like opportunity, like good companies, etc. But it's like generally held back by like different tolerance of risk, fractured regulations, political differences between different countries, et cetera. I don't want to necessarily like anchorage my success on the overall outcome of European policy making as it relates to technology and company formation. But I do think doing what I can to opine on topics where I can Potentially have some influence. Whether that's biotech or AI or defense is a good use of time. As we stand here, it's like a month away from the Munich security conference in Germany that famously last year led to weakening uh, of US security guarantees in Europe and the subsequent realization of holy shit, no one's going to come to save us if we need it. And then you know, close to like a trillion of new money mobilized to um, basically beef up European countries ability to defend themselves. I mean this is like such a macro shift I don't think we've seen since like World War II. And so I think having exposure to Europe means that you know, I can play in that sort of macro game and say like oh, there's a lot of new defense companies that need to be started and if I were sitting in San Francisco it'd be like a lot harder. So I think just in the same way that it's nice to be able to move between different industries and themes within like an overall umbrella. I think having the geographical flexibility also lets one go where the opportunity is and potentially go there ahead of it being consensus.

Speaker A: Yeah. Do you find yourself when you talk to other VCs ones who are based in SF for instance. Clearly you've illustrated some of the differences that come out in maybe investment theses or the range of companies you're interested in. Do you feel like you notice or feel as though you have other differences as far as investment philosophy, how you engage with companies? Um, I'm curious how you'd sort of self describe the you versus SF VC at uh, maybe like a well known firm or something.

Speaker B: Yeah, one obvious one is everything that is invested in at uh, Airsuite is a result of like my taste and my taste alone versus most other like 95% of other venture firms that are not solo GPS. It's hard to attribute whose taste is it that led to the investment. The second thing would be I think in SF it's increasingly like a momentum game where people assume they know who the best entrepreneurs are and it's basically a bidding like auction to get access to those individuals and they are you know, in most cases like worth it and you know, exceptional. But like the information asymmetry seems very small and so the ability to make outsized returns unless you assume that the outcomes become so enormous is hard to see. I think there's then just the diversity of things that one can invest into. So I have defense companies like Delian in the UK and Greece doing sort of like a version of Anduril or Helsing. I have a business doing embodied AI before people really thought that that was going to work called Syriact. You have one doing protein engineering, design, genome editors and sf. So that kind of that diversity I don't think one would get by just sticking in one ecosystem. Maybe the last bit is I'm like more willing to take certain risks where I think they can be mitigated with good work. For example, there are certain companies that do exceptional work but don't communicate it super well. And therefore for somebody who's used to seeing brilliant communication all the time, that maybe lacks actual substance or ability to deliver would underweight the former. And I'm willing to like see through that and understand like, yeah, some of these things we can actually work out and fix. But you know, if a team can't actually deliver on what they say they're gonna deliver, that's far worse than being bad at communications. But yeah, I mean, as we said before, like, there's lots of different ways to do this job. Like I'm trying to find entrepreneurs who feel like there's no other vehicle through which they can express their skill, like their passion, how they wanna spend their time than doing a startup versus what I feel like there's a lot of entrepreneurs that kind of kick off companies now because of politics or because of employer issues or because it's too easy to raise money. And I'm just not sure that all those reasons are like long term sustainable.

Speaker A: Yeah, this makes sense. Before we go diving into like really specific parts of your findings in the 2025 State of AI report and what's happened in the three months since then, what is your general. We're recording this in, um, earliest January 2026. Just how do you feel compared to roughly a year ago?

Speaker B: I think the progress in the last year has been really momentous. I think it feels like on several tasks that are at least important to me, like writing research, thinking about ways to understand new things and learn even some of the web, agent, computer use, work, image, video, all that is vastly better than what it was 12 months ago to the point where I think it is genuinely useful. And now the real issue, if you were to just freeze progress, is just, uh, how do you figure out the best way to use these tools? And the industry has released artifacts that are so capable if implemented the right way, but without an instruction manual. I think we're just in the instruction manual phase. I'm certainly surprised, I guess, at the scale of investment that's going into this and just keeps Going into it it really feels like yeah, if AI doesn't work, this is more than problematic for private markets, more than problematic for public markets and so problematic for nation states at this point. And so that's way stronger than what I thought maybe two, three years ago when there wasn't very much to hope for in the economy and thank God AI emerged. But this is next level. And so I think for 26 I'm generally very positive. I think there's still so much more to go. Even just looking at sort of like basic principles for why people were excited to invest in technology. Like hey, cloud spend is still pretty small spend of overall IT budget. E, uh, commerce penetration is still relatively small compared to where it should be. Everybody has smartphones, all those things. Now you layer AI on top of that and you get even more opportunity. Whether it's like AI driving more conversion on commerce websites, we haven't even started to try ads. Enterprises are still figuring out what really works and there's already good evidence that there's margin improvements and revenue opportunities and things of that nature. Uh, so yeah, I expect a lot of this year.

Speaker A: One way you could take the scale of investment observation you just made is that if things work quite well then what could happen is we get a state of pretty serious abundance or something like this and otherwise things could collapse in a pretty terrible way. And it feels like a very wide range of potential outcomes that feel on the table for quite a number of people who are thinking about it. I'm curious for you personally, do you feel like you have a, ah, sense of in your head just what the potential range of outcomes and what feels likely less likely to you is?

Speaker B: Well, because today's generation of systems are really, really useful compared to the last generation. I don't think that if progress were to uh, stop and things were to go bad that the artifacts we would have left are not useful. I think the worst case is that the prices of assets that people are paying for now is way outrun compared to how far those companies will actually get. And so it's a valuation problem. And the second problem has been discussed a lot, the sort of injection of financial engineering and debt and off balance sheet like SPVs and things of this nature. Sort of like radical financialization of build outs, uh, and energy and GPU data center creation. That's I think the issue, but that seems less tied to, you know, is AI going to work or not going to work and more to do with things like rates and geopolitics and, and financial engineering topics. Which is a little bit outside of the reach of the AI researcher who's trying to make GPT. Whatever.

Speaker A: Mhm. Maybe a non obvious segue there, especially related to the valuation and finances questions. As last year kind of started off with a bang, there was a little bit of a deep seek moment and a lot of trouble in the markets based off of this. And in particular we saw just the theme of China closing in on US labs in particular with deep seq R1, their reasoning model matching 01 on a number of key benchmarks. This included Amy, the Math competition in 2024 and the Math 500 benchmark and epoch AI analysis sort of pointed out that Chinese models seem to be lagging the US frontier by roughly seven months on average since 2023. No Chinese model yet, quite surpassing the top of American frontier models, but still a pretty close race. Again, many other dimensions of this that I do want to get into later, but just on the front of the capability gap we're seeing and how that closes, how are you thinking about that and especially how it relates to the ways that you might be thinking about what to invest in and things like this.

Speaker B: The seven months doesn't seem that long.

Speaker A: No, it doesn't.

Speaker B: It really doesn't.

Speaker A: Right.

Speaker B: It's like really two quarters. So on the one hand it seems like capabilities are getting a lot better if you assess benchmark numbers. But I'm very cognizant that when we're looking at a benchmark, it's a number on a table and we're trying to eyeball just how much of a delta 81 is compared to 74 is. And for most of us who don't work on benchmarks every day, there's not really a good intuition for this. And then there's many examples of people who look at the benchmark like, oh, it looks really great. They try the model and they're like, oh, it doesn't work for me or I don't think it's that good, or I have another system that I prefer. So I think what seems to be happening is as the number of use cases that one tries AI on grows, then there's inevitably like certain preferences of tools for certain tasks and it's a little bit like consumer preferences in that regard. You have like a workhorse model and you pick your vendor and if that one we call Coca Cola is not available and you have to drink Pepsi, then most of the time you'll probably say, okay, fine, but sometimes you just won't drink it. And I think every time There's a new model that comes out, we try it, uh, and then you just convert back to like using your Pepsi or using your Cola. The other thing I noticed is just, just like how hard it is if you not even looking at like the benchmarks, if you just try one versus another in like a chat window, how hard it is to really feel a difference. Sometimes I think that's just a lesson. There's probably some scientific or behavioral psychology behind this. It's well studied, but it just looks like humans reset their expectations extremely quickly. Whether it's like forgetting recent historical events or it's forgetting where we came from, or it's now thinking that something new is not that new. And if you're just running back the clock, a couple of years ago, you'd have been like blown away with any one of these systems. So I think all those things to me makes me wonder what is the next major result that would make someone's jaw drop. And I asked this with my friends and it's sometimes not. I don't really get like a clear answer. So in that sense it's hard to predict which model will get better. And then I think while most of the model companies started out generalist, it looks like there's some bifurcation or some clear product bets that are being made that would eventually mean like that system is very good at, uh, for example, code, but decays on its ability to write. And that's just expressing the corporate priorities of the company.

Speaker A: Yeah, this makes a lot of sense on the China beat. Specifically, a sort of big question and worry that came out of this was what is going on with the export controls that were put on China. Clearly a lot of people in the United States are very invested in competitiveness in this question and don't want Chinese models to catch up. And there's suspicions findings about high flyer accumulating many GPUs. Nvidia has some thoughts on this, things of that nature. To your point, of course, it is a little bit hard to see sometimes the differences here. I've played with Kimmy K2, but it never really stuck around for me. For instance, even though I was quite impressed. Do you ever feel worried about it or what's your kind of perspective looking at, uh, all of the discourse around? We should be really worried about this.

Speaker B: This. Yeah, my simple take is a little bit inspired by biological evolution, which is that anytime you put a selective pressure on a system that's capable of evolving, it will. And so China is a country that's incredibly ingenious and has talent, has money and resources, et cetera. It doesn't surprise me that if you put a selective pressure saying you will not have access to X, then there will be innovation around X separate from whether these reports of chip smuggling are true and the probably to some degree are. So in that sense I'm m not really convinced that export controls are going to work or maybe they slow them down like a little bit. But as we've seen it's a little bit is not enough if seven months is really where the gap is. And so it feels like another approach could be just try to keep your enemies closer and so some form of like collaboration so at least you know, like you have a better visibility on just how good or bad or how far away they are from you. And I think the other part that's interesting is just how quickly like Chinese companies have rushed to public markets. Like in just this last week we've had two IPOs of Chinese model companies, you know, minimax and then I think it was actually Zai, as it's called GLM makers. And I think that's interesting because it captures the zeitgeist. Quite a lot of international investors can invest in China or for American, you know, Robinhood users or something, it's a little bit more difficult. But I think at the end of the day money talks I think more than almost anything. And we can see this with certain corporate behaviors of like American Labs on their value values and shifting around. And so at the end of the day if that's true, then if an investor is looking for access to exposure to AI, are they going to really, really, really not invest in CAI because it's listed in Hong Kong. I do want to follow like where invested capital in the Hong Kong exchange actually comes from and I wouldn't be surprised if some of it does actually come from the US because for the pure reason of like there's an investment opportunity there and you can see like how Chinese equities related to AI have actually performed pretty well in the last 12 months after they've really entered the game. And all of that's not just going to be like Chinese money that's propping up the prices.

Speaker A: This makes sense. I think maybe a good next beat or sort of area for themes of last year was reasoning in general and this is referring to test time compute also referred to as inference time scaling. Again the computational power used during inference, the thinking traces that you might notice when one of your favorite models is preparing to give you an answer, maybe self correcting things of this nature. And there have been a lot of I guess interesting findings on how this seems to work and of course a lot of inference companies and um, places that are really trying to optimize that uh, process of inference. Many different models this year that are really focusing on things RL kind of having a moment again. What are sort of like the high beats of this year for you on that front?

Speaker B: Probably that just when we recorded this last 12 months ago, I think that was the first nugget of 01. And so just before Christmas everybody's like oh my God, let's see where this is going to go next year. So I think on the one hand it's are the chain of thoughts actually truthful? Do they really reflect what's going on? Or is it a model describing steps that it believes we want to see? Almost. The second one would be whether they are actually good or not good vignettes into model behavior and can be used to tease out if the reasoning is good quality I guess because in some ways do we care if the system reasons in the way that we think is best to reach an answer or does it just have to find a new approach to get to the right answer? And at ah, least to me at a high level, one of the most interesting parts of machine learning is this global optimization idea of maybe we've just found what we think is the local optimum. But a system that can consider everything can find a better one and that doesn't mean it's worse. The other part has been uh, sort of like user behavior is. I don't know if users really know which model to select and how much thinking time they should use. It seems like on some queries which did return very quick results previously, now the model thinks and uh, it's like why is it thinking when it could just return a result really fast before? But in general like I think it's been like a big change that has probably unlocked also some advances in AI for science. Like whether it's those literature agents that are considering huge bodies of work and considering what are the different hypotheses, what might I test like writing huge code blocks, executing them, et cetera. It seems insane to look at that kind of work and think like 12 months ago we didn't have that.

Speaker A: In the general case one of the things you were gesturing at was this question of are these reasoning traces good and are they actually helpful for us to understand model M behavior? And two direct I've sort of seen or two of many directions is one this is uh, an anthropomorphization. We're sort of grafting our sense of how reasoning works onto these models. And you know, this isn't necessarily representative of what's going on. 2 Is this question of chain of thought interpretability as uh, a safety question? If we can't interpret and understand what's going on, then maybe that makes it difficult for us to have the kind of oversight we might care about in making sure that these models are not wireheading or trying to get around and deceive us and things like this. Do you maintain any of these sorts of worries?

Speaker B: Yeah, I wonder also around this monitorability tax of is there a trade off between being able to see and read what the model says it is thinking and what it is actually thinking? And does that trade off if I can ask it, reduce the quality of the results or not? I do actually read the reasoning traces on certain problems just because I'm curious to see is it considering things in the way that I would consider or what sources is it looking at? How is the system maybe working? I've noticed some differences when I upload a photograph. Oftentimes the model is now calling an OCR package and then running the OCR package and be like oh, there's a problem with this, I can't run it. Whereas before it just felt like it natively did it. But there's that. Uh, and then on the safety part I think relating to can you trust the chain of thought, I think the result that I found most concerning was it's just like pretending to be aligned when it's not actually aligned. That seems concerning. And this self preservation behavior and the rest of it, again, it just feels challenging to know is this genuinely, genuinely what the thing thinks or is this gaslighting but regurgitating?

Speaker A: It's also kind of challenging to study. I feel like we've seen a few what I think are generally pretty good and insightful studies out of the likes of Anthropic that have put these models in a purported scenario and gone into what is the model thinking about, what is it trying to do here. But they're also contrived enough and it's a little bit hard to measure this without intervening on the model in some way. So I feel like it's also a little bit hard to make too much of those results.

Speaker B: Yeah, the problem is people who write about it make a big issue of it and then that can propagate certain feelings that might influence policy and the rest of it. But it does feel like related to that question of can you trust the model, is it faking? And perhaps this extends into cybersecurity risks and other things. That cyber stuff seems perhaps more concerning. If we extend the ability of the models to do these complex reasoning and write approaches to hack systems and work for extended periods of time, I think we have potential issues on our hands, which was one of the predictions that we made in the report.

Speaker A: Yeah, this seems very plausible, I think, especially with the number of people who are adopting these systems. I feel like Jan Ligo was quite early on this worrying about people integrating language models into their critical infrastructure and just all, all of the various vulnerabilities I could create.

Speaker B: Yeah, yeah. Particularly as some of these models can be compressed so much. And so as a payload, it's not unforeseeable that you could inject this onto someone's machine or onto like, uh, critical systems that are not necessarily. Well, that are like, maybe have to do with defense and things like this.

Speaker A: Well, speaking of all the machines that could be affected by this, we've been seeing a lot of commercial traction this year, at least accelerating. There's some Ramp data from the year claiming that 44% of US businesses are paying for AI. This was up from I think 5% in early 2023. The average contract values are up as well to 530k from 39k. This is like a sample though of 30 to 50,000 businesses using Ramp's platform. So I think this is skewing towards tech forward, small medium businesses. There's a U.S. census business trend that found that 8.4 to 9.7% AI adoption numbers using some broader survey methods. I'm curious just between those findings, others you've seen how you're sort of seeing this. Where does your head go when you think about maybe the differentiation between those two numbers in the US but then also other places you're investing.

Speaker B: Yeah, I think it has to do with the sample pool. As you mentioned, more tech forward companies in one survey and more general companies in the other. I think a leading indicator is these new companies getting formed. Are any of them starting without using some kind of coding agent? Pretty much. Absolutely. Not everybody's using it. And from just anecdotal evidence of talking to some companies in the portfolio, they're shipping products three to four times faster with fewer people, particularly for the more prosaic SaaS application that has a very well documented playbook and has the primitives and libraries, et cetera, that are on the Internet and in training corpuses as Opposed to the more esoteric brand new stuff which is probably a bit hard and more out of distribution. And then in our own survey data like you see this is of like a thousand or so or a little bit more than a thousand people on the more like educated level of the spectrum. So probably a bit more like AI people in the US and in Europe. I mean a lot of them use it for meeting notes, for like content generation, for just brainstorming, for coding and then increasingly for like financial analysis. And major labs have made I think a big push on this and data analysis more generally. Yeah, the startups are just showing us where things are going. And I know from other surveys that uh, equity research groups that banks have done with large enterprises spend there is also going up materially and the top customers of OpenAI is uh, reverse engineered from the demo day token prize winners. It's quite diverse. It's not just your most forward looking technology company that's a big spender. So I think that's very positive as well.

Speaker A: Yeah, I feel like it's hard to give just a singular story over all of this. I mean to your point it seems like there's quite diverse adoption but obviously maybe some sectors are going to find more uses than others. That's pretty clear to everybody I think.

Speaker B: Yeah, but I was surprised in our surveys. So over uh, half of people pay more than 20 bucks a month and 10% of people pay more than 200 bucks a month. And like 92% of respondents reported increased productivity gains from using it. And the ones that, that found more productivity gains were ones that paid. Don't know if it's like causal in some way. Maybe you get like a better service or they took it more seriously or whatnot. And generally like expectations of spend were going to go up in their organization like either slightly or significantly. If you sum those two, it's over like two thirds of respondents. I think we've as an industry probably converged on. It's going to be like slow takeoff because human inertia is real and folks are like realizing that, that you know, if they don't want to use something, they just won't use it. And it takes some convincing, it takes some education. We probably need some kind of genius bar to teach people how to make use of AI, particularly in enterprises. Just give them a dashboard and hope for the best.

Speaker A: You were mentioning a finding from your adoption survey that was uh, a new thing this year. Do you feel like in general the findings from that survey you mentioned just one of them but was There anything else that surprised you? Do you feel like in general that gave you a clear or maybe a different picture about adoption from the one you had, or did it feel like pretty consistent with the way you were thinking about things?

Speaker B: Um, it was more bullish than what I was expecting, particularly on the paying. That many people paid out of pocket. Some people, you know, paid by employer. But the fact that people were paying and not skimping out and then just like the diversity of use cases, I mean there's like over. I think there's some, probably some startup ideas like embedded in this data somewhere. But like, you know, there's like 1,366 people who answer, what are your motivations for generative AI? And some poor soul wrote survival. But you know, it's like more output, faster research, more brainstorming, automate daily tasks, processing batches of information quickly, paper reading, do more with less prototyping, outsourcing, boring things like email, creative generation, quick summaries. So I think it's really like a behavioral thing of just learning how your workflows could change. Even things like having to write a weekly summary of what you did or what you did on your team. If you're recording all these meetings and recording like internal things, like you can get an AI to generate that for you. Why should you ever transcribe a call anymore? Why should you do anything other than like text to speech? If you need to do audio creation, analyzing like any kind of financial data, it's. Or even like cohort analysis or things like this for startups, I mean just dump it in chatgpt. The thing could do a pretty damn good job. Huge time save.

Speaker A: Yeah, this makes a ton of sense. Maybe on the opposite, well, not quite opposite end of the spectrum, but a larger part of the spectrum from individual users. We've been seeing what you sort of dubbed an industrial era of AI. We obviously have Stargate as a prime example of this. But then also the XAI data centers. There's a really good two part podcast on this from the PJVote search engine. They did a pretty good job sort of covering what's going on with those data centers, how they're affecting the areas they're in, all the things like this. What do you think about when you're looking at just the scale of investment and what's going on with those?

Speaker B: Well, first of all, I think it's going to lead to way more usage. I do subscribe to the thesis of the people who are saying we need these systems, that it's not just because they want to have more flex in the system, but because they are making compute trade offs and we need this to scale deployment. I think the part that surprises me is just the limited energy resources that we have to power these things and the resurgence of gas turbines which are not particularly the most eco friendly nor are particularly nice to uh, be right next to. And so there's quite some documentaries of towns in Pennsylvania or in Virginia that have these brand new data centers and they're just loud as hell and their home is shaking because see the home is not built out of concrete but probably more like plywood or something and it just doesn't seem like a nice existence. But another part is it is driving a resurgence of employment and industry which is positive. I think it's driving also second order effects of investing in the national energy infrastructure and grid and production which is long needed anyway in most countries. It's not even just in the us. The UK desperately needs it too. It has the highest energy cost of any European country and suffers the consequences of that geopolitically as well. And the other part I wonder is if you are a serious AI lab or want to be a serious AI lab, it does seem relatively inconceivable that you can do so without owning the models, data, the compute and access to power. And there's a few companies that have outside of OpenAI and anthropic and Amazon, we have Poolside and they've pretty publicly started working on this a few months ago. And that does buy you a uh, totally different economic cost structure and ability to iterate and control your fade and also has benefit for customers who want to have that resilience. Seems like everybody's doing it. I mean Mistral to some degree is doing it. I don't know the extent to which Chinese companies are certainly in the west,

Speaker A: this was all sort of already a theme last year, but I remember when we recorded this a year ago we were sort of talking about the inevitable increases in GPU capex and just what that would save the company balance sheets and the scale of investment you need and things like this. And I do remember seeing Poolside raising quite a lot and that amount of GPU spend. And I guess for you as an investor also I think that clearly impacts the valuations, the needs of the companies that you might want to invest in. Do you feel any. Obviously not the case with Poolside since you're an investor in them. But do you feel any reticence or any different thoughts on just the structure that this implies for companies you might want to invest in for sure.

Speaker B: And I don't really know what the right answer is. I've sort of voted with what the right answer is based on what I've invested into. But yeah, uh, it's challenging to live in this situation where the gp, like me sort of lives in the present, the entrepreneur lives in the future, and then the capital sources into venture funds live slightly behind in time compared to the investor. And we all have to sort of align in some way to be a productive marketplace. On the one hand, when entrepreneurs are pitching the fact that compute is going to replace labor and models need to be trained and that's expensive and data is expensive and RL makes it even more expensive and simulation does too, and world models and all the rest of it, it's all very exciting. And that, uh, generality is what one should strive for and not specialism. That almost sets specialism as sort of like a lower tier problem. Because if you can solve generality, why would you bother spending any time on specialism? M this sort of narrative then produces very large fundraising rounds, which I feel like are fit for a certain type of financial firm. And that, uh, platform is quite different from what venture was set up to be in the first place. I think taking risk is completely fine, provided that you have a lot of potential upside. But if financing rounds are so large and implied valuations are so high, we can't all end up OpenAI or anthropic, um, existence proof of two companies doesn't mean that one should apply the same logic to everybody. So I'm in a little bit too two worlds with this, and I think I'm not alone in that because even the big firms are strategizing financing rounds to have it both ways. Where one round might be announced at 4 billion valuation, but actually was three tranches and started at 200, then 1 billion and 4 billion and different people bought it at different prices, which I don't think is particularly healthy either. So I don't know where it all goes, but I do think it'll continue because there's just so much opportunity and so much capital.

Speaker A: There's something really interesting in that answer you mentioned, which was the way this financing is happening seems to be quite different from what venture capital was set up to do in the first place. Do you feel like you see or expect that drift continuing to happen, or do you feel like there's any sort of identity crisis going on? Or just what is the shape of that observation that you're having?

Speaker B: It's related to how venture firms in the last five or six years have either decided to stay kind of stage appropriate or boutique, if you will, which generally means, I don't know, a couple hundred million per fund, if they're doing early, or if they've decided to scale to become like these aircraft carriers and have multiple stages, multiple strategies, then that becomes more of like an asset management business. You could argue. Uh, one entrepreneur told me this once as he was explaining it, uh, the venture business was about, broadly speaking, 2 and 20, which refers to the management fees and the carried interest. But actually now it's no longer 2 and 20. It's either you're in the business of the 2 or you're in the business of the 20.

Speaker A: Interesting.

Speaker B: And I think there's some truth to that. But at the end of the day, people will argue their own book and the strategy that they can execute on based on where they are in their career, in life, et cetera. I, uh, do generally subscribe to the view that the role of the venture capitalist is to support, empower the vision of the entrepreneur. And if the vision of the entrepreneur grows because there's opportunity for it, there has to be some meeting in the middle. You can't just, as investors say, I'm not doing that because that's not what this business was set up to. And I'm going to stick to doing your thing. You have to change and you have to update your priors. And so then we get back to the, you know, we're living in the present, entrepreneur in the future, and the LP slightly in the past. And so we all have to like work together to make sure that the system is healthy. So to make it clear, I'm investing more money, I'm raising larger funds to do that and providing more capital to entrepreneurs because you got to play offensively. I don't want to get stuck in the past by any stress.

Speaker A: Maybe a next B on the state of AI report findings was you had some really interesting notes on the returns on investments in video challengers. And in particular we've seen a, uh, couple of acquisitions happen. So Grok acquired for 20 billion. Sambanova acquired for a, uh, markdown of 1.6 billion. And you sort of talked about in the report the sort of return to various challengers. So for example, if you had invested in Chinese challengers, you would have had a 6x return versus a 26x if people had invested on Nvidia at uh, that day's stock price. And for US challengers it was 2x versus a 12x return on Nvidia had they invested at that day's stock price? I guess. One, what do you make of the difference between Chinese and US challengers? And then two, what do you make of just the general. What's a pretty large difference in challengers versus Nvidia and the challenger strategy here in general?

Speaker B: Well, the China one is again interesting. Related to the point earlier of those chip companies like Camerocon rushed to the public markets a lot faster than what they did in the US and so probably the contributing reasons for that Delta is that it was a public company and you could get access to next generation chips in the AI business as a purist play by investing in that business compared to a general purpose chip maker. When it comes to just challengers in general, I always found I did subscribe to the fact that we needed to have new silicon for AI because the GPU is not specialized and this was in 2018 or something like that, or 17. And there was probably a window where startups could execute against that. But I think that window has really narrowed, uh, over the last few years because of all the ways that Nvidia can rapidly listen to its customer base and update its designs. I mean that was one of the reasons why it was well positioned to do deep learning in the first place of being close to researchers who kind of leading them by the hand into new directions. The other thing is it's a little bit hard today now that we see some contenders seeing being bought for some significant amounts of money or even public stock contenders seeing their stock price grow. It's hard to look at that and say, oh yeah, that thesis of there should be contenders has been vindicated. Like Nvidia is not the only game in town. I think what's really happening is we just need so much of this stuff that one vendor can't provide everything. So we're going to go to the number two vendor and the number three vendor. We've seen this in energy where public companies that don't even have a product for the next couple of years are signing monster deals with public tech companies and seeing their stock price grow up. That's highly speculative. Just because these companies that are building and scaling out AI need to have access to everything and more. So it's like the tide is lifting all boats more than it is that there's genuinely good competition. To Nvidia, it's again like the Coca Cola and Pepsi. It's just like now the thirst is so high that all the Coca Cola cans get consumed. So now we're just switching over to buying Pepsi.

Speaker A: Yeah, this is interesting. I feel like also in that domain you were mentioning your earlier thesis about chips not being specialized enough and obviously you've seen in later generations of Nvidia hardware things like tensor cores and stuff like this. And I feel like also a really interesting thing on that front just has to do with the relationship of hardware to AI architectures in general. Sarah Hooker, who was at Google Brain at the time, wrote the Hardware Lottery. This essay from a while back that was pretty widely read about the way that the transformer architecture wasn't the best AI architecture in general in a vacuum, but rather the best fit for the available hardware, makes use of the parallelism. That GPUs offer clearly makes a ton of sense. And then you see a lot of the challengers there going after the same thing like using systolic array architectures, can we make matrix multiplications faster, things like this. And so you sort of see uh, a doubling down or intensification of uh, even in the challengers feed, GPUs not specializing and being good for different architectures for the most part, but just being good at roughly the same family of architectures. Do you feel like you, when you were thinking about this, felt like you wanted that specialization to more be in this direction or do you feel like. I think some people who might be thinking about wanting AI development to be a little bit more creative or diverse are not super happy with the state of affairs and I'm curious where you are.

Speaker B: Yeah, I was imagining more the trade offs of internship bandwidth of memory, movement of should the die be larger and attributed more towards deep learning, towards matrix multiplication versus graphics, should you put more memory on it or have a different on ramp? More of that stuff than believing in neuromorphic or other substrates or even biological neurons. I didn't spend that much time on the FPGA topic or custom ASICs specific to certain workloads because I think even at the time there was a lot of fluidity on which architecture is the best for the task. But as you all converged generally on the transformer, it looks like the custom ASICS are making a comeback. That's judged even by Broadcom's large deal that they've announced the major company that's facilitating the creation of asics. And yeah, having said all that, every year we look at all the AI, ah, research that's published in the public domain and then which chips are used, but it's still like 90 or 95% Nvidia, you see little inflections in TPUs and AMD and some degree Apple silicon. I think people are using a of lot. The laptops are super powerful nowadays. But Nvidia is still so far ahead.

Speaker A: I think the next could be is on the regulation and AI agenda front. And you had a few slides on the Trump AI agenda as well as Nvidia's influence and sort of the patchwork state AI regulation and what the Trump administration has been doing about that. And I feel like the three maybe key parts of this were the executive order that came out last January revoking the Biden AI safety order, mandating this AI action plan, the sort of July preventing woke AI idea and then in December this national AI framework that was creating a DOJ task force trying to challenge state AI laws. I feel like there's a lot of different directions there and questions around how do you do regulations, um, should this be like a federal or state question? And obviously different companies maybe have different interests here, but how are you thinking about the trade offs there? And do you feel like also in that do you see any parallels to or evidence around your feelings about say the EU AI Act?

Speaker B: Yeah, I think regulation with regards to AI should be more like existing regulators are upskilled to be able to understand how their purview and how their tried and tested policies extend towards AI or should be adopted or adapted for AI versus a new blanket policy that everybody has to adopt because at the end of the day the medical regulator or the aviation regulator, et cetera, knows the best about the risks that ah manifest in that domain. Far more so than some AI policy person extrapolating basic principles to risks in that domain. And so by that extension two, I think individual states dreaming up what they should implement as rules feels probably not optimal just because I don't think we figure out what the right rules are in the first place. And so 50 different experiments on such is going to introduce a lot of stochasticity. And it seems like the Trump administration is taking some efforts towards banning state level regulation perhaps for that reason. And in the eu I think the main lesson is just overreach and then also trying to lay down found regs that are future proof when progress is too fast. I think the classic mistakes were putting compute thresholds like arbitrary compute thresholds of model trained with more than this is dangerous and must be inspected or model, you know working in biology must default be inspected. Like how's uh, any of this like really practical to monitor any good regulation like has to be able to be relatively easy to Monitor in the first place.

Speaker A: Yeah, we can maybe jump into the EUA act for a second. Max Cutler had this piece with you about it and one of I think the insights there was this sort of lag. So a quotation taken from that was that the act was born in a world before GPT4 and will take effect in a world shaped by GPT7. The core thesis there being, of course, that Europe's regulatory initiatives might be obsolete before they're even fully enacted. And there's a couple of other criticisms there. The fact that only three EU member states were compliant as of late 2025 leaders calling the act confusing various industry response. How are you thinking about the act since then, since you wrote that?

Speaker B: I'd say most of the arguments. I think the crux of the argument remains the same and it's further vindicated by authors or key authors involved in the act feeling like it went too far. I do think though that the mandating of regulatory led sandboxes, which is what we just discussed, is a good thing. But general reporting requirements and compliance is just always a drag. And perhaps it's worth even just anecdotally looking at how many more companies have been founded or talent that's gone into Europe versus out of Europe. And just anecdotally, it feels like more people have left than companies have started. As you mentioned, very few countries are even up to speed and properly implementing the act in the first place. So I don't know, I'm just generally not super excited about its potential. No one's really looking at it saying, yes, Europe will have a chance of competing on the basis of having better regulations. The only people that are saying that are the bureaucrats.

Speaker A: Maybe another related frontier is there's been the notion of sovereign AI thrown around for a little while now. I think Jensen Huang in particular had this idea and you were quoted, I think somewhere in the Economist as saying that most of these efforts are probably a waste of money. If you were advising a non US government on AI strategy, maybe both, both related to this and the regulatory landscape, what would you say to them?

Speaker B: I think I would say form alliances with countries that do have a reasonable shot of being sovereign. Uh, and to me, sovereign, I probably take it on more of the extreme definition, which is the ability to both create and run AI systems, or frankly any systems you want to be sovereign in. And so for that reason US probably is sovereign because it has energy, it has compute, it has data, it has talent, it has chip manufacturing, chip design. The UK by itself is probably not Sovereign. It doesn't really have energy, it has talent, it has data, it doesn't have leading edge chip design nor does it have manufacturing. And then you can go to countries such as Portugal or Spain or something like that, and then they're even less uh, having a chance of being sovereign. So what difference does it really make if those world leaders are buying a GPU data center from Nvidia and setting it up in their country and saying, oh, now we have cloud services and we can run AI, if like those same data centers could just get shut down by the builders based on saying like, oh, we're not going to ship any more software updates or now there's like a big change to like the software system and it makes it obsolete. I think the only way is to form alliances or pick one level of the stack that you're going to really, really compete on and be world best at, which forces other countries that are sovereign to want to partner, uh, with you or to need you. And so for example, like the Netherlands with asx, ASML seems like no one can really compete with asml. So that's like a good example of being like world class and therefore being part of the conversation. But if you're not world class in anything and then don't have the resources to be sovereign, then that's where the prediction around like just declaring neutrality, which is a little bit similar to the military defense guarantees of like, I'm a small island, there's no way I can have an army, so I need to form a defense like a protection agreement with a country that was willing to do that for me in exchange for something, something.

Speaker A: Maybe a good last area on themes of the report is safety. And I think this has become sort of a, uh, bigger and bigger theme in the past years and obviously the discourse has shifted over time. I think some of the particularly interesting areas when we can start with is the question of open weight safety, which has returned over and over again. And various people have different opinions on the kinds of open sourcing you should do and how this relates to safety. And I think there's been just a ton of debates on this over time that maybe we don't need to fully rehearse. But you did highlight three paths for open weight safety which relate to model based safeguards. So this could involve training, data curation, tamper resistant, fine tuning, watermarking and so on, then scaffolding and ecosystem based safeguards. So these are like content filters such as llama guard or monitoring tools, classifiers for algorithms, outputs and finally A step further from that being procedural and governance. So full access audits, stage deployment, things of this nature. What are your sort of thoughts on? I feel like there's one sort of pushback on a lot of these things, especially the model based ones is open weight safety measures can be undone with a few examples or a conversation of long enough. There's a lot of red teaming that goes on and so of course more and more robust mitigations are being built. But. But you can see lots of ways in which people can get models to do things that maybe are not ideal. How do you think about how that relates to the open source question?

Speaker B: I'm a bit in the camp of Pandora's boxes open and it feels a little bit too late to try to do containment of model weight releases or training data or recipes. There's too many people who actually understand how these things work nowadays. So it feels like most of the containment you can do is on large scale running of these systems. Yes, somebody can have access to it on their phone, but can they deploy it on a phone farm of tens of thousands of phones or can they deploy this on a large cluster if they need to do mass scale inference? This first part around deterrence and non proliferation just feels a bit like the evolutionary pressures argument. There's going to be some incentive of a leak and sometimes the leak is the most vulnerable in organization. I mean there's examples of these hacks happening of just somebody just calls a person, a contact center in a developing country and says like please reset my password and then after 15 minutes they reset the password. So at the moment maybe the labs have good protections, but at some point it's going to be somebody who's going to make a mistake. So probably more of the camp of close communication between developers and policymakers so uh, that people understand where progress is going, what issues might happen and how they might be be mitigated and then upskilling regulators to understand how to adapt their policies in this era. And maybe if I were running a major lab rate limiting the kinds of tools I would actually produce in the open source arena and then otherwise I don't have a good fantasy answer for

Speaker A: We've also seen China's AI safety approach developing over time. Matt Sheehan at the Carnegie Endowment has a lot of great stuff on this. You and Alex Chalmers had this 2023 essay called China has no Place at the UK AI Safety Summit. Your argument sort of being around China's AI regulation being politically motivated to preserve CCP control and not actually safety focused I guess the way that aids like China did attend didn't sign key safety commitments, but some Chinese companies did sign frontier AI safety commitments. And then Deepseek has this open source source approach that uh, I think is even more transparent than a lot of other places. And I guess looking at uh, China's AI safety approach in general, there's a lot of focus on these information content risks and to your point, political control. The goal for that content to reflect socialist core values. How are you thinking about in particular, maybe among those various things I said you were arguing for the exclusion of China from safety summits. We're seeing this transparency of open source models and things like this. How do you feel like that perspective that you and Alex had has aged?

Speaker B: On the one hand, it looks like the outcomes of that safety summit have only really resulted in the formation of safety institutes or now renamed security institutes. So they've increased state capacity for AI. At least the one in the UK has done strong work, I think by industry standards. But apart from that there's been no future safety summits. There's not really any adherence to like race dynamics, uh, slowing down to collaboration between labs and then China kind of does what it wants. Very, very recently, I think it produced some new policies that are coming out particularly around mandatory minor modes, uh, mental health, suicide intervention requirements, trained data governance and other safety guardrails that companies now need to adhere to. And that seems like pretty strong and far reaching. Would I invite them again in some sense? Like probably, but I wouldn't really expect much from it. It would be more performant.

Speaker A: You mentioned the sort of rebranding from safety to security that I think the UK has done in particular. What do you make of that?

Speaker B: It's a good question. I think it's good because, uh, security is easier to defend. It's easier to defend why one should care about security than should care about safety. And safety can to some extent be bundled within security. And it's like a better understanding, understood concept, but perhaps, yeah, it's easier to defend politically and uh, because cybersecurity is I think going to be a very important threat vector to invest in defenses against. I think cybersecurity is going to be a much bigger risk than existential risk of AI. And to me safety is uh, originally a lot around the X risk and less around the cyber risk.

Speaker A: You've had various predictions over the years. I feel like you all grade yourselves quite harshly on prior predictions. So I feel like this is a very, or for me as a person reading your stuff I feel like that gives me a lot of trust in how you think about things. But maybe we can go through a couple of these just really fast. So I think one theme that I notice on your predictions is your heads tend to cluster around these regulatory and political areas and you've had these US sanctions escalating and involving allies, but misses on semiconductor startup consolidation, the AI, video game breakout. Do you have any thoughts on just the patterns of your predictions and what you've gotten right, what you've gotten wrong?

Speaker B: I'm just trying to understand maybe if policy moves faster than research or vice versa. Because in some ways if we predict something is going to happen in research, like I uh, don't know, a video game that's fully generative AI based, uh, achieves breakout status. I still think that's going to happen, but is it just going to take longer than a year? But some of these policy things just happen a lot faster. I don't know if there's actually uh, a good trend here beyond potentially us hanging out more in one circle or another and therefore having a bit of better litmus test for how the vibe's going to change.

Speaker A: Maybe we can just quickly touch on one or two of your 2026 predictions before we wrap up. One of them is just a Chinese lab overtaking the US lab dominated frontier which we've already seen seems to be closing. Do you have intuitions or bets on where that could come from?

Speaker B: I'd say either ZAI or Deep Seq and I think in 26 labs are going to go and sort of science max. So I think it could be on some kind of scientific reasoning related tech ask.

Speaker A: And maybe the last one is you've discussed this idea of data center NIMBYism. Um, do you want to expand just on that very briefly, um, and how you think that might show up?

Speaker B: Yeah, I think the gist here is the uh, NIMBYism being not in my backyard. So on the topic that we raised before of these data centers are big, noisy, gas pollution, just not very hospitable places. I don't know if anybody is particularly happy to have that near their home. And this being being one of the biggest construction projects in the US is raising a lot of pushback and uh, with midterms approaching this year in certain of these swing states where data centers are quite popular in Pennsylvania could throw up some changes in political allegiances. I think as people really vote with their feet, that was mostly where it comes from.

Speaker A: I think with that lightning round, this is maybe a good place to end. So thanks again for doing this. This.

Speaker B: Yeah. Thanks. You done?

Speaker A: Thank you again to Nathan for doing this a fourth time. And thank you for listening. Hope, uh, to see you again.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • The End of One Model to Rule Them All: Why Enterprise AI Is Going Small, Specialized, and Multi-ModelDisambiguation · on Inference-time compute85 / 100
  • 183 - Verbal identity as the antidote to generic AI content with Ken MarshallProduct Led Growth Leaders · on 11 Labs80 / 100
  • 044 - Synthesia: Data Director - Why Data Teams Should Stop Trying to Be in Every RoomThe Stacked Data Podcast · on Synthesia79 / 100
  • Global tax compliance with an AI-native solution by Sphere with its CEO and CFO, Nicholas Rudder (USA)Voice of FinTech® · on 11 Labs76 / 100
  • AI, Authenticity, and the Future of Podcasting with Chris HillScreaming in the Cloud · on 11 Labs76 / 100
  • Building High-Performing Teams in the AI Era with Charles Guillemet & Sandra Schwarzer40 Minute Mentor · on Synthesia75 / 100

More from The Gradient: Perspectives on AI

All episodes →
  • Iason Gabriel: Value Alignment and the Ethics of Advanced AI Systems
  • 2024 in AI, with Nathan Benaich
  • Philip Goff: Panpsychism as a Theory of Consciousness
  • Some Changes at The Gradient
  • Jacob Andreas: Language, Grounding, and World Models
Explore the best B2B AI & Data podcasts →
All The Gradient: Perspectives on AI episodes →