The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Unsupervised Learning with Jacob Effron
Unsupervised Learning with Jacob Effron artwork

Ep 92: xAI Co-Founder Unpacks the Future of Model Development

Unsupervised Learning with Jacob Effron · 2026-07-31 · 1h 4m

0:00--:--

Key moments - from our scoring

Substance score

67 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality12 / 20
Guest Caliber17 / 20
Specificity & Evidence11 / 20
Conversational Craft13 / 20

Igor Babushkin's journey from DeepMind (AlphaCode, StarCraft) through OpenAI's reasoning work to xAI's Colossus has positioned him at the forefront of AI breakthroughs. In this conversation, he unpacks why the December 2024 moment - when coding agents became undeniably powerful - sparked his decision to leave xAI and start River AI. He articulates a clear concern: AI benefits are concentrating among a few closed-source providers controlling access through APIs and token-based pricing. River operates as a research company with three core bets: the River API (a reinforcement learning and fine-tuning service similar to offerings from Thinking Machines), personal AI agents trained end-to-end on individual users rather than average populations, and local hardware capable of running frontier models in homes and offices. Babushkin discusses moving AI verification beyond coding and math - domains with clear reward signals - into scientific discovery, material science, and everyday consumer applications where feedback loops must be engineered differently. He explores how personal agents will need access to sensitive data (documents, voice, context) and argues this data must stay local for privacy. The company explicitly rejects the assumption that inference compute must live in centralized data centers controlled by OpenAI or Anthropic.

Key takeaways

  • →Coding agents became undeniably powerful in December 2024, making it clear that agents will transform domains beyond software engineering - from scientific discovery to consumer productivity.
  • →The next frontier for agent improvement is solving verification problems in non-coding domains like material science and physics, which requires closing the loop with real-world experimental feedback.
  • →Bifurcation is already happening: frontier models serving a few specialized users (costing millions per inference), while everyday AI serves mass consumers who don't need maximum capability.
  • →Personal AI systems must be trained end-to-end on individual users with continuous feedback (happiness signals, wearables, behavioral observation) rather than averaged across populations.
  • →Running inference locally on consumer hardware solves control, privacy, and latency problems that centralized API-based models cannot address, enabling voice and video interactions impossible with data center latency.

Guests

Igor Babushkin

Topics in this episode

xAIcoding agentsPersonal AI agentsAI SafetyRiver AIDeepMind AlphaCodeColossus modelReinforcement learning and fine-tuningLocal hardware inferenceScientific discovery agents

Questions this episode answers

What inspired Igor Babushkin to leave xAI and start River AI?

After witnessing the power of coding agents in December 2024 and doing angel investment in AI safety companies, Babushkin became convinced that distributing AI control and benefits - rather than concentrating them among closed-source providers - is an urgent AI safety priority. He left to build technology enabling individual control over AI models.

How does River AI's vision of personal AI differ from current commercial models?

Current models are trained to behave the same way for all users (pooling feedback from everyone). River's approach trains models end-to-end to behave differently for each individual user, adapting vocabulary, communication style, timing of outreach, and recommendations based on continuous learning from that specific person.

What are the three main bets River AI is making?

The River API (reinforcement learning and fine-tuning), personal AI agents customized per user rather than averaged across populations, and local hardware capable of running frontier models in homes and offices to give users control and privacy over their AI.

Why is local hardware important beyond just user control?

Local hardware enables lower latency for voice and video interactions, improves privacy by keeping sensitive personal data out of centralized data centers, and creates better consumer experiences than data center-based models while keeping users fully in control of their AI.

What is the next frontier for improving AI agents beyond coding and math?

Scientific discovery and physical-world experimentation (material science, physics, rocket engines) where agents need real-world feedback loops, plus everyday consumer use cases requiring non-verifiable reward signals like user happiness that models themselves must learn to evaluate.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode contains a solid mix of forward-looking frameworks and concrete technical details, particularly around training methodologies, architectural choices, and business model transitions. However, it includes significant stretches of conventional wisdom (AI progress, the need for alignment, centralization concerns) that a sophisticated operator would already grasp. The most novel density emerges in discussion of non-verifiable domain training, hardware localization, and post-training economics, but these comprise roughly 30% of the conversation; the remainder covers well-trodden ground.

the biggest unlock would be just new ideas around how to set up the training such that it can handle much longer time horizons, um, such that it can handle m non verifiable rewards much more easily
what if we kind of break that assumption and we allow the model to behave differently for each individual that it's serving

Originality

12 / 20

While Igor articulates some genuinely contrarian positions - particularly the claim that proprietary model providers are in a structural bind due to capability saturation and regulatory pressure, and his thesis on distributed post-training as an alternative to centralized API models - much of the episode recycles standard framings: the Sorcerer's Apprentice metaphor, alignment concerns, and the coding-to-science progression. The personal AI customization angle is relatively fresh but underdeveloped in terms of novel mechanisms. The central thesis that companies should own their models feels somewhat inevitable rather than deeply counterintuitive.

you might actually have to keep the model private because they're starting to cross this critical threshold where now you really have to think carefully about whether you can give anyone access to the model
I think it's actually not the best place to be as a proprietary model builder

Guest Caliber

17 / 20

Igor Babushkin is an exceptionally credentialed operator with direct involvement in three tier-one AI efforts (DeepMind's AlphaGo/StarCraft/AlphaCode, OpenAI's reasoning work, xAI's Grok/Colossus). He's not a career podcast guest or pure theorist - he has hands-on experience building frontier systems at scale and has just launched a new company in the space. He can speak with authority to technical depth and organizational dynamics. His credibility is substantive and earned through execution.

He was at DeepMind where he led a lot of the work around Starcraft as well as AlphaCode. He uh, was at OpenAI. We're doing the early work on reasoning. Then he was a co founder of XAI where he did some of the heroic work on Colossus
I was a big inspiration behind XAI as well. So we were all really fascinated by this idea. Like well at the time LLMs, uh, weren't really capable of solving hard reasoning problems

Specificity & Evidence

11 / 20

The episode suffers from a concerning lack of concrete metrics, dollar figures, timelines beyond vague references (e.g., 'less than two years' for xAI, '120 days' for Colossus), and named examples. Igor discusses River's three bets, coding agent improvements, and training dynamics but rarely provides numbers - no latency figures, no accuracy deltas, no revenue or cost basis. The Colossus anecdote is evocative but light on technical specifics. Claims about model progress and market dynamics are stated confidently but backed by assertion rather than data.

within less than two years we're able to, to get to the frontier
we're able to fit all of the weights of the model onto a single chip, onto a single device

Conversational Craft

13 / 20

Jacob Efron demonstrates solid interviewing fundamentals - he asks follow-ups, probes Igor's reasoning, and occasionally challenges claims (e.g., on US vs. Chinese open models, on slowing down AI). However, the conversation often accepts Igor's framings without deep interrogation. Jacob misses opportunities to pin down specifics (What exactly makes Cursor's data superior? What's the actual throughput constraint on rollouts?), to probe contradictions (How does River's distributed post-training avoid the same data moat problem he attributes to incumbents?), or to push back on assertions (Is the 'bifurcation' thesis actually evident yet, or speculative?). The conversation is intellectually generous rather than adversarial.

Just awesome to talk to someone who's at the forefront of the space
Yeah, but you think people like the recipe is kind of known and it's just literally about running that experiment?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B74%
  • Speaker A26%

Most-used words

models91model43data38agents34different32world30feel30coding25training25better23today22open20figure19best18openai17research17

Episode notes

Igor Babuschkin, co-founder of River AI and formerly a co-founder of xAI, joins to unpack a career that spans nearly every major AI lab: he led the StarCraft and AlphaCode work at DeepMind, joined OpenAI's reasoning team years before o1 shipped, and co-founded xAI, where he helped stand up the Colossus data center in roughly 120 days and reflects candidly on what it's actually like working with Elon Musk day to day, plus what the Cursor acquisition actually unlocked for Grok's coding models. He also discusses why he left xAI to start River AI, the three bets behind it, and why he's betting on local hardware, not just software, for personal AI. On the enterprise side, he tackles whether companies will actually train their own models or if it's just a cost play, and makes the case that proprietary labs like OpenAI and Anthropic are facing a real business squeeze. He's skeptical that stacking specialized RL domains generalizes the way pre-training scale did, and is candid about the uncomfortable reality that today's frontier open-weight models are almost entirely Chinese.

Full transcript

1h 4m

Transcribed and scored by The B2B Podcast Index.

Speaker A: Welcome back to Unsupervised Learning. I'm Jacob Efron. We had an awesome episode today with Igor Babushkin. Igor has been at all the right places at all the right times, uh, leading a lot of really interesting AI work. He was at DeepMind where he led a lot of the work around Starcraft as well as AlphaCode. He uh, was at OpenAI. We're doing the early work on reasoning. Then he was a co founder of XAI where he did some of the heroic work on Colossus as well as helped get the models up to, uh, you know, where they are today. Um, it was awesome to get to talk to Igor about his new company, river AI. The bets they're making around personal AI for companies, for consumers, uh, as well as local hardware. We talked about all his takes on where the ecosystem's headed, what's required to improve models beyond coding into non verifiable domains, and when he thinks we'll get there. We talked about why he thinks the closed source model providers are actually in a really hard place right now as businesses. And we hit on his reflections on what AI models mean for the future of the world and the policy implications of this as well. Just awesome to talk to someone who's at the forefront of the space on, uh, all the questions that are top of mind. I think folks will really enjoy Igor's perspective. Without further ado, here he is. Igor, thanks so much for coming on the podcast. Really excited to have you here.

Speaker B: Yeah, thank you for the invitation. Jacob, we're excited to chat with you.

Speaker A: You've pioneered so much important research in your time at DeepMind. OpenAI. Xai, you, uh, were early to like so many of the things that I feel like def, uh, this era we have today, verifier based reasoning, inference, time search and you know, you recently left XAI to start your own company. River really focused on personal AI. I feel like there's so many things that our listeners are going to be excited about. To get your take on the episode. I want to make sure they understand river. They get your reflections on this. You know, being at the epicenter of all these things that have happened in AI, as well as your predictions for kind of where we're going. You know, as I was thinking about where to begin, one thing I found really compelling is that you wrote some really interesting fiction earlier this year. Um, for those that haven't seen one piece was this dystopian look at how models ended up uncontrolled running the world and it had this really interesting ending where your protagonist kind of realizes that they dreamed all of this and none of this has happened yet, but they still make the decision to take that first step and begin. And then kind of interesting contrast. Your second piece was really this love letter to the machines that come next. Uh, and so maybe to start, I'm curious what inspired you to write these?

Speaker B: Yeah, I think that's something pretty big that happened to all of us software engineers, AI researchers around November, December of last year, which is that the coding agents suddenly became uh, so powerful that you can't ignore them. Previously, some people use them, some people, uh, didn't uh, use them. Um, but from that point onwards, that latest iteration of um, cloud opus models, uh, suddenly everybody agreed that yes, these models are very powerful. Yes, they, they make our jobs much, much easier. And also while they're doing so much of the software engineering, um, now that we used to do that we used to, used to enjoy. So it was a real transformation for all of us. And like many folks, um, I spent large, uh, part of December, uh, rewriting all of the coding projects that I've ever um, uh, wanted to do in my life or have done, um, using the coding agents. And um, you get this incredible sense of like, wow, anything's possible now. And also these agents are going to continue to improve. They're going to um, to get more capable. And what will happen with the world at large as these models are getting stronger and stronger. And that kind of inspired um, the story. It's very much, um, um, a modern version of the Sorcerer's Apprentice. So the Sorcerer's Apprentice is very, uh, fascinated with magic and starts to use it, but then it kind of runs away from him. So kind of like in the Disney short, um, or in the original uh, form. So it's very much a, uh, version of that in my mind. And I think we're all becoming the Sorcerer's apprentices Now with uh, LLMs and agents becoming uh, more and more capable. And I think it's not going to stop at coding. So I'm very, very interested in figuring out um, where our agents are going to go next, um, what else is going to be transformed. And um, the more I think about it, the more it feels like everything will be transformed. All of computing really is a potential is up for grabs when it comes to agents because, um, the ability to deeply understand, um, what's happening in your life, at your workplace, what's happening in the world at large, and then taking action based on that, um, is something that you can apply to any problem, to any kind of um, thing you might want to do as a business, as a consumer. So that's why I'm so interested in personal AI, because I feel like that's the next iteration, that's the next manifestation of what the agents are going to be able to do for us. The impact will only grow from here.

Speaker A: I mean obviously it feels like that December moment you talk about resonated so strongly with so many folks. And I think actually even a lot of researchers responded to your story being like this really hits hard. We had that moment in November, December on the coding side. And I think a lot of folks are trying to figure out what's required to get agents to work really well in other domains and kind of maybe asking, well, is there something special about coding and math or the most verifiable domains out there? Are we to see the same kind of progress in other domains? How do you think about what needs to be done to kind of like get to, you know, even more general purpose agents?

Speaker B: Yeah, I think there are actually many next steps that we could be taking. Uh, from here there are many different directions that open up. Coding was sort of the central path, uh, the obvious low hanging fruit, um, because it's verifiable, because um, it touches so many different domains, it unlocks so many different things that you can do with the agents. But from here there are a few different options of where to go. I would say scientific discovery is a very clear one. We're also starting to see the first uh, sparks of uh, agents being able to solve um, real serious problems in math and the sciences. And I was a big inspiration behind XAI as well. So we were all really fascinated by this idea. Like well at the time LLMs, uh, weren't really capable of solving hard reasoning problems. But we felt that this was the next big breakthrough and that was a big inspiration behind um, xai, where we started it. So I think agents will continue to improve very rapidly on complex coding projects and also math problems, anything that's verifiable. So with math you can even formalize um, your theorem and your proof using tools like Lean. And it gives you a good reward signal that helps you improve agents. It also helps you tell if you've actually found the proof as these get more and more complex. And seeing that actually work for real, uh, is uh, really fascinating because uh, AI researchers, we've thought about this for some time before, like maybe formalization will be useful in the future and so on, but actually seeing it happen in real Life um, is really quite something else. But beyond this, um, one huge bottleneck is data. So if you want to do great breakthroughs, um, in the physical world, whether it's material science, whether it's new fundamental physics discoveries, whether it's building a better rocket engine, the type of thing that Elon might want to do with uh, the agents, um, uh, that's where we often need um, to run experiments in the real world. So you need to figure out a way to close that loop with the real world environment. So the agents, when they have an idea what kind of material to design, they need to get some feedback as to whether that was a good idea. Did it work, did it not work? So that's uh, a big frontier, um, right now I would say, and using agents for scientific discovery. So this kind of direction of science and discovery, I think that's a big uh, thing that is happening at the moment. That's something a lot of people are focusing on with the agents. And there's a potential there that we'll get these highly super intelligent AIs that are vastly more capable than any human or even any collection of humans at creating scientific breakthroughs, at making decisions, maybe at predicting the future as well. So I would also broadly categorize that in that type, ah, of type of direction. And then I think there might be other directions as well. So what is the AI going to be that the everyday person is going to use? They probably are not going to use the super AI that costs a million dollars to run once on a question. But what is that AI that everybody, um, so I think we might see a bit of a bifurcation, uh, there actually where um, there are the super AIs that only a few people have access to, but there's also the everyday common AI that you and I will be using all the time to be more productive or to help organize our lives better, just lead better lives basically. So very much. Um, so the reward function is how much does it help the human feel better and get more things done, feel better about themselves? Um, so I think there's also tremendous potential to improve things for humanity, improve our everyday lives. But um, those agents might not actually need maximum capability. They might not need to be able to prove the Riemann Hypothesis.

Speaker A: I mean it already feels like that bifurcation is happening right in terms of uh, the most frontier models being really valuable for a lot of coding and mathematics tasks. But you know, for, for many people, uh, that, that are, you know, ChatGPT users, uh, you know, it's not Inherently clear that the more complex models have actually, you know, the ones from a year ago were, were totally fine for, for most of what they wanted to do.

Speaker B: Yeah, yeah, exactly. And um, I think it's important for the individual to also feel the benefits of AI because increasingly it feels like there are a few large companies that create these powerful AI models that control access to them, like get to say like how they will, how they will behave. And for me it's just very important that everybody feels like they're actually personally able to get value out of AI. And that's a big inspiration behind where I started, uh, river, uh, recently because um, to me it feels like potentially a technology problem. So we've made certain assumptions about how AI should be built and distributed to people. You should do large pre training runs, you should set up an API and charge people by the token. Um, and that's sort of the model everybody um, knows about, everybody follows. But um, it doesn't have to be that way. So we can come up with new ways of building AI models, of customizing them, of distributing them to people that um, maybe will put them more in control of their experience, will put them more in control of what the AI um is doing. And um, it's a research problem, it's an engineering problem. So it's something that uh, I find very fascinating. So that's why we decided to start this new company.

Speaker A: When did it kind of crystallize for you? Because obviously your previous experience had been at some of the largest closed source players with those kind of models. And so uh, when did this kind of all come to be?

Speaker B: Yeah, so when I left um, uh, Xai, uh, in 2025 decided uh, to do a bit of angel investment in AI safety companies. So I got to meet a lot of different founders in the AI community and AI safety in particular. So people um, having new kinds of ideas around how we could um, help make AI more beneficial for humanity in the long run or to better control the downsides of AI or the risks. And um, there are a lot of amazing ideas um, out there and um, I uh, found it all really fascinating. But at the end of the day I'm just sitting around waiting for all these folks to make progress on their mission and to succeed and got really impatient. So I just um, felt that I wanted to start something new myself so I can help push on this direction. And for me, um, helping distribute the benefits of AI, helping distributed control over AI is um, a part of AI safety. It's maybe the, for me the most urgent thing that we need to do right now because, um, power over AI is concentrating quite a bit. And uh, especially in the US you can feel that people aren't too happy about that. They, uh, don't necessarily trust the AI builders or, um, feel like they have their best interests in mind. So I think it's important we show some ways in which people can, can feel the benefits, can feel like they're in control of the whole thing.

Speaker A: Just wanted to take a quick break to say that if you're enjoying the conversation, the single best way to support the show is hitting follow or subscribe wherever you're listening. It makes a real difference in helping us get the best guests and helping others find the show. Now back to Igor. Well, maybe an interesting, uh, time to just talk about your vision for like the River AI as a company, the product and you know what, how you hope to kind of achieve that.

Speaker B: Yeah. So when, uh, we started the company, we kind of looked around like, where are these technological opportunities, like what are these ideas that have been underinvested into that could change the way that we use AI, that we build, um, AI and let's uh, run this as a research, each of these as a research project. Let's figure out if um, this is something that we can make work and then ship to everybody, um, in the world. So we have really three different projects that we're working on at the moment. So one is we realized, um, the tools to help companies take control of AI are pretty much already there. You just have to make it work really, really well. So that's why we started the river API, um, which is, uh, you can think of it as a reinforcement, learning and fine tuning service, um, similar to other, um, players like Thinker from Thinking Machines, um, but we have our own unique take on it. And the one thing that we're really proud of is our engineering. So we really optimize, uh, this kind of platform as much as possible. Um, try to make, uh, training as cheap and as reliable, as scalable, um, as we can make it using these skills that we've learned over the years. Um, so that's the first bet that we're making on the Rover API. The second one is around, uh, personal AI and personalization. So we feel like there's an opportunity to align, uh, AI agents much, much better with the individual than what we're currently doing. So today we're training AI models, uh, to basically work well for the average user. So you're literally pooling all of the feedback, um, all of the results from all of your different users or during training, and you're telling the model behave the same way for everybody. And our idea is, uh, what if we kind of break that assumption and we allow the model to behave differently for each individual that it's serving? So your AI agent will speak differently to you than my AI agent speaks to me. And they're going to have different preferences, different behaviors. This area of research is still pretty early, so there are a few different directions we're going down here. Um, because obviously all the agents that are out there that are commercially successful, they're all trained in this sort of average population kind of way. But, um, we feel like it's going to feel really amazing once the agent truly learns directly from you and gets better every time you interact with it. So, um, something that I'm pretty excited by. And then the third big bet that we're making is on hardware, because right now the assumption is that the only way to get access to the most powerful agents is through access to a data center. So companies like OpenAI and Tropic, um, they bring up really, really large data centers. All of the inference, all of the training happens there. And as an individual that wants to use AI, you have to go through their APIs. There's no other way they can decide, um, how it's restricted, where they get access or not. So really the ultimate move here is can we, uh, bring the inference compute, uh, locally to the users? So what would it take to actually be able to run a frontier model in a small device that you can have in your office, that you can have at home? Um, very much a research project, because today you're very memory limited there. So we're actually able to fit all of the weights of the model onto a single chip, onto a single device. So that's something that we want to figure out, like, how do we bring the best models in the world, um, to the individual so you can really feel like you've got control over it, it's in your home, your data is protected, um, as well. Um, so those are the three bets that we're making. And we're very much running this the way that back in the day OpenAI and DeepMind will run. So we're thinking very long term. Um, m. We're taking some riskier bets. Um, we assume there's going to be research, there's going to be some new ideas that will be needed.

Speaker A: No, it's fascinating. I think each of those three bets is really interesting. And so maybe to kind of dig into Them. We'll start with the last one maybe on the hardware side, obviously you have a really broad scope of things, uh, you're seeking to accomplish here is the motivation there. On the hardware side, obviously, ultimately the most controllable AI is one you can run locally. Right. That no one can kind of dictate whether you have access or don't have access. Is that kind of the motivation, or is there other kind of tangible benefits you see to end users that made you increase the complexity level even further by adding the hardware component to it?

Speaker B: Exactly. So, um, really the control aspect is the most important to us. So we'd like people to be able to take control of their AI experience and really feel like the model is theirs and not anybody, Anybody else's. That's a big motivator. But then as we started to work on this, we also realized there are a few additional benefits. So having, um, this box in your home locally with you could uh, give you a much lower latency, for example. So it opens up a lot more possibilities when it comes to interacting with the agents via voice or via video input, video output. Um, so there are actually some new possibilities for consumer experiences that open up once you go with this form factor. So that's something I'm very excited about because it might even make the whole experience much better than what we're getting from the data center world. Yeah, And I think just uh, the privacy is also going to be uh, important in the personal AI world. So this personal agents that know you deeply, they kind of observe everything that you're doing and then help you proactively. They're uh, going to get access to incredible amounts of sensitive data about you. So basically all of your documents on your computer, um, you know, um, everything that you're saying in the room, uh, could be stored by a personal agent because it just makes it more useful if it has that context for you. And sending all of that data out to a data center where there might be, um, some security leaks or that your data might be trained on. Even if you don't like that, um, that's really suboptimal. So we're also thinking of it from a, uh, privacy point of view.

Speaker A: I think part of the vision of the company is everyone having their own model that's continually learning to improve and be customized to that person. Folks think about how these types of models may be built. There's like different schools of thought. One of them is, you know, going into the weights themselves and updating the weights and other people saying, hey, you Might just be able to do this through like a really good memory system and prompting. You know, how do you think about that? And uh, is that still like, you know, uh, are there a bunch of different research directions you all are pursuing or do you feel pretty convicted in the uh, in the approach?

Speaker B: Yeah. So for us it's pretty open which of those um, are going to be the key, the key tools to use. But I'd say from my experience building AI models, the one thing that you really, really need is end to end training. So you actually want to train your agents to make use of their, the personal information to make use of their memory systems if they have one. Um, that's really crucial. And uh, today often what we do is we um, take off the shelf, um, proprietary uh, agents which have been trained a lot on coding but not necessarily trained on over long time horizons helping an individual using their memory to help them. So how do you maximize somebody's happiness as the reward function? That hasn't really been done yet. So for me that is the key aspect because I've kind of observed um, over time the models that were trained for a particular task um, have performed the best and you really want to train end to end on the actual signal that you want to optimize which would be individual's happiness.

Speaker A: Are you going to basically uh, have the models running in the wild and hook it up to some self reported happiness or a wearable I guess to close the loop?

Speaker B: Yeah, that's really something we got to figure out. But you could also just try to observe did I end up helping uh, the user and the models are getting more and more intelligent so they might be able to figure out oh actually I made things worse today, negative reward, I made things better, positive reward. So the possibilities are growing there. It's moving into this non verifiable domain that's still very challenging. So we need to use um, models to figure out how you're doing. But um, increasingly with models becoming more intelligent is getting possible.

Speaker A: As you kind of envision like the future state of river, like what are the kinds of differences in models that you think of as canonical examples of? If you and I had different models over time, obviously there's some set of capabilities that if the general models had, everybody would be happy. And then there's different things folks might want. Uh, what are the examples that you think of? Top of mind.

Speaker B: I think the examples are really endless and it's just hard to imagine it because we haven't experienced it yet ourselves. But it could be really subtle things from um, maybe how much text, um, the model writes, when it replies to you, to what kind of vocabulary it uses, is it more casual, is it more formal? Um, so all of these different aspects of the answer, um, today we have one, uh, sort of fixed set of, fixed set of responses, how the model is going to respond, one fixed style, um, in which it will respond to you. Um, but in the future this can be untied and it can pretty much be endless because it will affect um, how it will help you throughout the day, like when should it reach out to you, in which situations is it being too annoying or is it not helping enough? You know, it will calibrate itself, um, to you. It will know um, what your taste is in terms of books, movies, music, and it will be able to adapt your day uh, to that. It's really um. The possibilities are actually endless. Um, to um, be a bit more inspired about what's possible, you can look at what recommendation systems are able to, to do today. So if you use an app like TikTok or YouTube, it figures out what your interests are, um, and allows you to go really, really deep on what you would like to see. Um, and the experience is totally different from giving the same set of videos to every single user. That's kind of what we want to create for the agent, um, world. And you can kind of think of it as a recommendation system, agent hybrid, which has never been done before. So that's something we're a bit excited about.

Speaker A: You could imagine obviously the parallels in the enterprise space. And you said you're kind of starting the business with uh, working with enterprises on models. There's obviously been a ton of dialogue in the last months about owning your own models and the importance uh, of training your own models. It feels like today a lot of people are doing that really for cost and speed, like finding ways to run, you know, smaller models. How do you think that changes over time? Like do you think folks will, will actually try to push capabilities forward or like, you know, even. Obviously a lot of, there's a lot of parallels to what you talked about and people having different preferences, companies have different cultures, ways of doing things. Like, how do you imagine the parallels there?

Speaker B: Yeah, I think today, um, a lot of people go to the proprietary models first. They have the strongest capabilities, they're used by everybody else. That's a big reason to, to start out with them. And uh, then they might do custom models for um, speed reasons, cost reasons, maybe privacy, uh, could also be an aspect if you actually have to run this on your own infrastructure because you have maybe other customers that have more sensitive data. So those are the main reasons that I'm seeing. Um, but I think a lot of the reason why people go to the proprietary models is that um, the field is moving so quickly. If you wanted to learn how to train your own models internally at your company and sort of master uh, deep learning, master reinforcement learning and uh, get the most out of your internal data that you have, um, you have to build up a team, you have to build up infrastructure and um, it can be a many month project to be able to surpass the proprietary models by that time.

Speaker A: Right in time for the next.

Speaker B: Exactly. There's a new model that's even better. So you could have just gone with the proprietary model and um, I think that's been the way things have been going for the last ah few years but expect that to change um, now that um, open models are starting to become extremely capable and we're um, sort of um, starting sort of harnessed all the low hanging fruits in terms of what you can do without deeply deeply knowing um, basically the data and the problems that the company is solving. So if I'm an expert in a certain area, my business is building a certain SaaS tool like an HR platform, uh, for example, can imagine that I have the best um, data, I have the best understanding of what the customers need and I thought about the domain quite a lot, much, much more than Ananthropic or OpenAI would think about the product that I've built. Um, so really those companies are in the best place to decide how their agents should be working, what they should and shouldn't do, how to best help the customers. And um, so far that hasn't really um, hit yet because we're still riding the wave of stronger and stronger proprietary models. But as soon as that runs out and um, then you have a choice. Either you give all of your data to entropic to OpenAI for them to post train their models better or you start to do it yourself um, to try to squeeze out additional um capability and I think a lot of people will opt to go for the second one because you're kind of giving away the whole foundation of your business. Otherwise we're giving away the Ray thing that gives you an advantage over all your competitors.

Speaker A: You think the proprietary models do slow down because obviously there's endless amounts of data budgets there seem to go purchase data uh, across a bunch of different domains. I guess I could see a few different arguments but one might Be take uh, a bank for example. You can actually go hire a bunch of people that were ex bankers and they can label a lot of data for you and do a lot of, do a lot of things in different environments and suddenly your model gets pretty good at finance. And I think it's an interesting question of one, do we expect the model progress to continue in these other domains? And two, how unique is actually an enterprise's data versus the army of data labelers that exist out there?

Speaker B: Yeah, so training uh, of these models actually has diminishing returns. So um, the stronger we want to make the pre trained model and the model Overall, the more GPUs we have to put together, the more high quality data we have together. And um, AI companies have done an incredible job at scaling up. We're putting in so much more effort at gathering the data, tuning the models, improving them again and again, adding new innovative algorithms and capabilities, um, into them. But eventually something will have to give. So at some point you can't cover the entire earth with, with GPUs. Um, so eventually there will be a bit of a slowdown in terms of what uh, the AI companies are able to do in terms of the capabilities of um, their models. And I would argue we're starting to get close to that point. At the same time there's another thing that's happening which is um, these models are starting to get so capable at the frontier that um, you might not want to release them anymore or you might even get criticized or there might be some regulatory action that prevents you from releasing the next generation of the model. Um, at the same time open models are getting stronger and stronger. So I think as a proprietary model builder you're kind of starting to get squeezed in a little bit. You make the model too good, you're not allowed to release it, but then open source is just right behind you, getting better and better, um, every month. Um, so I think it's actually not the best place to be as a proprietary model builder. So I think the way out for OpenAI and anthropic is to innovate. So the same thing that they've done the past few years that's allowed them to be in the place where they are today. So they have to come up with some really new fundamental ideas about how to make the models more useful, um, how to um, make everybody feel the benefits of what they're building, um, explore new domains, explore new ways of having impact, generating revenue.

Speaker A: Because you think that basically the capability improvements are slowing down.

Speaker B: Yeah, exactly. If they don't then um, you might actually have to keep the model private because they're starting to cross this critical threshold where now you really have to think carefully about whether you can give anyone access to the model. And um, then on the post training side um, the efforts that's being made is really incredible. So you're basically assembling um, all the world's experts on every single topic and then having them work with you to centralize all that knowledge flows into one set of weights and then gets delivered to, to all the customers. So um, scaling that operation up is also incredibly challenging. There are also some diminishing returns there and I think that the way out there is to go distributed. So there are all these experts out there in the world already for their particular domains. So they're inside of the existing companies, they're uh, at the universities, um, and so on and so forth. And I think we should give people an opportunity to basically independently create improvements to models and, and essentially post train them locally for themselves, for their company, um, for their team, for the individual um, rather than trying to capture all of human knowledge essentially and then sell it by the token. So just um, my personal philosophy and uh, that's the type of thing that we want to do at river AI figure out how do we allow people uh, to do their own post training essentially locally.

Speaker A: It's interesting because I feel like we're in this moment where you buy a lot of data in a domain and you focus on that domain and that then adds those capabilities and you're kind of going one by one and it feels very different than the uh, pre training world where it's like you get to some certain threshold and then you start to see massive generalization across the board. And I think some people have posited that the same thing would happen in post training in RL eventually. I don't know, you're adding domain 30 or 40 or 50 and suddenly it becomes like you get a lot of benefits from having everything in the same place versus you know, just training on, on you know, finance in this place or health care in this place. Like do you think we see that generalization and that benefit of generalization down the line for RL or is it like going to always make sense to just take some really strong pre trained model and like you know, uh, apply one domain specifically to that?

Speaker B: I think we're seeing generalization in the sense that these models have such incredible knowledge of the world and they have such a wide set of skills that they can now go into Any domain, um, and bring in all the knowledge from everywhere else. I think coding is a great example of that. And then I think whether there is a lot of synergy between all the different things that are being post trained, I don't really know to be honest. I feel like there are so many specialized domains, so many specialized skills out there that you only really need in one place. I uh, wonder how much of that uh, really is contributing to the overall intelligence of the model. A lot of the intelligence seems to come from coding and math training and things like that at the moment.

Speaker A: Yeah, that seems like a really interesting question for how much it makes sense for enterprises to post train their own models. Right. It's like there's pharmaceutical companies that have data sets of drug trials that pretty proprietary. There's chip design companies. Uh, there are certain things that you're like yes, those companies have quite unique data um, that probably isn't out there. Maybe it's less clear to me if your average bank actually has data that isn't just going to be that eventually your closed source model providers get access to.

Speaker B: Yeah, we might see some of those businesses struggle if they're not able to retain their competitive advantage, if they're not able to leverage their own IP to make their own models uh, better. So I think it really depends on the industry and that's kind of how you can guess uh, where things might go and how the transformation might happen. I would say the worst place you uh, can be in is if all of your knowledge is on the Internet. If everything that makes uh, your company unique and special is already out there and can be, can be scraped. That's a very dangerous place to be.

Speaker A: It kind of exactly ties into the overall thesis of river which is like you know, keeping this intelligence local, you know, the kind of ways you're interacting with and using these models.

Speaker B: I'd say we have a choice actually. Um, right. I think uh, as a company you can choose to give your data to somebody like Entropic and OpenAI in a way by, for the APIs, for other means, for uh, maybe selling your data to them even if you want to turn a profit on uh, it. Or you can start uh, to leverage it to build your own models. So I see these two different paths. One obviously leads to more centralization of the intelligence and one leads to more distributed benefits of AI behind all the

Speaker A: stuff you're building today it seems like anyone that's trying to build their own models or fine tune models is using the, the cutting edge Chinese open source models. And obviously there's been a big discourse these past months about uh, these models. Um, obviously I imagine you think uh, cutting edge open source models are really good for the world. What do you make of the fact that this ecosystem is all Chinese today and then the government talking about potentially banning these models?

Speaker B: Yeah, I think first of all it is good that we have powerful open models um, out there. So uh, it really lowers the bar for anyone who wants to get into AI, who wants to customized models, who wants to use models um, to their own advantage and really have that control of their own destiny. Um, at the same time it does put the Chinese labs in ah, a kind of a preferred position where they also can exhibit some control. So for example they might um, stop releasing their open weights in the future and everyone relies on them. Um, that's not a great thing for the U.S. uh for the U.S. economy. Um, they could also um, change the licenses to those, to those models. So we often see that already that there are certain terms in there. If you're a company that has more than X revenue then um, you need to enter into a partnership agreement where the rules might be a bit different. So that's another way to, people can exert a bit of pressure and um, stay in control of the open weights. Um, and then um, one issue uh, that's more uh, theoretical but um, it is theoretically possible is um, for the waste to somehow have backdoors or behaviors that we would find undesirable.

Speaker A: How worried about that are you?

Speaker B: Yeah, today I'm not very worried at all because my sense is that if um, these kinds of things existed we would um, already have um, at least one example of it happening in the wild. We haven't seen any reports of it. Also the technology that I'm aware of for kind of planting these behaviors, planting these backdoors into weights is not advanced enough today to really uh, avoid detection and not leave any hints. So I think we're kind of early on this um, right now and I'm not too concerned. But in theory the concern um, is there. And over the next few years as this tech becomes more sophisticated, um, and you can actually, and maybe LLMs and agents become more important as well in all kinds of domains including the military. Um, um, it is something that we should have on our radar. So I would say it's really critical that we have the best um, open models in the US So that's something that I'm very supportive of and I think we should all figure out um, how can we get a US team to train the very best open model, period. Not just the best US Open model, but the best, best one, uh, worldwide. Um, so yeah, definitely, um, super supportive of that.

Speaker A: Yeah, it seems to be an interesting business model. Question of how you end up ah, throwing it. But it's like such a public good that maybe find a bunch of folks, uh, I mean Nvidia already is kind of subsidizing this stuff quite dramatically to make it possible. And so maybe you'll see, maybe it's

Speaker B: going to be us um, as well. Right. So um, we've had that experience with Xai of uh, starting from scratch and then within uh, less than two years we had Grok Free and then Grok four as well, which were very much at the frontier. So um, I'd love to um, train the best um, US Open model. But um, right now we're very interested in um, how to build a business around these open weights so that the company can be self sustaining and we're making good progress on that. But the RL API with the other products that we're building.

Speaker A: Well I'd be remiss if having you here not to ask a little bit about some of the past roles you've had because I think you've obviously seen uh, you've had a front row seat, uh, and kind of led a lot of this stuff that's happened over the past years. Maybe to start you've said AI progress took off faster than you expected. Is there a moment or two that really surprised you these past uh, I don't know, three, four years I think

Speaker B: the crazy thing is I'm not too surprised by the progress. It's just uh, crazy living through it. Uh, it's totally different from when you imagine it. It's just so different from actually living it, uh, for real. It's almost like being in a dream. So uh, when I um, started to get interested in AI, I was a physicist. I was working at the Large Hadron, uh, collider, was uh, doing a PhD, uh, there. And in physics they teach you basically to really think ahead. I think some of that rubs off on the physicists as well and they start to think about how will the future look like. And uh, so when I realized that for example AlphaGo, uh happened so DeepMind was able to uh, create a go, uh AI that's very, very powerful. Then for me that's when it clicked that amazing things are going to happen in AI and I better switch over you.

Speaker A: And a lot of prominent researchers uh, made the move from physics. It's very interesting how many folks started in that world.

Speaker B: Yeah, I think it's just a great uh, way to learn the ropes about how to think about these tough problems and um, how to solve problems from first principles. So it's a great preparation for doing AI work. Um, and I decided to join uh, DeepMind as a research engineer which is really um, fun because I got to see both the algorithms side and how to solve these research problems, how to train the models and also how to build distributed systems, um, how to improve efficiency of the training of inference and so on. Then um, worked on text to speech generation with Wavenette, um, and then got really interested in reinforcement learning and ended up being a tech lead for the Starcraft uh, project at DeepMind. So it was a really challenging um, RL uh project. Kind of in some ways it was ahead of its time 100%.

Speaker A: I mean that seems to be a much more complex environment than uh, a lot of things we're optimizing RL on today.

Speaker B: Exactly. We had to figure out how to deal with that complexity and um, some of the strategies we've developed um, we could immediately um, see how they could apply to language models as well down the line. So for example with starcraft we did a lot of imitation learning. So we got all these um, replays of people playing StarCraft at various skill levels from beginners to uh, almost professional players. And then um, we started with imitation learning and only then switched to rl. And um, the rewards here are verifiable because it's announced that two player game, whoever wins gets a positive reward. Um, if you win you get a negative reward. Um, and um, the combination of those two imitation and then uh, large scale training for imitation and then reinforcement learning to further improve the policy, that's really, really powerful we found. And that's still the way that these agents are trained today. So really after the Starcraft project um, at DeepMind we actually realized that um, coding is going to be the next big thing. So it was a bit early uh, at the time. So it was way before uh, there were any commercial sort of capable commercial code models um, out there. And we kind of saw it as a natural extension of Starcraft. And so we started this project called AlphaCode UM at the time and um, I had the strong feeling that models um, that basically LLMs were not performing well on code at all at the time. And I thought the reason was that they're not able to think. We don't give them enough time to think about the problem basically. So um when humans solve tough problems they have to think for a moment or two and um, something magical happens that allows them to understand the problem much better and uh, uh, to solve it. And we didn't have a mechanism back then for LLMs to do that kind of reasoning. And so I really felt like we've got to do some research um, on this and um, who's working on that. And I realized OpenAI, uh, has just started uh, up a reasoning team uh, exactly to figure out this problem. So uh, joined OpenAI, uh moved to California and um, yeah was both involved with the reasoning effort there but uh, also helped out on some big pre training runs together with uh, Greg Bachman. So kind of got to see the large scale pre training side there as well which was really fun.

Speaker A: A lot of this reasoning stuff seems to have worked post really large scale pre training. And so it seems like in the early days of the reasoning work at OpenAI there was almost just this conviction like this is going to work as models scale and get better. Which is really impressive intuition.

Speaker B: Yeah, exactly. And then at the time we didn't have the solution just yet. The solution ended up being essentially the O1 uh model. And um, the research that happened later at OpenAI to unlocks the basically this reasoning approach where the model uh, has thinking tokens that it's allowed to generate before giving the final answer and then use reinforcement learning to basically improve its thinking process and thinking ability. Uh, so that ended up being really the key that we were uh, looking for. That happened a few years uh later. And um, then um, at OpenAI I just felt.

Speaker A: Like you know, DeepMind had a who's who of heavy hitters. I mean I think almost everybody did a uh, did a, did a stint through their or Google brain at some point. You know, I think as people have reflected on it it's like maybe you know, there was more, you know, different kinds of bets made or more of a belief in kind of reinforcement learning and, and you know, less of a, of a focus. I don't know. You lived through both of them. So I'm curious like how, how you, you know, uh, uh, how you thought about that.

Speaker B: It's really hard to uh, say honestly because there were all these extraordinary thinkers at uh, DeepMind and I think some of the right ideas were kind of floating around for sure. But uh, I think what helped OpenAI was that they were a smaller team and they were able to get great alignment within themselves, great conviction that this is the right approach. I think it requires kind of that Critical mass of uh, talent density and people thinking along the right lines, uh, which they were able to do. So even as the underdog, they were able to make huge progress. And um, I think to us at the time it was pretty clear that OpenAI was going to win in many ways. And then already at that point I started to think like um, wow, that's kind of uh, unfortunate if there's only one AI company in the world that uh, controls the most capable models that uh, decides what gets done. And we would even think about how do we uh, set up UBI for everybody so that uh, humanity can be paid while OpenAI, um, uh, essentially controls the models. Which just didn't sit right with me because I felt like I was always a big open source fan and big Linux users over the years. So I felt like that's the way uh, to distribute the benefits uh in the end philosophically if you think about it, the pre training of the model happens on all of humanity's knowledge, all the text that we've created out there on the web. It's uh, really a commons, it's really something that humanity as a whole should own. Um, uh, the fact that nobody technically owns it and you're able to just download it basically allows you to compress the knowledge, to refine it, to uh, turn it into something that's actionable, that's useful uh, in the form of the model weights. Um, and that's really what the AI model builders are doing in my opinion. And to me philosophically it makes the most sense for those pre trained checkpoints to go out and be free for people to use just um, thinking from first principles what's fair here. But um, obviously the AI labs are putting in tremendous effort in the form of compute, in the form of all these experts that figure out how to um, set up the algorithms, how to train the model, how to innovate um, in the space. Um, so there's definitely not huge benefit that comes from having them involved. But I feel like we're not accounting for um, um, all the human ingenuity that went into the text on the web.

Speaker A: But obviously the fact that all that data is open allowed you in the case of xai in short order to get back up uh, to pretty close to the frontier. And I feel like that uh, that's been interesting to see that play out. How did that come about?

Speaker B: Um, yeah, so basically I um, ran into Elon, um, randomly happened to be at the Tesla office and talking with the team there about what they're up to. Um, and he happened to be around so I got a chance to talk to him which was uh, pretty fascinating because he's ah, such a great thinker as well, so thinking very deeply about humanity, technology and everything in between. Um, and um. Yeah, so I think we kind of saw eye to eye on all these things like the idea that maybe there should be another competitor in the world. Um, another AI lab that has a bit of a different mission, uh, different philosophy around um, how the models should be developed and um, basically how they should behave when they're out in the wild. Um, um. So we kind of had that connection from then and then. So when later on he wanted to start uh, something new, he reached out uh, to me and that's kind of how Xai, um, started. Um, and then from then on I talked to some old friends, uh, that were looking for something new to do and we managed to get a pretty good group of people together. And I think that's critical for any uh, AI effort. You have to get that critical mass of talent, uh, the right kinds of people that want to work pretty hard and do something great. That was uh, really the number one, the question of whether we were able to do it with the plant, on whether we got the people together. Um, um. And then from there starting pretty much uh, from scratch, um, we sort of slowly built up the company and uh, trained the models. Um, and then within less than two years we're able to, to get to the frontier.

Speaker A: We did some pretty crazy things, uh, famously like getting Colossus up and running, I think in what, 120 days and I guess what lessons did you take away from just the speed at which you were able to do a lot of that stuff?

Speaker B: Yeah, I think it just requires a tremendous focus and thinking outside of the box. That was really very much driven by Elon because he saw the need for us to bring more GPUs online very quickly, to be more competitive, to train grocery in particular. And so he decided uh, let's build our own data center. We don't want to depend on anybody else. And um, he's really the best person in the world to do that, um, from what I've seen because um, he's able to um, see the big picture of what needs to be done, what are the individual components and steps and then uh, sort of uh, in a very focused manner, uh, figure out from first principles how to accelerate every aspect of the project and then is able to go out and um, uh, uh, do things that other people might not actually be aware I mean did you

Speaker A: think it was crazy when it was like first proposed on timeline perspective?

Speaker B: Um, I did, but I also um, knew Elon well enough at that point that I knew that um, we'd get it done. So I also had tremendous confidence. Uh, and um. Um. Yeah. And then um. Really the uh, key to making clauses work was just questioning how everybody else was building data centers. Because out there, um, when we talk to people, how long would it take to build a new data center and put this many GPUs um, inside? We would get quotes way above a year. Um, and um, the reason is because uh, they had to bring up a new shell like a new data center building and uh, had this very waterfall like process for building it. And um, you also often deal with uh, subcontractors of subcontractors of subcontractors. So a um, lot of efficiency gets lost or can get lost in the process, uh, there. So Elon has the totally opposite mindset. We want to do it ourselves, we want to be the general contractor. And here are all these um, ways of solving our problems that other people haven't even realized are possible. So it's kind of how to um, find a glitch in the matrix on every single aspect of the project. So yeah, really fascinating working with him on this.

Speaker A: I'm sure our listeners would be super curious. What is it like working with him? And obviously how do you compare that environment to obviously you've worked in DeepMind, OpenAI3 probably very different environments and cultures.

Speaker B: I'd say that the reason um, I was really interested in working with him after that initial chat is that I could sense the respect that he has for engineers. So as a research engineer it's kind of um, jumping between research and engineering quite a bit. Um, and um. Uh yeah, I wouldn't say that uh, there's necessarily uh, a double standard at this AI Labs, but definitely um, um. Uh, you've got titles that kind of reflect your research scientist. You're a research engineer. And I uh, just felt really respected by Elon and uh, felt like we were seeing. I tell you, he was another team member know that I was working with. He focuses a lot in engineering meetings. Um, so really talking directly to the people doing the actual work um dives tremendously into the details of everything that the company does. And you can feel the respect for the engineers, the respect for the problem that they're working on. And um, yeah that was a really fun um, aspect of it. And also um, he has this tremendous energy and every day he Comes to work with, with a smile and you know, wants to tackle all the hardest problems, uh, once again from the beginning. And um, that energy, you can feel it in the team as well. And I think that's, it's. Uh, they come in with tremendous energy and um, and motivation.

Speaker A: Since you left the, the company said some pretty interesting things. Obviously getting acquired by SpaceX and then you know uh, acquiring Cursor. What would you make of the Cursor acquisition?

Speaker B: Yeah, I think it's an incredibly smart move. I think it's, it's uh, basically allowed him to jump way ahead on uh, coding models. So Cursor both has a great team that's very familiar with coding. There's a product that a lot of people out there use and I've also used Cursor quite, quite heavily. Um, and uh, there's uh, a lot of data that you can use to uh, re. Kickstart your training and um, uh, tremendously improve over what all the states where Xai was at.

Speaker A: Yeah. And so is that coding data, you know, from the real world, from people using it, like you know just that much more valuable than kind of the stuff you get from data labelers and then like the, the other suppliers.

Speaker B: I think it's the quantity of it is um, is really, if you have it in quantity that's really, really um, valuable. Um, there's also a tremendous value in building um, uh, building the right kinds of RL environments. So basically setting up these verifiable tasks where the agent is tasked with doing something, fixing a bug in the code base and then afterwards you check did it actually do it correctly and also did it do it in the right way? That's um, really what's um, been able to deliver a lot of improvements in uh, coding models. Um, and that's kind of the other side of the coin. You need kind of one both types of data in the process.

Speaker A: Yeah, but obviously it's way easier to figure out what environments to build. Having seen these things deploy. I'm sure way easier when it's actually your own model in uh, the, in the cursor harness. But uh, even Claude or whatever in the cursor harness probably still really valuable.

Speaker B: Yeah. And obviously I don't have any insight into what kind of data the Cursor team brought in and how they were able to make such a huge leap in model performance. But I'd say it's super impressive how quickly they managed to ramp up the coding abilities of the GROK models. And I'm sure we haven't even seen the biggest jumps yet. So there's some additional models that are coming out.

Speaker A: What do you feel like the biggest constraint is on continued model progress right now? Is it like, is it data? Is it, is it like just um, you know, kind of getting, yeah, get all this expertise that lives within organizations that hasn't been fed in or how do you think about that?

Speaker B: I think that the biggest um, unlock would be just new ideas around how to set up the training such that it can handle much longer time horizons, um, such that it can handle m non verifiable rewards much more easily. And I think there's a big, there's a lot of inertia and there's sort of a big energy barrier to getting these things to work and it's kind of easier to continue to iterate on the uh, coding regime. So um, I think one trap here would be just to continue to make the coding better and better. But I think what people really want to do is figure out how to move into those longer um, term domains. Uh, so for example, um, one quantity that's um, that can bottleneck your training is how long does it take to roll out your agent directories. So if you're um, in order to update the model you first sort of do trial and error. You have your agent attempt many, many solutions, you collect the rewards and then you can use reinforcement learning to update the weights of the model. But if your um, rollout takes 24 hours, you're going to have a really hard time um, training the model because every single training step will, will take very long. So we have to figure out ways of slicing up uh, the work that the rollouts, the work that the agents are doing and so being able to incrementally train them. Maybe agents have to guess do they do well, did they not do well based on incremental progress? Um, and then moving into all these domains where we don't have the rewards, where we can't have a unit test or we can't have a proof, um, uh, or a verification of a proof. Um, so that's really where people should be pushing today if they want to improve agents.

Speaker A: How close do you feel like we are to cracking some of the nonverifiable domain problem? I guess maybe back to your uh, AI scientists and the ability to kind of um, have a novel physics discovery or something like that. Do you have a gut intuition as to how close we are to some of that stuff?

Speaker B: I think that the models are so capable now that they can often make amazing judgments about whether something has worked or not. But we still haven't seen great examples of people um, implementing that at scale to improve the models further. So I think that's something that we might see um, now in the next few months, next 12 months or so, people actually pulling off these um, LLM judges, the approaches with LLM judges where really there is no verifiable reward. But we're still able to learn a lot from real world interactions.

Speaker A: Yeah, but you think people like the recipe is kind of known and it's just literally about running that experiment?

Speaker B: Yeah, I think people have maybe haven't figured out all the details of the recipe yet how to optimally do these things. But um, the thing that's been helping is just the models themselves becoming so strong um, now that you can often trust their judgment. They're often very reliable. Um, it's determining uh, what's happened. So um, that's kind of something that lifts, um, that's the whole, whole thing.

Speaker A: Yeah. I'm curious. You've obviously, uh, it seems like had conviction early in a lot of ideas that prove correct and obviously super early to reasoning and a bunch of these things in the last few years or maybe in the last year, anything you've changed your mind on that you kind of actually felt pretty strongly would happen in the AI world and has gone in a different way.

Speaker B: Well, the thing is, um, I think nobody could predict, um, how transformative the coding agents would be once they crossed a certain threshold of um, ability. So I was expecting kind of smooth progress on coding models. They just continuously get better and eventually they'll be used more and more. But this sudden explosion of usage, um, that was really something else. And now I'm starting to think maybe we'll see more of these transformations um, in the future.

Speaker A: Yeah, there's just like a step change moment where something, I mean I feel like everyone felt that around like November, December, um, and it's even hard to articulate even what it is, but it just like it just you know, starts working so much better. And you can imagine. Yeah, unlike other domains you'll reach some sort of uh, point where that becomes, becomes uh, apparent. But it's kind of hard to predict. Yeah, I think what's fascinating is even the folks closest to those models, like it was hard to predict a priori, like that's when the breakthrough was going to happen. I was thinking about where to end and one thing, you know, uh, maybe to tie it back to where we started, which was really around you kind of writing this fiction and thinking about like, you know, where the world's headed. And there's obviously these are questions that motivate river and you're founding there. I feel like you've written the question, obviously keeping many people up at night is if machines can do all this work, what's left for people?

Speaker B: I think the future, if we want to stay relevant in the world as humans, um, we want to figure out the right kind of symbiosis with the machines. Um, so today there is a bit of um, um, symbiosis going on. If I'm using a coding agent, it takes care of a lot of the low level details, takes care of tasks that I kind of understand at a high level. But I might not want to type out the code myself. And I still get to do the architecture, the planning. And um, I think that's a great mode to be in where human and machine are working together. But there's a big risk that the balance will shift over time. So as the coding agents and as agents in general become more, more capable, um, the incentive is uh, always to let them take more control, let them do more and more. Um, and this is kind of what the story is about because once you let the agent make all your decisions for you, you're not controlling your life anymore, you're not controlling the situation, um, anymore. Um, so I think what we should try to do is figure out how to keep the symbiosis going. And for me that requires new advances in alignment. Uh, so the kinds of alignment that we do for the models are often now around kind um, of learning from human preferences. So a lot of this group of humans, they like this kind of behavior, they don't like this kind of behavior. And we modify the model to comply with that. Um, but I think the future, in order for humans to stay relevant, it will need a much uh, deeper alignment, much deeper integration between human and machine. And uh, ultimately maybe there'll be something like neuralink that allows you to, to control large amounts of intelligence just by thinking. But I think there's still some time away. So in the meantime, um, we can align the models better by changing the way we train them, by innovating on the algorithms on the technology. And that's kind of what Revo AI is trying to do. Um, so I think that's um, uh something I'd encourage anyone who's in AI research also. Our intentions and helping us rather than uh, sort of taking over and replacing us.

Speaker A: Yeah, no, and it seems like in this context, you know, helping us is Is, is still giving us a role, basically, even if it's, if it's maybe suboptimal in some way of, of like, you know, maximizing whatever paperclip objective there is. Uh, you know, it, it is maximizing at least, uh, you know, human flourishing. And so that.

Speaker B: Exactly. And we should train the models to maximize human flourishing. That's um, that's another takeaway.

Speaker A: There was obviously this letter, uh, over the last 24 hours that many researchers signed, I think it was released yesterday, calling for the government to consider slowing down AI. And there's all this kind of concern around just um, how fast is all moving. Is that something that resonates with you? Um, how do you think about that?

Speaker B: Yeah, I think it's incredibly difficult to slow down at this point. AI has become such a big thing. Um, the US economy is very reliant on AI making, making further progress as well. There are international competitors out there. You know, China is quickly catching up in terms of AI capabilities and they might or might not decide to slow down as well. So I think it is really, really difficult to get people to slow down. I don't know how realistic, um, that kind of, um, the ladder really, really is in uh, that sense. Obviously I would uh, be, would love to be surprised if we do manage to slow down a little bit, do more work on alignment. I think that would be, that would be a good thing. But um, if we're not able to slow down, I think we should accelerate on technologies that can improve alignment, the AI models that can improve safety, um, because that's going to quickly become very relevant for the models that we deploy.

Speaker A: Yeah. And it's the best way to do that. Just more open source models and more effort within the closed source providers or if you, I don't know if you had a magic wand, what would help increase uh, that amount of work.

Speaker B: I think open models that are close to the threshold of where they might be start to become dangerous are really, really valuable because it allows anybody in the world to try out their ideas for how to align these models, control them so they're not so powerful that you can do real harm with them. But they exhibit some of these things that will become dangerous in the future. Cybersecurity capabilities are one of the biggest ones, um, right now. Um, so I feel like we don't want to just have the few thousand people at the big AI labs thinking about these problems. We should have as many people as possible trying to figure out what are the algorithms, what are the tools that we need to Build to make these things viable.

Speaker A: I mean, you've obviously been closer to this stuff than anyone. What's your current probability that this all ends up? Okay?

Speaker B: That's a good question. I mean, the question is, okay, for whom at this point? Because, um, potentially, um, AI, um, uh, could amplify inequality in the world and there could be people who benefit greatly from it, and then there could be people who kind of feel left behind. I think that's the most immediate, um, urgent, uh, risk. And then, yes, there are risks in the future around, um, can the AI models take over? Can they take control away even from the people that have built them, that uh, are controlling them today? And I think that's also a risk, but it's so hard to make predictions about it. I feel like that's going to be in the further, further future. Um, and the more immediate problem for that I can see is are we able to bring everybody else along on the journey? Yeah.

Speaker A: Well, last question for you. You had in your, uh, you know, in that piece you wrote the prayer to machine God, you had this beautiful line, I think that was, you know, uh, like the tears you didn't expect when the model first spoke back. And so I have to ask, was that from like a personal experience? Have you had something that like, felt like that?

Speaker B: Yeah, I put, I put that in there because, um, uh, I still remember how magical it felt to first get these models to work. And back then it was very, very modest compared to today. I think this is something that all the early folks in AI felt that when the models first started to do some pretty simple things, but things that you couldn't do with any other machine, uh, before it was, uh, just really, really magical. So whether it was just, um, writing a python function that can generate prime numbers or something like that, just felt like this is unreal. So, uh, yeah, I wanted to capture, uh, that feeling. I'm sure that's something that a lot of folks that were early in AI, um, felt. But now looking back, um, it all seems so unreal because the models have advanced tremendously. They're really affecting, uh, the entire world at this point. And it's gone so far beyond those little experiments that we, that we used to do when, when AI was more of a hobby.

Speaker A: Totally. Well, I think that's a perfect place to end. Uh, EREV has been fascinating. Thank you so much for, uh, for coming on the podcast and uh, and sharing all this.

Speaker B: Thank you Jacob. Appreciate it.

Speaker A: I'm Jacob Efron, and this has been unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's, uh, a nights and weekends project in addition to my day job as an investor at redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so please consider doing that.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Utilizing AI internally to iterate faster and empower smaller teams to upskill w/ Vivek Raghunathan #263The Engineering Leadership Podcast · on coding agents96 / 100
  • AI-Powered GTM: 3 Founders → $30M ARR with Amos Bar-JosephGrowth Activated · on coding agents95 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on coding agents86 / 100
  • #2307 Eric Ries: Why Anthropic Won and How To Build Incurruptible companiesStartup Stories - Mixergy · on AI Safety86 / 100
  • Beyond SpaceX: War, AI, Orbital Infrastructure and green utopia | Mark Boggett | Seraphim SpaceFund Shack Private Equity Podcast · on xAI82 / 100
  • Anthropic Code Leak: A Rare Look Inside Frontier AI | EP.52Hidden Layers · on coding agents82 / 100

More from Unsupervised Learning with Jacob Effron

All episodes →
  • Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI Today85 / 100
  • AI Vibe Check: Lab Wars, Why APIs Might Vanish & Future Predictions73 / 100
  • Ep 89: AI Research Legend’s Honest Assessment of Where We Are80 / 100
  • Ep 88: Unpacking DeepMind's Quest for SuperIntelligence with Demis Hassabis' Biographer82 / 100
  • AI Vibe Check: Chinese Open Models, Distillation & The Hugging Face Breach
Explore the best B2B Engineering & DevTools podcasts →
All Unsupervised Learning with Jacob Effron episodes →