Unsupervised Learning with Jacob Effron · 2026-06-03 · 1h 14m
Key moments - from our scoring
Substance score
60 / 100
Five dimensions, 20 points each
Lukas Kaisel, Transformer paper co-author and former researcher at Google and OpenAI, explores whether current transformer architectures have hit a fundamental ceiling or whether better generalization is possible. The core tension: transformers achieve remarkable feats with chain-of-thought reasoning and reinforcement learning - solving research-level math problems, coding at near-intern competency with Claude and o1 - but require massive data ("trillion tokens") to learn concepts that humans extract from far less. Kaisel argues this isn't purely a scaling story; something fundamental may be missing, though the evidence remains intuition rather than proof. He discusses how agents (like Claude) have transformed his research productivity by 5-10x, examines the unexplained jump in coding capability last Christmas (involving unknown combinations of harness changes, post-training tweaks, and pre-training improvements), and addresses why projects like Waymo's self-driving struggle with highway construction zones despite millions of simulation miles - a generalization gap humans don't face. The episode will interest researchers, AI leaders, and operators assessing whether transformer-centric scaling or post-transformer architectures (pursued by labs like NEO) represent the path forward.
Kaisel argues transformers require exhausting all surface-level patterns via massive data before learning true concepts, whereas humans extract concepts from far less data. He suspects something fundamental - possibly recurrence-based architectures or a different learning mechanism - could enable better generalization, though he acknowledges this remains intuition rather than proven fact.
He reports approximately 5-10x productivity gains: reproducing old papers took three weeks before Claude, now takes two days. Beyond speed, agents improve his mental workflow by letting him think in terms of machine learning concepts (losses, batches) rather than low-level code implementation details.
Kaisel is uncertain. The jump involved multiple simultaneous changes - harness improvements, post-training tweaks, and new pre-trained models - making it hard to isolate the driver. Unlike the clear architecture shift from RNNs to Transformers, this jump remains "a little hard to pin down."
He suggests the issue may require architecture, data, loss, and optimization tweaks working together - not just the transformer itself. The fact that humans handle construction zones identically across city and highway contexts suggests transformers may be missing something fundamental about concept learning, though better training methods could potentially close this gap.
Kaisel suggests long-horizon reinforcement learning on weeks-to-months-long research problems, aggregated across thousands of researchers, might enable this. However, current RL methods require impractically long rollouts. Alternatively, a post-transformer architecture with better concept handling could be necessary - humans somehow learn to do research without needing 200 practice problems first.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of genuinely interesting observations - the LLM-learns-concepts-last analogy, the jagged generalization point, the RL verifiability spectrum - but is padded heavily with hedging, repetition, and statements like 'it's a feeling' and 'I don't know.' The ratio of novel claims to filler is moderate at best.
Americans will do the right thing after exhausting all other options and LLMs. They will learn a concept. They will learn it. But after exhausting all other options, you need this trillion tokens. You need to learn all the surface level things. And only when that doesn't explain something, they will finally learn the concept.
it generalizes, but in its weird, alien way. And that just doesn't cover some ways that I can generalize.
A few framings are genuinely fresh - the LLM-last-resort learning analogy, GREP as a hack-that-works for long context, and the verifiability-as-spectrum argument - but the dominant themes (transformer vs. post-transformer, scaling persists, Anthropic bet on coding while OpenAI did chat) are well-worn AI discourse and largely un-challenged.
GREP is our solution to long context is let's write bunch of stuff in files and give it access to grep so it can find and tell it to write index files and it's like a little library. And of course to me as a normal researcher, you told me five years ago, that's not a solution, that's a hack.
every hole you have, you can kind of plug by hammering on it. But it would be so nice if you didn't have to
Lucas Kaiser is a co-author of 'Attention Is All You Need' with senior research stints at both Google Brain and OpenAI - about as credentialed an ML practitioner as exists; the transcript partly squanders this with hedging but the underlying expertise is genuine and surfaces in specific technical observations.
nobody talks to me about the paper I had before. Attention is all you need, which is you don't need attention. I had the paper at Neurips the year before. Active Memory wasn't quite good advice
the GPUs we research transformer on, they had nine teraflops and we had eight GPU machines. So in absolute scaling you could say be like um, 70, 80 teraflops for real on the machine. Um, so now I have under my desk something that's like five of these machines in one gpu
The episode scores above average on specificity thanks to concrete GPU benchmarks, a quantified productivity comparison (3 weeks to 2 days), named companies (Harvey, Waymo, Gemma), and a specific failure example (Waymo highway construction zones); however, the majority of substantive claims about future trajectories and architectural alternatives remain vague and unsubstantiated.
I knew it took me about three weeks to get to a runnable state. And with codecs I could get there in two days. So it's about, let's say a week to a day.
the 5090, it's about 200 teraflops... the GPUs we research transformer on, they had nine teraflops and we had eight GPU machines
The host shows flashes of good follow-up instinct - pressing on why meta-level RL hasn't been tried and asking why Anthropic won coding - but largely defers when the guest hedges repeatedly, accepts vague non-answers, and closes with a puff-piece softball; the overall interview dynamic is collegial rather than rigorous.
Why doesn't that work today? Like, I'm sure people have tried that.
Why do you think Anthropic was the first to be really successful on the coding side?
Computed from the transcript - who did the talking, and the words that came up most.
This episode with Lukasz Kaiser, co-author of the seminal "Attention Is All You Need" transformer paper and former researcher at both Google Brain and OpenAI, is a wide-ranging conversation about the fundamental limits of current AI architectures and whether transformers will continue to dominate or eventually give way to something new. Lukasz brings a rare dual perspective: deep belief in how far the current paradigm has taken us (he's an enthusiastic daily Codex user who's seen 10x productivity gains in his own research), while maintaining genuine intellectual humility about whether transformers can truly generalize the way humans do. The episode weaves together questions about data efficiency, the non-verifiable RL frontier, the coding agent revolution, the open vs. closed source gap, and what the next architectural leap might look like: all filtered through the lens of someone who helped build the foundation the entire field is standing on. (0:00) Intro (1:12) Transformers vs. Human Learning (8:37) How Do We Get Physical World Generalization? (10:52) What Comes After Transformers (13:59) How Much Have Agents Improved Lukasz's AI Research Productivity?
Transcribed and scored by The B2B Podcast Index.
Host: Is reasoning enough to get to generalization or is another method needed?
Lukas Kaiser: It does feel like there is something else that possibly could generalize much better.
Host: Why do you think Anthropic was the first to be really successful on the coding side?
Lukas Kaiser: Anthropic made this very good decision to focus on coding. OpenAI was like, we're doing ChatGPT hard way. Anthropic made this decision was that they just could not compete.
Host: What's your kind of gut intuition on the gap we'll see between closed source and open source models and whether that widens or shrinks in the next few years?
Lukas Kaiser: I think it's a fair question, but
Host: Lucas Kaiser is one of the authors of the Transformer paper and has had amazing roles at both Google and OpenAI on unsupervised learning. I got to ask him all the top of mind questions of what's happening in AI today. Of course we had to talk about the Transformer, uh, and how he thinks about its persistence and whether it will remain the dominant architecture and what its shortcomings are. We also got his thoughts on what changed in the fall to really make coding models so much better and why Anthropic was really first to code. We talked about with the future research directions that he's really excited about. And we also hit on a bunch of things around how he thinks the ecosystem will evolve from open versus closed source model to application companies. I think folks will really enjoy this episode with a top researcher whose research, uh, really set off a lot in the space. Without further ado, here's Lukas. It's a pleasure to have a Transformer, uh, paper co author on the podcast. I feel like you've been at the forefront of so many major changes in the AI, uh, world and our goal is really to get your thoughts on all the questions around the AI, uh, frontiers today. So I really appreciate you coming on the podcast.
Lukas Kaiser: Thank you very much. Thank you for having me.
Host: I can think of kind of no better place to start than generalization, right? It feels like that's the question in the air right now. Um, and I think in November I heard you say, you know, basically the big, this big question of is reasoning enough to get to generalization or is another method needed? And I'm wondering, I guess you said that, you know, maybe six, uh, months ago now, which is, you know, dog years in the AI world. So years ago, uh, how has your thinking on that question evolved since then?
Lukas Kaiser: If we take the current transformers with reasoning, right, and agents, and they have access to a shell and Stuff they can do amazing things, right? It's incredible how far we've gotten. Like two years ago even. Not to mention before Transformers. I would have never believed that. You just take this next word, predictor, give it then chain of thought and RL that and tools and that it will. I know every day spend hours talking to codecs in my case or other people. And it works, right? You talk to it about hard problems at work and it makes sense and it implements things. So that's incredible. On the other hand, there is this feeling, um, that it is not quite like us, that it's not quite at the edge of what's that we all feel that it possibly should be even better, right? That we can generalize from less data, uh, somehow make bigger leaps, get these concepts from way less. I recently have this saying that people say Americans will do the right thing after exhausting all other options and LLMs. They will learn a concept. They will learn it. But after exhausting all other options, you need this trillion tokens. You need to learn all the surface level things. And only when that doesn't explain something, they will finally learn the concept. That's not how we learn. We just get concepts from. Sometimes we make them up and they're not great. So it does feel like there is something else that possibly could generalize much better, that could possibly have this like a bit of a different form of understanding, more like long term. Um, but it's a feeling, right? And every time we try to put our thumb on seems to evaporate or more like maybe it doesn't even evaporate, but it's like the transformer just catches up, right? It was like. So both sides in this time have grown, right? Like Transformers have gotten even better, but the case for something else has also gotten even better. I would say there's now a number of labs that pursue post Transformers and people see interesting results. There are certainly interesting things out there. So who wins? I still don't know, to be honest. I think there's good arguments for both sides and it will be extremely interesting to see how this goes.
Host: I think it'll be interesting for our listeners. You know, you, you obviously I think at talk ah in near Con more recently like alluded to this like whiff in the air, right? Uh, that there's something that, that that's happening in progress that's inspired like these NEO labs and other folks to spin out and you know, work on things that are maybe alternatives to the dominant architectures that are being worked on within the labs. What Is that feeling, is it seeing some of these early results or, or what is it like? Or is it just like researchers intuition? Like there's maybe making it a little more concrete for uh, our listeners.
Lukas Kaiser: I think a lot of this is intuition and you need to beware because it's like a lot of this happens in San Francisco at parties and people talk to each other. So it may be, or on podcasts. So it may be that it's self fueled to some extent. Uh, but I think there is a part of it that's very fundamental. I mean, Yann Lecun has been saying something like this for years, way before now, which is the models we have in a long, long history, they were meant, they're called neural networks because they were meant to imitate our brain. But they don't really, they are quite different even if they may have some similarities. And if you look at how humans learn what we can do, it is quite hard not to say that the, from much less data we can do much more than our current models. So it feels like there is this fundamental ability that we as learning machines have that our models currently don't. So fundamentally there should be something there, not just a vibe. Now you can say as a counterargument that these models always had a trillion tokens to train on and people never do. So we just didn't optimize them for training with less. And if you, you know, if you had the same amount of compute but limited data, you can tweak transformers to do much better than they do today. So you know, it's like some people say, why would you. Right? We have the data now, it's a big enterprise, but it does feel even when we, even when we try to push with as little data as people or. Well, it's also like we get a lot of data from visual things from moving in the world. We take actions. So it's very different kinds of data. Uh, it's not truly comparable. That's why it's hard to make a very firm scientific statement about it. But there is this feeling that we have not exploited all that is there, uh, in machine learning. And obviously the exciting feeling is that maybe if we find out what's out, it could make what we have even more amazing. Maybe not, maybe it vanishes when you have that much data. Who knows, right? Uh, but it's definitely extremely interesting to me as a researcher and I think to many people. It's like, I mean transformers were fascinating, right? They're great reasoning is uh, I mean it can Solve research math problems. I'm, uh, sure you've heard about the recent erds things and I was a mathematician before in my life, so this is extremely exciting. I never thought a computer in this time frame will talk to me about mathematics at a high level. As a real researcher that exists now, and this is insane. But then as an ML researcher, I'm like, okay, but we haven't really figured out this learning. There is this feeling that it learns, certainly, but it needs so much data. It needs so much compute. This feels like we're not quite there yet. Now is this only a feeling? Is this a vibe? It seems to be reality to some extent. Right? But we'll need to see the research
Host: appeal of figuring that out. Makes a ton of sense. And I think other folks might look at it and be like, well, so what if it's not like people, right? We have the data, we have a method that works. Um, you know, obviously there's going to be some areas where there is limited data, like, you know, medic, like, you know, drug development and other things where learning from more limited data would be really helpful. But so many problems that exist in the world actually aren't that data constrained, right? Sometimes I feel like these sides almost like talk past each other, right? Like people at the labs will roll their eyes at Yann Lecun or something like that.
Lukas Kaiser: I think this is fair to say. But on the other hand, given how quickly and you know, with the whole investment in AI, the problems that are not data limited get solved very rapidly. So very soon all bottlenecks that remain will be quite data limited or already are becoming. And in particular it does feel that to work well in the physical world, you do need to solve some part of it, at least, because the physical world, if you train on one robot hardware, it doesn't quite scale data the way that the virtual or text worlds, uh, or Internet worlds do. So in the physical world is a
Host: sizable chunk of it though people are certainly trying, right? With simulation data and with egocentric video data or cheaper sources.
Lukas Kaiser: I mean, I'm a huge fan of Waymos, right? They have always this joke, like people say, where are my self driving cars? Well, I drive them. They're here. But then they just canceled the highway driving, right? Because they couldn't deal with some construction zone again. And it feels almost like they have had this construction zone things for years in this. I'm sure there's millions of miles in simulation and quite some in real driving. And it still can't generalize to the construction Zone on a highway. This feels just off. Right. I mean, I don't know what exactly didn't work there, but I certainly know no teenager has this problem.
Host: Yeah.
Lukas Kaiser: Or no human. Right. We have many other problems. Right. But not that we can drive in the construction zone in the city, but not on the highway. Construction zone is a construction zone.
Host: And do you think that some of this stuff will be or could be solved within the transformers? And I guess, what are you kind of looking for, I guess in the next few years to get a better answer to this question?
Lukas Kaiser: The exciting part in ML research is that it is so broad. Right. You never know whether you need to tweak the architecture or do you need to tweak the data or do you need to tweak the loss or do you need to tweak the optimization process? Um, and they're fair arguments for all. And on top of that, it might turn out that you need to tweak all of them to some extent. Right. It's like transformer is great, but it's also great with the next word prediction, loss. Right. Or you can make it work with rl, but you need the chains of thought. Or it's like these puzzles only work when you click them together. So it is possible that if there is a new thing that there might need to be tweaks to everything. But it's also possible that parts of transformers will survive, for example. Probably attention will be somewhere there. Right. Um, but maybe you need other things to it. Yeah, like maybe. You know, I've started my machine learning life with RNNs and, and I certainly hold recurrence deep in my heart. I, I like it as a construct. It feels and reasoning kind of brought it back because every new token you produce, it's the same weights that that we currently produce it. So, so in some sense it's, it's back. But, but, but it does feel that this RL is like very sparse losses and do so much. Um, but it works, right? And every time we try to do recurrence in other word ways, it somehow does not seem to click yet. But there is always the question, how hard have we tried? Uh, I don't know if you or the audience knows. There are models like TRM and hrm. It's like very small models that turned out to do very well on problems like Sudoku, but also rkgi. So they're a little bit toy tests, but they do quite well. I think a lot of the post transformer architectures are trying to merge this with LLMs. It's interesting, certainly, right. It's like the pure transformer can't do so well on it. But you add some recurrence, you add some bit of architectural tweaks, maybe a little different loss, and it does really well. So even on the small scale you can do a lot. Uh, but then will it generalize to the language and give you the things you want? Well, it will be very interesting to see and luckily there is a number of labs that are trying. And the other thing though is this year we have the agents and to me this is a totally. It's the biggest change in the way I work as a ML researcher in, I would say the last 15, 20 years probably.
Host: I don't know if you try to quantify it, but how much more productive do you think it makes you?
Lukas Kaiser: Oh, I can really well quantify it because I tried recently just on the private machine to reproduce, uh, a bunch of papers, like old papers that I was always quite interested in. Um, even some of my papers that I lost code for. Uh, and at least one of them I tried to reproduce before. And I knew it took me about three weeks to get to a runnable state. And with codecs I could get there in two days. So it's about, let's say a week to a day. That's already whether it's a 10x or a 5x. Maybe I could have been faster back then, but um. But it certainly changes your rhythm because you can just take on things. It's also, you know, I can just start three things in parallel and let it go. While when I was doing, I would just do one thing usually. So it both makes it faster and makes it more parallel with. Which is, uh. But when I do private things, not in the production repo, I basically stopped looking at the code. Uh, a friend asked me, do you think you're less sharp now? I gave it some thought and I think it's actually to the contrary because due to the fact that I don't look at every class name every small function, but I still know that these agents can go off the rails with this paper. It runs something and there were some aux losses and it just added. It just thought it should have another auxiliary loss and it was totally off the charts and out of place. So you need to have a, uh, full control in your head of what exactly is it doing, what is the loss. But you don't need to have control of what's the name of the class, you know, what are the exact words in the function. It's quite impressive that you can trust the agents to be trustful in that they're really implementing what you think. But I mean, sometimes we check and um, they are. But since you need to have a full control in your head of what's actually running, machine learning wise, what are the losses, what are the batches? Then I feel it gives me actually more mental control over what I'm doing than it was before because before I would implement it. But sometimes in this time before running it, I would have to forget a little bit about what the big picture was exactly and focus on the little things. Debug, then go back to the big picture. By that time, maybe I forgot some detail and then I would remember it when it was wrong. Now it's this beautiful thing where you can just be in this flow. You just think machine learning wise, what's supposed to happen. You tell it, verify it, and it's happening. It's not just about the time saved. It just makes the work so nice. It's a mild psychosis, I guess among researchers. We just can't stop.
Host: OpenAI very publicly said. Hey, our goal is kind of, I think a research level intern, uh, by November of this year. Know as someone who plays around with codex all the time in your research, does it feel like you're. You're close to that or how are you feeling about that milestone?
Lukas Kaiser: It does feel like close to an intern. But you need to be very carefully checking. Like as I said, it can just add you a loss that you did not ask for because it seems reasonable to it. Um, I don't know if interns do that. Maybe sometimes, I guess sometimes when they're creative, we. But like I try sometimes, you know, it's like I will just let it go for the night and I give it the goal, you know, make a better model for this lower perplexity. Um, that never works. It will just start doing some very trivial tweaks that are not really interesting or useful. So it's certainly not at the level of a researcher.
Host: Yeah, dude. What's the path forward to make it better?
Lukas Kaiser: There is it. It goes back to our question. For a long time I worked on long context in machine learning. Even before Transformers, you could say on memory and so on. And then we worked on it with Transformers and the context got longer. We got a million tokens, which huge given what attention does. But now with agents, it really does feel like GREP or rip. GREP is our solution to long context is let's write bunch of stuff in files and give it access to grep so it can find and tell it to write index files and it's like a little library. And of course to me as a normal researcher, you told me five years ago, that's not a solution, that's a hack. Right? But you know, machine learning everything is a hack of sorts. So dropout is hack. We don't judge, right? We take what works and it works. Amazingly, it works. And you add a little bit of rl, like for example compaction. If there is one reason I like codecs over cloud code, it is compaction. You can go on with the thread and it's good at compaction. Why is it good at compact? I don't think there's like very mysterious, right? People prompted it well and then put some RL on it to just make. And if you told this to me like some years ago, that the long context where you just irrelevant that it can use tools and find stuff in files and then summarize good enough to keep the. I would tell like, okay, that's a band aid. That's not a. Like it doesn't feel like this deep thing. But we don't judge solutions by how they look, we judge them by how they work. And it works really well to the point of can it become a researcher? Well, some people would say, well maybe, no, maybe you'll need this new architecture. Maybe you'll need a post transformer thing that has concepts that are bigger and follows goals and, and it's a fair argument, right? Currently it feels like it cancels, but then there are other people who say, well, well, well, you'll have your conversations with Codex for a month and then you're going to prompt it to go over them and find meta patterns, write us to some files and just think how it can use them. And you know, maybe if you have some data, over a thousand people and do some RL on it, it will start behaving like a researcher. It's. In some ways this is how researchers learn, right? We look how other people do research, we do a bit of our trials, see what works better.
Host: Why doesn't that work today? Like, I'm sure people have tried that.
Lukas Kaiser: I know, I don't think people have tried yet very hard. It, you know, it's like some people do some prompts and they work for them. Um, it is important to, I mean to me the Codex era started like this year or Christmas, right? It's, I mean codecs existed before and we used it and cloud code existed and we also used parts of it
Host: But I think everyone felt at Christmas.
Lukas Kaiser: But it seems like only the newer. But it's not just the models, it's also the harness and the some tweaks. So you know, it's barely half a year and there's still many people if you go a bit outside of the, you know, our SFAI bubble that totally don't get it. They're like, you know, you're a little psychotic. But why? Right. And I think it's a fair question. But it started working very recently. We don't truly understand like it was not like a big pre training that changed it that much even though big pre training came too. But when we went from RNNs to transformers, it was very easy to attribute the change to this. While now I feel. And then there was reasoning which clearly is important in rl. But the change last winter, last Christmas, it's a little hard to pin down. I mean harness changed and the little post training change and then new pre trained models come which of course made things better but. But it felt like a big jump which is not that easy to pin down what did it. So it's a little bit messy. We improve everything all the time. But then because it works and it feels so important, there is also the necessity to just bring it to people, make it work everywhere, promote it. There is this competition going uh, on. So I think in all of this people did not have truly the time yet to think how do you really do this meta level and people are starting. But it also feels because the meta level is something you research for a week and then you get some patterns and start applying them. That feels like RLLync. This needs to take weeks. Our current reinforcement learning methods and luckily need to run basically all rollouts on this. And if your rollouts are weeks long, then your training starts to be months long and that all becomes a little impractical. Which maybe is an argument that that the post transformer that the human side has something to learn. Because clearly humans can do research over years and they do this once in their life or twice. Some mathematicians spend 20 years on one problem. That's their magnum opus and that's it. So they did not have 200 problems 20 years long before to learn from and somehow they manage. Um, how does this work? It's a fascinating question. Clearly with some relevance to this. We haven't figured it out. But on the other hand now we'll gather, since a lot of people work with it, we'll gather a lot of the data on the weeks to months long Humans, someone will run this RL and it may just turn out that, that it gets you.
Host: Yeah.
Lukas Kaiser: Further.
Host: So it's such an interesting point because basically, you know, like you're saying as folks were scaling pre training or as folks were scaling the original sort of reasoning models, it was kind of uh, it was straightforward or at least made sense like the vector you were scaling on and then this kind of big advance we've had in Codex and cloud code over Christmas. If you don't actually know what the source of that is or not, you're not fully crystal clear on it, it's very hard to uh, then determine what you should be pushing on to continue to improve these capabilities.
Lukas Kaiser: Yes, it's a little, it's a little confusing. I mean the fact that I don't know doesn't mean nobody knows. I think maybe some people have uh, like stronger opinions on what exactly pushed it through. But, but I don't think it's that clear at this point. Uh, which, yeah, it's just, I mean on the other hand, it's been improving for a while, but something happened.
Host: Yeah.
Lukas Kaiser: Because yeah, it did not feel possible to do this and now it does
Host: on this kind of current scaling regime. On the RL side, I think one question a lot of folks have is we've seen obviously tons of uh, coding improvement and math and these kind of verifiable domains. And I feel like the two big questions around RL continue to be, you know, how well is this going to work on the non verifiable side? And then also, you know, the extent to which we'll, we'll get generalization and not having to keep, you know, do tons of data in each space. Maybe we'll take them one at a time. But starting with the first, you know, how do you think about the problems that need to be solved on the non verifiable domain side and you know, any inklings as to which spaces might be, might uh, be next beyond code and math?
Lukas Kaiser: I do think there is uh, there's been a fair progress on the nonverifiable side. If you look for example at like things like Harvey and Law or things in medicine, they're not verifiable, but there's a lot of parts of them that are verifiable. So there's been good progress on that. And I think, I mean GDP VAL is one benchmark that in some sense benchmarks things like that too. And I do think there is really good progress and there's really good incentives to make, make progress in These domains, I'm not sure if it's fully fair to call them nonverifiable.
Host: They're certainly not as perfectly set up as coding in math.
Lukas Kaiser: Right, they're not coding in math. Right, but math. I think people overstate how verifiable math is. Coding is fairly verifiable in the sense programming competitions are verifiable. Once you go to front M end coding and stuff, it's also not that verifiable. But, uh, still, the proofs are not that easy or clean. I mean, you can do lean, but most of the math, at least from GPTs, it's not formalized, so it's not that verifiable. So it's a spectrum. And then things get less and less verifiable. I had this pet project of translating poetry into polish, which seems fairly not verifiable. But then you run these models as verifiers, and, you know, they get a fair bit of stuff. They get like, rhyme and things, and they can get cultural references. So it turns out, uh, once you read how people have verified things before, you can get to some level of verifiability. Um, but then, I mean, I think what this poetry thinks, that was also what it was meant to show is you can verify a lot of things and then have. Still have kind of no taste. And. And it's. I mean, since it's not verifiable, it's not so easy. If it were easy to describe in words, then it would be verifiable, but it doesn't mean it isn't there. You read this, and there is something in your brain that reinforces this idea that there is something they're missing. M M We have driven ourselves into this whole basically on purpose, because what is reinforcement learning? It tells you whenever you have a teacher, a validator, someone telling you, this is good, this is bad, I can train against it, and I'm going to get good. And that's what the models do. So, you know, whenever I will come and say, look, I don't think this does this very tasteful in something. Then, um, someone will tell, okay, show me. And then it will nail it. And I think some people even run rants that go basically against, like, for image generation, you can ask, is this beautiful or not? Okay, not verifiable. But you just get a bunch of people who, during training, click, this is beautiful. This is them. And lo and behold, the images start to be more beautiful. So the verifiability thing is very weak. Right? You can. It's Just a uh, very sparse signal. When you ask, you can ask people, is this nice, is this not nice? Which then, you know, how do we, you know, why do I think this is not very tasteful, right. Clearly some of my experiences and some way that I have processed it make me say this statement. So why does the model not say it? There are two possibilities. One is that it hasn't seen enough experience that would make it do it. And the other is that it's not processing it in the right way. So I believe in both actually. But even with the way it is processing, if you just put more experience, you ask a thousand people to tell it, then it gets better. So every hole you have, you can kind of plug by hammering on it. But it would be so nice if you didn't have to, right? Because also every hole you plug stops being a bottleneck and then the bottlenecks that emerge are again the holes that you have not plagued. Um, um, so we're in this interesting circle, but if we had this method, this brain like method that would just not have so many holes that need plugging, then wouldn't this be great?
Host: Does that kind of imply that any problem area that someone does focus on under the current architectures can be figured out? It's just that to your point, it probably is, is far more, you know, requires curated data and far more manual than, you know, a potentially more beautiful way of doing things down the line. But there's not like a set of problems or a set of domains that you're like God. Under the current uh, RL methods, you know, that would be too hard for uh, the models.
Lukas Kaiser: It does not feel so. But you do need to take economics into account, right? I mean currently to make these models work really well, you need to start from a fairly strong model which is fairly big and expensive. On top of this, it's usually closed. So you can't um, really you do it. I mean There is the RL Fine Tuning API which I quite like from OpenAI and some similar ones, but you don't truly have like full access to it. So even with the API it can be a little hard. And even on top of that, the investment you need to make into the data and things, it's substantial, right. You couldn't do this. You'd need a company, you need some contracts you'd need, which, you know, if it's important enough, it's a fair method. But then, right. Wouldn't it be great if you could just talk to the model and uh, it would work on Its own.
Host: Does it feel like there's any signs of general capability improvement as you do? M. You could imagine a world where it's like, okay, we'll start with code, and then we'll do math, and then we'll, you know, do this for legal and healthcare. And you could. You could tackle each of these one by one. Even if you're not getting any sort of generalization across or, you know, ideally, I think that maybe the hope would be at some level of. Of having done reinforcement learning on a bunch of different domains, maybe similar to pre training at. At some level, like, generalization emerges or something.
Lukas Kaiser: But I think generalization emerges in reinforcement learning.
Host: So you think already the models get better across the board?
Lukas Kaiser: Oh, yes, they certainly do. Like, if you look at, like, I think law is simply not in the RL pipeline at all, and you talk to Harvey or someone, and they say it either emerges or they need a little train. Just a few touches on top of it, and it suddenly catches it. So there is definitely generalization, but the generalization doesn't seem just to go as far as. Or it just works in not the ways that we would hope. Sometimes it doesn't generalize even from math to other areas of math. If you look at even the IMO right now, it seems so far away that models were imo, but it would have some types of exercises. For a long time, it was geometry that it just couldn't crack. It would do very hard. It would solve very hard problems in other domains. But in geometry, you were like, oh, okay, it has no spatial understanding. And then it just saw more data and started cracking it. But not spatial understanding data or physical. Just smart geometry problems. Um, but it has this jaggedness, right? It will generalize from here to here, but not to something that seems very close, but somehow in this representation of these chains of thought is not like it's close to me, but it's not close to the model, right? So it's not like it's not generalizing. It's generalizing, but in its weird, alien way. And that just doesn't cover some ways that I can generalize. And it's possible that with more data, it will just cover more of this space. M. But I also understand people who say when it's like that, it's very hard to trust, to commit to it because, you know, there may be this spike that it just hasn't gotten. So you need to be on the lookout for problems. And as I use it as ML researchers, I think it keeps me very honest because I think it keeps me sharp. So maybe this is good in this way but, but it's not good in a capabilities way because you just hope that doesn't have these sharp edges. Right? For now it does.
Host: You mentioned some of the application companies obviously that benefit from these models getting better. And I think there's like this big question of if you're an application company right now, should you be working super closely with one of the labs and sharing kind of all these evals and like understanding you have of the domain or you know, is that like actually, you know, are you better off kind of building almost your own model based on information versus you know, sharing it back? I'm curious how you think about the room for applications on top of uh, the core models.
Lukas Kaiser: For now, what is certainly true is that the bigger and better your pre trained model is, the less of these sharp edges you get and generally the easier all of your life becomes. Whether you do RL on it or fine tuning on whatever bigger model, things just get easier. It's insane how this has continued to be the case. We have uh, I don't know, you remember like a year ago, two years people were saying oh LLMs are that SLMs are the future small models and we have amazing small models like the Gemmas that recently were like few billion. Remember GPT three people said oh you go, you do no zero shot learning under 100 billion. No, we have like you know, three B models that are so that's all amazing. But if you really want to solve big problems, easily adjust ah to your data and context, there just doesn't seem to be anything like a really elephant model. Uh, but they're of course expensive and hard to use and even harder to train.
Host: One thing I think would be interesting for our listeners is I think something that's maybe less obvious to folks outside the cutting edge is just what's enabled by new generations of hardware. Right. And so I'm wondering if you'd speak a little bit about to. I mean you know obviously it seems like for, for, for certain things as, as we, you know, uh, as we waited for like blackwall chips to come online, it's like hey, they came online and like the models got better. And it's always hard to tell like how much that is, is just yes, you could now uh, do lots of things on the hardware you couldn't do before. How much of that was just timing correlation but maybe just speak to, to that and I think it's kind of relevant to this conversation around like are these architectures just going to get better as the hardwares get better.
Lukas Kaiser: I mean hardware gets better and hardware, it's easy, it's flops and memory access. Right. So you need memory fast enough to feed the flops. Um, but it's a very simple can call it performance. And uh, I recently so I got a personal computer, I got one for myself and they bought a 5090 GPU and it felt like, oh, you know, it's one GPU and like you're under your desk. What can this do? Um, so I get a little bit of some tests and, and it's just insane to think. So the 5090, it's about 200 teraflops. I mean it says 400 but some are turned off on BF16. So the, the GPUs we research transformer on, they had nine teraflops and we had eight GPU machines. So in absolute scaling you could say be like um, 70, 80 teraflops for real on the machine. Um, so now I have under my desk something that's like five of these machines in one gpu, which is much more convenient than, but I think we used like around 10 or so. So you could do all of transformer research on this few thousand dollars GPU under your desk that you know, you could have in your kitchen. Like it's, it's just, it's a normal little tower and oh, okay, it's a few years. It's not even a decade though. So. So it's quite amazing what, what they can do. And, and now we run everything in BF16, but of course you can go lower even in precision, especially with MOEs. Then you pack more um, in inference. This is amazing. So our ability to run these models has dramatically increased and it increases the things you can research. You can now run so many interesting ways now it does give you the ability to just. Oh, and Also there's more GPUs right in the world. Like the big labs are building out so you can train huge models on a huge number of very fast GPUs. And Nvidia has kept the pace. And then TPUs at Google have kept the pace. They're really speeding up very quickly. Um, and their numbers are growing and then it's a very parallelizable process. So we can now train much bigger models much faster. That is amazing. I do still think that the even more interesting things is that we can do more research. And. I remember when I was joining Google, people were talking about how much flops do you need to do something like the brain. And it's a very vague question because to really simulate every neuron in the brain, it's maybe impossible, maybe still very much. But people for decades have been doing these estimates and they always felt like somewhere between 1 and 100 petaflops. And I remember back then we were like, okay, so this is going to take a few decades for us to get there. Now you can buy a single gpu, so that is quite insane. You have this one thing and then of course you can on the cloud get machines with, with many of them. So potentially you can run like a year of worth of human processing in a day, uh, at a cost. Right? But it's not a cost of millions, right? It's a cost of hundreds to thousands of dollars if you believe you can maybe figure out this algorithm. Like, I mean it's questionable whether we have the data that people have. Some people are trying to do recordings of kids, right? There's a question how, uh, there's a lot of questions, but we're getting to this level where someone at a university will be basically able to run like a childhood. If you have an idea for how the brain learns, you'll be able to run it in a few days. The whole 10 years of learning of a human being and see if it works or doesn't maybe if you know how to evaluate it. I think this is even more powerful than the fact that we can build these huge models, which is also powerful because they will help you implement this all. And we're getting this loop where I always felt limited with RNNs for example, because they're very sequential. So if you just run them like in Torch, they're very, very slow. Right. But you can write a special CUDA kernel that makes them very fast. But writing CUDA kernels is awful, right? You uh, really don't want to do this except when you can have a unit test that does exactly the same thing as your slow thing and an agent that writes them for you. And they're not yet amazing at it, but they're already do it and you know, bigger model will probably be so good that you'll be, you'll just say, you know, use this hardware as best as it can be and come a few hours later and here it is. So the bottlenecks that were like because the hardware did not fit your idea well, the hardware is still the way it is, right? It can't do anything you want. It still needs to be parallel, but it can do much more than it could do before. Because you can just ask pages to write kernels for it.
Host: Yeah, it's so interesting because some people will say, God, without the scale of compute that exists at only a few places, it's so hard to do. Maybe you can do basic research, but ultimately the rubber hits the road on seeing whether these techniques scale. Right. And you need to be in a lab to kind of experience that. But it's awesome to hear your kind of bullishness around, uh, the opportunity for academia and hobbyists and folks that are just messing around with, with single GPUs to be able to uh, to, to contribute here.
Lukas Kaiser: Well, I think especially if you believe that there are some radical changes that you should do.
Host: Do you think it's more likely than not that that's the case?
Lukas Kaiser: Like I guess it's uh, it depends on the day on my, on case. I do, uh, you know, research has always brought us beautiful things. There is no reason to think it won't. Um, but then the techniques we have also seem to work so well that it just feels mind blowing to. It will be a big mistake to not push on those two. But luckily there's, you know, there's enough labs. I feel the whole thrill of being an academic, I was an academia before I joined the labs, is that you can go wild with your ideas, right? You, you can't go, you can't scale up that much. But, but on the lower scale, which now is not that low, you can go really wild. You, you, you can try, you know, beautiful ideas that, that are totally out of the current paradigm and you should, that, that, that's, you know, that's the fun of being a researcher. And then, well, you know, not many won't work. Some will work in the small scale and not scale up. But I mean at the scale the current eight GPUs machine are, I mean, sure, there will always be ideas that work up to a certain scale and don't work further, but I think you're at much higher level now than it was like five years ago. Because five years goes was really amnesty tiny things. There was a lot of tweaks that were just really small scale tweaks. Now you're getting even on one machine. You're getting to a scale where it's not tweaks anymore. It's like um, I privately use Nanochat from Andre. Yeah, it's a GPT2 level model that you get in a few hours on one box. Right. Um, these boxes have gotten a bit more expensive these days, unluckily, but, but you know, a new generation of GPUs will come. The older will get cheaper. It's a. Yeah, it's just quite astounding what you can actually do. And yes, not all of this will scale, but the fun you can have on the way.
Host: Yeah, totally. I guess one more research frontier I'd, uh, love to get your take on before we shift gears is multimodal models. And I think you, I think on a previous podcast you said we haven't made a ton of progress there. Do you still feel that's the case? And what's your kind of current state of the union on, uh, the multimodal world?
Lukas Kaiser: So people are certainly making progress. Maybe this goes a little bit towards jepa, but the way we do multimodal and Transformers or even with diffusion models, it's like in the end you predict every pixel of these things around. And if you think of me being here in the environment, and I think humans sense an amazing amount of information every second or less than Microsoft. But we can't act. Our neurons are slow, they have hundreds of millisecond processing. But we get all these senses everywhere, all the time, and we somehow manage to learn from this insane stream M without maybe, you know, like predicting every pixel auto regressive. Like it's, it's both like way more parallel and like much larger. So I feel like the models we have, they have not truly done justice to this yet. Maybe it needs new research. Maybe. I mean it's. But they're also very similar. I think Thinking Machines has recently had this like multi stream Transformers and it feels so easy, right? I mean in a Transformer you pay attention to the previous tokens. You could have a bunch of streams that do this, right? That feels like an easy tweak to the architecture. But you know, maybe it's an easy tweak, but just an amazing tweak. Because I always, when I work with like codecs and you know, I just forget something, I say it. But then it's executing some bash command, so it needs to wait for my thing to steer it and it takes three minutes. And I'm like, this is just so not interact like it should just. And then you can have the side thing and there's a bunch of hacks again that kind of make it feel better, but it feels like of course everything happens everywhere all at once for us, right? We see, hear, talk all at the same time. Um, that should be how our models behave now that there is a bigger lab putting pressure on that maybe it will come, uh, but yeah, it does feel like we do multimodal without it, without all of these truly architectural changes to be parallel and absorb, uh, you know, transformer can't currently at the speed it does absorb, uh, a high resolution image every millisecond. Right. Because it splits them and they're so sequential in this that it just doesn't work. That feels somehow wrong. Right. It's like we shouldn't be putting these tiny patches there which should like just go in, be processed somehow. Um, so I don't think we have like on this deeper level gotten there yet. But on the other hand, I think feels like a lot of people are working on it then for coding. I mean, does it matter all that much harder to say Totally. Well, I'm sure it will come.
Host: I'd love to kind of switch gears and maybe just talk a little bit about your time at OpenAI and your kind of journey there. Because, uh, obviously it's been quite the eventful past years and maybe, uh, there's a few moments everyone kind of thinks about. And so I'm curious to get your perspective. But maybe just like on the OpenAI side, the company's had some very public moments. Um, and I'm wondering what were some of the difficult decisions that really defined the company? I guess in your time there?
Lukas Kaiser: I wasn't there for the earliest things. I think for my time there was this big question at, uh, some point whether to pivot to reasoning. And I feel it was very brave of the company and the leadership and all of us to actually take this plunge and say, yes, reasoning will be as important as pre training. Our models will be reasoning models. They will be launched. And at the beginning it was like the reasoning models were not that chatty somehow. Personality was harder, they were slow and they still are to some extent. And it was like, should you ever do it? Maybe people just prefer chat models, but OpenAI was very good at taking this hard bet and saying, yes, we're going to launch it, we're going to go this way. Well, try to figure out how to manage. There were two lines of models at the same time. That's obviously awful. You want to unify this. The unification took a lot of time because everything is moving. It's a very hard decision. But now we wouldn't have possibly all of these amazing things we have if it didn't push on it. And it feels like even some bigger labs still have trouble catching up to the RL quality. So there is some win that you get when you commit to things and you Know, I wonder these days, you know, OpenAI has since then grown probably like 20 times or something like this, become a much bigger company. All of the labs have, uh, I mean Google was big even before, but everyone anthropic has become big. Having been at Google before for a long time, I think it's much harder for a big company to take wild bets like that. Right. Because you have so much more to lose because you have processes like, it's just harder. I just hope OpenAI retains this ability and the other labs, uh, too, because the current techniques are amazing. They get us very far. But if there were sparks of post transformer world, would these labs be able to jump on it or would they be on the more conservative side?
Host: Now it feels like with reasoning there were some early sparks, but obviously not a ton of data. And then it was kind of almost ah, a. I've heard it articulated as like a religious belief that this is just going to work. If we double down on it and
Lukas Kaiser: we don't have the successor yet, or at least I don't know about it, but with the hope that it will appear, will you need a new lab to push on it or will. I mean, I think if anything, OpenAI is good at wild bets.
Host: It's obviously interesting to see this whole, this whole trend of neolabs. Right. And folks like Jerry Torick spinning out and saying that like it's almost, you know, easier to do this work outside of uh, a large lab. Right. And make one kind of strong convicted bet.
Lukas Kaiser: Yeah, it's a fair point tool. Right. Uh, but then, um, you know, you start looking at the GPU numbers and it's a little sad when you're outside of the lab. It's hard to get them and they're very expensive. So, so, but, but, but then GPUs are not everything. It's, but it's, it's quite nice to have this whole ecosystem.
Host: Right.
Lukas Kaiser: You have both these little labs now and the big labs. Yeah, we'll, we'll m. You know, I, it's, it's so fun because being in this, this AI little bubble here, you clearly see that there is a ton of competition. That change is coming. We have not exhausted even on the current paths. There's still a lot of techniques to do. There's a lot of data and improvements and bigger models to train in. And then there's all these new things that are bubbling. Maybe they're not ready, but they're very actively pursued with good resources. And then I feel like you step outside of San Francisco. And people treat AI basically as if it was like from the last year before Codex and would never change again. And then now that is a wrong way of treating it. It has all. I mean, to me, coding agents have been such a reveal that it's hard to, uh, hard to get over it. I call it AGI. You know, should call AGI what they want. We may get one day past AGI the way we got past the Turing Test, right? We don't really argue about the Turing Test anymore. Is it passed? Is it not passed? But who cares? Um, these things that they code with are clearly intelligent and coding and the. Any disputable.
Host: You know, obviously the AI coding wars are like, quite fierce right now. Like, you know, what do you think ultimately will determine, you know, which of these AI, uh, coding products ends up, uh, you know, how do they become better than each other? Uh, and, you know, uh, how do you see this, the next frontiers for like, codecs and Claude code?
Lukas Kaiser: You know, I think coding market is good to have two big enough to have two programmers in it. I think the bigger question will be how well do they go to other fields, right? I mean, coding is great and it's important for us, but you could do the work of many people. And currently Codex, I tried to recommend it to some friends, but, you know, it used to start with the question, what is your GitHub repo? Well, well, that cuts off a lot of people. Now it's a little bit friendlier, but it's still called Codex. You know, like, uh, even so people kind of don't hear, this is your accountant tool, right? And totally in contrast to ChatGPT, where you just said something, I think Codex takes a little bit of getting used to then clot even more. Uh, I would say if you go on the code side. So I think there is some question, how do you get this power to people in other occupations and places? And that may be the more important question.
Host: Anthropos tries with Claude Cowork and basically making a friendlier version, uh, of the core code products.
Lukas Kaiser: I certainly feel like the abilities are there. As an emojis person, I feel like this one obviously can do these things. They obviously can do Excel. They obviously can do this or that or that. But then of course they. Again, I watch them as a, like a hawk. How, uh, to. Like there is some level of skill that you need to put in to, To. To get this. Now this is totally a learnable skill, but I understand that people are busy in their lives and don't necessarily want to learn this. So you need to smoothen it. And in some way, there are some fundamental things that I don't think will allow you to, like, just let it run, not watched.
Host: Yeah.
Lukas Kaiser: Uh, you don't think you'll want to do this, but on the other hand, I don't think you would even want to do this even if it was super good at first. Right. You need to gain some trust. And so the question becomes, how do you convince people to start putting some effort into gaining this trust? It will pay back. But there is a hump, um, on
Host: the coding side, like, why do you think Anthropic was the first to be really successful on the coding side?
Lukas Kaiser: I think Anthropic made this very good decision to focus on coding. This was at the time when OpenAI was like, we're doing ChatGPT and great. I mean, chat is great. And I think part way Anthropic made this decision was that they just could not compete and. And Chat, but they made a very good decision on what else to do. And this goes back to, you know, AI goes through these upheavals. Right. You need to put a bet on something that is not what is today, even though the things today. Like, isn't ChatGPT amazing? Of course. Right. It was the most amazing AI of 2025, but clearly not of 2026. And maybe in 2027 we'll have another thing. Uh, so things change quickly if you put a good bet on something else you can. And it's not like OpenAI didn't do coding. We did. Right. And that's why it could catch up reasonably quickly, but it was just not the focus. I mean, these companies are tiny. You grow to a billion users, you have stuff to do just fall apart.
Host: You mentioned, uh, uh, this kind of, like, almost tension between, you know, nailing the stuff that's working today and then, you know, keeping other areas open so that if there's a glimmer of hope in a different area, you kind of double down on that bet. And I'm wondering what you make of that. You know, obviously, OpenAI, I think, very publicly, has gone in this, like, focusing moment now. Right. And you've seen it in the results of Codex and, um, you know, maybe slashing Sora and some of the other things that were. That were different. Yeah. How do you think about, like, kind of navigating that tension of. Of, like, really nailing some the here and now versus like, keeping these other embers open? Uh, that could potentially be really interesting down the line.
Lukas Kaiser: It's a matter of culture and size and money and perspective. Famously. Google, right. Google is the lab that will keep all its.
Host: I think some people have been quite critical of Google for this. Right. Missing, you know, missing your, your invention, uh, uh, not being the ones to capitalize on it, you know, but then
Lukas Kaiser: it, it works for them. Right? It works because it, it whatever good comes out, it's very easy to catch up because you already have a strong team in, in. In it. Right, yeah.
Host: Do you think they've caught up? I feel like there's a lot of discourse claiming that, you know, they're still a bit behind.
Lukas Kaiser: I think they've caught up in the ChatGPT world. They haven't yet caught up in the. I mean I don't know if you've seen Anti Gravity Tool. Yeah, I opened it after IO and it. I couldn't tell which one is codecs and which one is nt.
Host: Of course there are a lot of funny uh, tweets about that.
Lukas Kaiser: So that's great. Um, I tried to do some of my codexy things with the new 3.5 flash and it just doesn't work right. The barrier that wasn't Christmas. I feel like it hasn't crossed it yet to me, but it will. So if you're very broad, it's can make it safer later if you need to catch up. But then you may not get the immediate win of entropic and coding. You're just the first to nail it. And great that there are labs that just go and are the first to nail it. That's exciting and I feel like that's how it should be. Um, and OpenAI had a good culture of making bets, but now it is also a bigger thing and it has some. ChatGPT has a billion users. It's important for many people in the world. And a Google search has 3 billion users. It's important for many people in the world. You don't want these things to be hampered totally. Uh, you should go fast. But this breaking things is not so good and I actually feel like quite good if the labs don't break everything on the way.
Host: I guess a lot of people wonder about the kind of gap between closed source models and open source models. And it feels like there's two distinct things pulling in different directions. One is it feels relatively easy to distill models and you've seen uh, uh, a lot of claims around folks doing that on the Chinese open source side with the closed source providers. Then on the other hand it feels like more and more of these models, even in the big labs, are getting too big to serve, so they have to be distilled, uh, within the big labs themselves. What's your kind of gut intuition on the gap we'll see between closed source and open source models and whether that widens or shrinks in the next few years?
Lukas Kaiser: Yeah, it is not that easy to predict, I feel. So bigger models are better. Um, you can distill them, but the distilled models are never quite. They're great, especially if you need a model for some money. Uh, but they're, uh, not quite as good as the big models. I, you know, I just said like the 3.5 flash, I could not quite feel it's on par with 5.5. Maybe because it's a distilled pro. Right. Maybe, maybe you just need to wait for the pro. So even within the, like I, for example, I don't remember when I have used the mini model, the, the minion animals. I think they're very good, they're very useful. I just haven't used them in a while. Right. So. And whenever I use them, they're fine until they trip and cost me so much time. I go back to the people. And so you can distill things and you know, when the open source can distill or not distill. I mean, you know, labs just try to not make you distill everything naturally, but I think they also, like, don't fight you to death. Yeah. It would be very sad if open source had like, models that are very, very far behind. But I don't think there is a risk of that. There's enough, uh, companies and now there's like, you know, notions of. I also very much understand, like, if you're a country, right. Do you want to depend, like, say your police stations or hospitals run AI to help you do administration? Maybe you don't want to rely on one company that may just have an outage or, you understand. So there'd be a lot of people who want sovereign, they say models, Right. Even if they're slightly weaker, maybe the tasks are not so hard. So I think there will be enough incentives to have open models that they will exist. And, um, there will be a very good incentives for the labs to still keep ahead. So, you know, people, people keep paying for the. So. So it feels like, uh, a state that should persist for a while. But you know, it's famous last words and AI and tech, you can say things and they may turn out. I don't want to make future Predictions,
Host: of course, but what job is a podcast if not to try and force you into them? But uh, no, that all makes a ton of sense. Um, we always like to enter interviews with kind of a quick fire round where we stuff in a bunch of broad questions at the end. And so maybe to start, I just love. What's one thing you've changed your mind on in the AI world in the last year?
Lukas Kaiser: Uh, well, definitely I did not believe that they will be like a intern kind of thing so fast. And I have definitely changed my mind. I actually used to not talk to AI very much every day. People were always like, so how do you use ChatGPT? And I was like, yeah, I don't know. I asked it one query yesterday and one 3 days ago and I was always like I'm not going to talk to my computer very much. And now I do about work. Uh, so yeah, like, yeah, I also did not think I'm going to like not use an editor for programming and now I don't. I just tell it to change the code and yeah, that, that was a big update.
Host: Uh, that's awesome. Um, I guess, you know, as you worked at these models more closely these past years, have your concerns around like the existential risk of safety around these models, have they gone up or down these past years?
Lukas Kaiser: I don't think they have changed very much for me. I was always on the, you know, not too worried but, but also we should not be complacent side. And I still feel, you know, with all the skills they have now with programming and so on, I still feel the small risks, right, the risk, the risk that they will hack some of our systems, make the grid go down or things like that. I still feel these are the risks I would focus on right now. Not to say that external risks don't. You know, it's good that there are people thinking about it. It's good to have some guardrails. It's in the end good to, you know, we should be able to turn off these data centers if we so decide and have control over all of that. But I don't feel, even though the models have become much better, I don't feel any threat from them.
Host: On the lab side. It feels like the buzzing news of the last week was that like, uh, you know, Andrej Karpathy was going to Anthropic, uh, to work on rsi, right. As a team there, like, what do you make of that?
Lukas Kaiser: You know, I am part of this psychosis, right? It's like you can do so much research with this assistant. Right. And it's amazing and you should. And you can make many parts of the systems better too, like much faster. So that's certainly true. But on the other hand, you think about these post Transformers things and the space of ideas is vast and unluckily most of them are wrong. That's why it's called research. And you need enormous luck and skill, but also luck to happen upon the right one. And we kind of feel like maybe it's somewhere there in the air, but it's research. It may be years away. And even with the best AGIs of the world, they're like human level, they're maybe researcher level, maybe they'll be like a 10x researcher. Uh, but for years there was a huge community of researchers trying to crack these things and they didn't. So it may be just very hard. And yeah, we understand very little about the human brain yet and we cannot connect it to our ML in any way. Great way yet. So on the one hand I think it's great, I think we'll see the current things getting better. But if you're thinking of a research breakthrough, it may just require something that even when you have this, if you're searching in a very efficient way, and even if you're searching some interesting ideas, that still doesn't mean you're going to find it. Right?
Host: Yes.
Lukas Kaiser: Just the space of all ideas is so vast that even very efficient searches can just not get there. So I'm not that, I'm not that worried existentially about this.
Host: One thing I found interesting is I think, you know, if I'm correct, like all of your, your Transformer paper co authors have, have gone on to start companies. Right. And I'm wondering if that was ever something you thought about or uh, certainly
Lukas Kaiser: ask about it many, many, many times. Um, yeah, I, um, Well, I, I'm, I'm very happy that I didn't so far. I thought both my time at Google and my time at OpenAI have been great and, and, and, and it was a privilege to, to, to be there and, and to, to, to be able to do the work. I, I love technical work. You know, um, everyone who thought maybe they won't need to spend so much time on the company work and it feels like they had to, you know, but, but, but then sometimes companies do amazing things.
Host: That's been a fascinating conversation. I want to make sure to leave the last word to you. Anything you want to point our listeners to or thoughts you want to leave them with. Uh, the mic is yours.
Lukas Kaiser: I. Thank you. I just want to. I think I said it already, but I just want to repeat. I feel this time now that you have powerful GPUs that you can put under your desk and coding agents that can really help you push them to their limits. And the time where all the big things are pushing the transformers and great they are because they're amazing. But there is this whiff of possibly other things. I think it is, um, still and again, the most exciting time to be a researcher in machine learning. And I want to encourage everyone to just go and try their ideas to learn from others. Uh, if anything, I feel like we should publish more of Wild Things. I feel always a little sad when so many papers are about like, oh, we took a pre trained model and rl'd it in a slightly different way. It's good, but, you know, you don't need to catch up with what is there. You can just do new things, even if, even if they'll start smaller, even if, you know, maybe it won't work the first time. Um, you know, nobody talks to me about the paper I had before. Attention is all you need, which is you don't need attention. I had the paper at Neurips the year before. Active Memory wasn't quite good advice, but you need to explore the wrong things because they may lead you to the right thing. And, um, this is also what models are still so bad at, which I think Jerry is trying to push. Models are very bad at learning from a totally wrong direction to actually twist it to a right one. That's what we humans can still do very well. So we should do more of it. We should just do wild explorations even if they fail. And I feel now that, uh, if you put a lot of your own effort without an agent, it's then very hard to when it fails. I think with agents it's even easier. So I want to encourage everyone to research explorations fail when it comes to this is how we can get to interesting things.
Host: I love that. Well, I feel like that's the perfect note to end on. Um, thank you so much for, uh, coming on the pod. This was super fun.
Lukas Kaiser: Thank you so much for having me.
Host: I'm Jacob Efron and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a, uh, nights and weekends project in addition to my day job as an investor at redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so please consider doing that. And thank you so much for your support and listening. We'll see you next episode.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.