The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The MAD Podcast with Matt Turck
The MAD Podcast with Matt Turck artwork

Why NVIDIA Is Giving Away AI Models | Bryan Catanzaro

The MAD Podcast with Matt Turck · 2026-07-02 · 1h 23m

0:00--:--

Key moments - from our scoring

Substance score

61 / 100

Five dimensions, 20 points each

Insight Density12 / 20
Originality10 / 20
Guest Caliber16 / 20
Specificity & Evidence12 / 20
Conversational Craft11 / 20

Bryan Catanzaro, who leads Nemotron at NVIDIA, explains why the GPU leader is investing heavily in building its own frontier AI models and open-weight model families. NVIDIA's Nemotron 3 Ultra, released recently, demonstrates the company's dual strategy: first, to deeply understand AI systems architecture so NVIDIA can co-design specialized hardware and software for acceleration (now that Moore's Law is dead and progress depends on optimization rather than raw transistor scaling); second, to support the broader AI ecosystem by ensuring open technologies remain available to companies of all sizes - a business imperative since more deployed AI drives more demand for NVIDIA's infrastructure. Catanzaro draws parallels to the internet, arguing that AI, like the web before it, will be most transformational when developed openly across diverse industries and geographies. He addresses the distillation debate directly, countering concerns that restricted APIs from labs like Anthropic will slow open-source progress, and credits the Chinese AI community's openness for spurring global innovation. For infrastructure operators, enterprise AI leaders, and those evaluating open versus closed model strategies, this conversation clarifies NVIDIA's strategic positioning and the business case for open-source adoption.

Key takeaways

  • →Open-source AI progress is driven by demand from companies needing customization and data privacy for their specific use cases, not just community collaboration.
  • →NVIDIA builds Nemotron for two reasons: to deeply understand AI systems for designing better hardware/software, and to support the broader ecosystem where NVIDIA wins whenever AI is deployed.
  • →Moore's Law is dead and semiconductor progress now comes through specialization and co-design of hardware-software systems, which requires companies to deeply understand AI workloads.
  • →Open technologies enable different applications of AI across industries in ways closed APIs cannot, similar to how the open Internet enabled different transformations in retail, healthcare, and manufacturing.
  • →Companies derive competitive advantage by integrating AI with their proprietary data and business secrets, requiring control that only open-source models provide.

Guests

Bryan Catanzaro

Topics in this episode

NvidiaMoore's LawTransformersNemotronDLSSMegatronBaidu Silicon Valley AI LabGPU accelerationCUDATensorFlow

Questions this episode answers

Why is NVIDIA building its own AI models when it primarily sells GPUs?

NVIDIA builds Nemotron models for two strategic reasons: first, to deeply understand AI systems so it can co-design hardware and software for specialized acceleration now that Moore's Law is dead; second, to support the ecosystem by ensuring open-source AI remains accessible, which drives demand for NVIDIA's infrastructure and benefits its business broadly.

What is the current gap between open-source and closed-source AI models?

Catanzaro avoids directly ranking the gap, emphasizing instead that the entire AI field is moving extremely fast - progress over the past three months has been incredible regardless of model openness. He views the field's velocity as more important than comparing specific model capabilities between open and closed systems.

Can open-source AI progress continue if companies like Anthropic restrict model distillation?

Catanzaro argues it will, because the technology community will continue investing in the most transformational technology of the era, and progress won't be controlled by a few labs - bright people globally have good ideas, and community-oriented approaches have historically proven most successful for building transformational technologies.

What's the main business advantage for companies using open-source AI models over closed APIs?

Open-source models allow companies to customize AI around their core secrets, proprietary data, and specific customer needs while maintaining control over data privacy and regulatory compliance - something critical since AI value depends directly on data quality and contextual integration that only individual companies truly understand.

Is the progress of Chinese AI models primarily based on distillation from Western closed-source models?

Catanzaro, who worked at Baidu, rejects this characterization as false, noting that Chinese researchers are genuinely creative and inventive; while the tech community learns from each other globally, the Chinese AI community's openness has actually benefited the entire world by enabling companies elsewhere to build capabilities they couldn't achieve otherwise.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

12 / 20

The middle third of the episode contains genuinely useful technical content - hybrid SSM/Transformer ratios, latent MOE's 4x expert expansion, and the counter-intuitive quality-speed relationship in multi-token prediction. However, the first 20+ minutes are diplomacy-heavy open-source vs closed-source platitudes, and the organizational/safety sections trend toward abstraction. A solid but uneven density.

with multi token prediction the speed that you get is a function of the accuracy of your model, the more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works
using both of these together was actually better than using either one on their own. Um and that is independent of the speed benefit.

Originality

10 / 20

There are a handful of fresh framings - the kitchen/external stomach analogy for AI, a genuine singularity skepticism grounded in multifaceted intelligence, and the open-source-as-safer-because-of-sunlight argument - but the broader narrative follows the well-worn open source advocacy script and the organizational culture observations are standard startup-at-scale lore.

we, we have an external stomach, we call it a kitchen. Now we're creating an external brain.
the singularity is. Although it's an attractive idea, I think that it's really a wrong headed idea because it doesn't really um, take into account these other factors

Guest Caliber

16 / 20

Catanzaro is a genuine first-principles practitioner: co-created CuDNN, launched the Megatron project pre-transformer-hype in 2017, worked at Baidu Silicon Valley AI Lab alongside Andrew Ng and Dario Amodei, and now leads Nvidia's frontier model research. He speaks from lived engineering experience, not thought-leader positioning.

I published my first paper, Training Models on the GPU. And people asked me why I was there. People said, this is not a good paper for icml. We just do fancy math here.
when Andrew Ng asked me to go uh, build the Silicon Valley AI lab with him, uh at Baidu, I thought oh, this is a great opportunity

Specificity & Evidence

12 / 20

The episode provides useful concrete details - active vs. total parameter counts (3B/12B/55B), 4-bit NVFP4 pre-training, 1M token context, 10-15 teacher models, 23-of-24 DLSS pixels - but is conspicuously absent of benchmark comparisons, comparative latency numbers, or hard evidence for the 'best open weights model' claim. Numbers are architectural, not evaluative.

nano is a 30 billion um ah total 3 billion active parameter model supers 120 and 12 and ultra is 550 and 55
23 out of every 24 pixels is uh, being generated by our AI model. When you're using DLSS

Conversational Craft

11 / 20

Turck lands two genuinely probing moments - pushing on the distillation-from-closed-models concern and the China copycat narrative - and asks clarifying technical questions that unlock real content. But he consistently accepts marketing framings without challenge (e.g., 'immediately became the best open weights model') and the safety discussion is allowed to close on an uncontested, sweeping claim.

To ask maybe a slightly cynical question, there is at least a part of the community that's wondering whether open source as an ecosystem, not Nvidia, but uh, in general has been progressing in part based on the ability to distill closed source models
Is what you just described, uh, called latent moe or is that a different concept?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A85%
  • Speaker B15%

Most-used words

nvidia55model55open43models41nemotron38build28important28trying26world25ideas25different24nematron24intelligence23technology23data23token21

Episode notes

NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? Bryan Catanzaro is VP of Applied Deep Learning Research at NVIDIA and one of the people whose work quietly underpins modern AI: he helped create cuDNN (NVIDIA's first deep learning product), co-invented DLSS, and named and built Megatron, the framework behind how much of the industry trains large models. Today he leads Nemotron, NVIDIA's family of open models - and Nemotron 3 Ultra, released just weeks ago, is one of the strongest open-weights models to come out of the US. Matt Turck sits down with Bryan for a genuinely deep conversation: the real business logic behind a chip company building its own models, the state of open vs. closed AI, and whether the US is falling behind China in open models. Then they go inside Nemotron itself - four-bit (NVFP4) pretraining, hybrid Mamba-Transformer architecture, mixture-of-experts, multi-token prediction, and multi-teacher distillation - all explained in plain language.

Full transcript

1h 23m

Transcribed and scored by The B2B Podcast Index.

Speaker A: If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. We build tools, we build external organs that help us solve problems. You know, we, we have an external stomach, we call it a kitchen. Now we're creating an external brain. What is the implications of an external brain? Pretty profound. Nobody actually really knows.

Speaker B: Hi, I'm um, Matt Turk. Welcome back to the Matt podcast. Open source AI is having yet another moment with powerful new models arriving almost weekly. And my guest today is one of the very best people to unpack it all. Brian Catantaro leads Nemotron, Nvidia's family of open foundation models. Now not everyone realizes Nvidia has a massive effort to build frontier AI models, but it employs hundreds of AI researchers. And Nimotron 3 Ultra immediately became one US open weights model when it was released just a couple of weeks ago. We begin this conversation with the state of open source AI and the race between the US and China. And then we go deep inside Nemotron 4bit training hybrid member transform M architecture, mixture of experts, multi token prediction and multi teacher distillation, all in plain language. And finally we get a rare look at how a modern AI research organization actually runs. How you get many brilliant minds to build one model instead of 100 papers. Please enjoy this awesome conversation with Brian Catantaro. All right Brian, excited to do this. It seems that open source is having a banner year. So you guys at ah, Nvidia just released Nemotron 3 Ultra, which is uh, an important moment and the best open source open weight model in the us. That was just a few days ago and then uh, Even more recently GLM 5.2 came out and that was another moment. So it seems that things are accelerating in open source AI. It feels like a great uh, place to start. What's your assessment about where we are and how wide the gap between closed source and open source currently is?

Speaker A: Well it's really exciting to see all of the energy going into open technologies for AI because we know that um, open technologies make it possible for people to innovate. You know the Internet is such a great example of that. Um, we actually did have closed Internets. I don't know if you remember things like America Online and Prodigy back in the day, um, and they were great, um, and open Internet has also um, been amazing. Right? Like so many different companies have been able to figure out how to transform their work, um, thanks to uh, an open technology. The application of the Internet to retail is very different from the application of the Internet to healthcare or manufacturing. But all of them have been totally transformed, um, by the Internet. Um, A.I. i, uh, believe is uh, also a very transformational technology and also a technology that needs to be applied in very diverse ways. And because of that, I believe that open technologies for AI are really fundamental. Um, and it's very exciting to see continued, um, investment and development of open technologies, uh, for AI from so many different organizations around the world. Um, uh, and uh, I hope that that continues.

Speaker B: And what's your sense for how far behind open source is compared to closed sources has been the big trend of the last few years has been this sort of narrowing gap. Do you think that open source is almost there or the bar keeps getting raised by the closed source models?

Speaker A: Well, I feel like this question, um, uh, it's maybe a tempting question because it's fun to set up kind of competition, but I actually feel like the whole AI community is moving very fast. Um, and, and if you look for example at the progress in AI, whether it's closed or open just over the past three months, it's been incredible. Um, and so if you're in a field that's moving really, really fast, I think that's more important than any particular gaps that might exist between different models. Because the most important thing is how is AI developing as a field?

Speaker B: What do you think the drivers are to continued progress in open source? Uh, AI is that the communities and big companies like Nvidia being behind it is that the global competition with China, what propels uh, open source AI forward?

Speaker A: You know, I think there's a number of things that, that are pushing open technologies for AI forward. One is just the demand. You know, there's so many organizations that want to customize AI and want to integrate it deeply into their work in a way that really requires open technologies for AI. And so, so I think um, the demand is certainly there. I think also it's just um, the best way to develop technology. Um, and we've seen this, you know, for, for many decades, that technologies developed in the open move quicker because we can all learn from each other. And um, in an era where we're undergoing the most exciting thing to happen in technology in our lifetimes, with the development and the deployment of AI, um, what else do computer scientists want to work on other than making AI awesome? And if working together as a community is the best way to do that, then that's also a driver that pushes the community towards openly developing technology.

Speaker B: To ask maybe a slightly cynical question, there is at least a part of the community that's wondering whether open source as an ecosystem, not Nvidia, but uh, in general has been progressing in part based on the ability to distill closed source models. And in a world where we seeing the anthropics and fable fives of the world starting to discourage distillation, do you think there is uh, a chance that open source AI progress may slow down in that context or as a result?

Speaker A: You know, in my mind there's no question that when the um, technology community decides to make huge investments in the most transformational technology of our time, that there's going to be rapid progress, um, and also that that technology is not going to be controlled by a small group of people. Um, because that's just not the way that um, the industry works. You know we, we um, do our best work. Um, uh, we have the most impact with our work when we're able to uh, each think about it in our own way and apply it in our own way. So um, you know uh, I love uh, the uh, closed AI APIs, uh, whether from anthropic or other people. I think they're amazing, you know, really, really impressed with the work that those labs are doing. But they're not the only labs in the world. There's lots of labs around the world and lots of people have a good idea. Um, it's not the case that there's only a few labs that have the monopoly on all good ideas. That's just not true. That's not how humanity operates. There's a, there's a lot of bright people on this planet. And um, you know the community uh of course cares deeply about this technology. It's obviously so transformational, has such profound impacts on so many things, um, that, that of course uh, many people ah, want to be involved in that. And um, so I think over time we're going to see that um, community oriented approaches to developing and deploying AI are going to continue to strengthen and be widely adopted because that's really the history of how we built things as a, as a human uh, species.

Speaker B: Do you think that is globally true as well? So um, you know, in particular with respect to China, this perception that uh, yes, a lot of people have great ideas around the world. Uh, however a lot of progress uh, from Chinese models were directly inspired or perhaps uh, generated through Distillation from the closed source models. Is that just uh, kind of like press rage bait or uh, from the perspective of a um, leading ah, AI researcher, you're very impressed by the novel ideas that come out of China as well.

Speaker A: You know, um, uh, perhaps unusually, I uh, actually did work at a Chinese company for about two and a half years. I worked at Baidu. I um, worked in the Silicon Valley AI lab, uh, along with Andrew Ng and as well as Dario Amadei. And um, uh, we all worked uh, for a Chinese company and saw how smart, hardworking, creative, inventive our colleagues were uh, at the rest of Baidu. And you know that experience has, has stuck with me. I um, think it's absolutely false to say that um, you know, uh, the achievements of some other country are all being um, created by sort of, you know, copycat mentality. It's just not, it's just not true. Um, now, do we all learn from each other, uh, in the technology community? Of course, you know, of course, of course we learn from each other. But um, you know, uh, I would say uh, you know, it's been a really good thing for the world that the Chinese uh, AI community has been so um, open with what they've been building. I think it's enabled a tremendous number of companies to build things that uh, they couldn't have done without um, that community. And I think it's also spurred um, technological progress throughout the AI uh ecosystem. So you know, I'm really grateful for um, the contributions that um, our um, colleagues in China have made over the years. And um, you know, I, I would love to encourage a spirit of openness amongst AI labs around the world outside of China as well. You know, I was really excited when uh, OpenAI released the GPT OSS models, um, ah, a while back. And then of course Google's been doing great work with Gemma. Absolutely thrilling to see that. Um, and you know we're pushing Nemotron along here at Nvidia as well. Um, so I think there's, there's a chance for um, uh, the rest of the world to catch up to China, uh, in the sense that um, we can understand the benefits of working together uh, as a community to build technologies for AI in uh, a way that I think China has frankly been leading.

Speaker B: Great. What is the case for a customer to be using open source models? Uh, these days? What is your fundamental advantage?

Speaker A: Every company is built around a secret. Uh, this is a secret that has to do with not just their intellectual property but also their platform, uh, which has to do with how do they interact with problems and customers, um, how do they think about solutions, uh, to what their customers need. And it is always the case that the value of AI is greater when it can be more tightly connected with those secrets. Because AI depends on data critically. So the more valuable the data that goes in, the more valuable the solution becomes. Now, um, every company, when it's thinking about how to deploy AI, has to think through what are the implications for the core secrets of our company. And, um, there's a lot of circumstances where, um, due to trade secrets or trying to think through the business model or even regulatory requirements, that there's data that you really have to treat very carefully by law. Um, and it is much better to do that when you are able to think that through and implement it yourself. Um, thinking about the integration of AI, the way that AI interacts with, um, customers, um, the guardrails that are put in place. You know, every company, um, has a specific understanding of its customers and therefore what, what the customer needs. And, um, the amazing thing about open technologies for AI is that they allow customers, uh, customization, right? So companies can think this through. They can build things that, that really matter for them. And, you know, I started out this conversation talking about the Internet and about how the Internet, the deployment of the Internet has been done in very different ways for very different industries. Um, and there's a lot of desire to do that. Um, uh, as we see AI change the way that we work and play throughout the entire economy, um, this is really spurring a lot of demand for open technologies for AI.

Speaker B: Great. I'd love to go into a bit of a deep dive into Nemetron. But before we do that, maybe a few minutes on your story, your background, what was your past to where you are today, um, including the Baidu detour.

Speaker A: So, um, I started work at, uh, Nvidia in 2008. At the time I was a graduate student trying to figure out, um, parallel computing for artificial intelligence. And I thought, um, Nvidia had a chance of changing the way computers work

Speaker B: AI, which was a lonely, presumably a lonely quest. Right. In 2008.

Speaker A: Oh, it was, it was very chaotic back then. People thought I was crazy. And, you know, I remember going to ICML in 2008. I published my first paper, Training Models on the GPU. And people asked me why I was there. People said, this is not a good paper for icml. We just do fancy math here. And I was like, well, but I think computing actually matters, matters a lot for AI. If we could train bigger models that had more capacity to learn. We could probably solve more problems. And um, they kind of nodded their heads in the doors like well we're not really sure why you're here.

Speaker B: Isn't a gpu, uh a thing for gaming as well presumably?

Speaker A: Right, yeah, there's also that. Right, which we continue to um, run into, uh, that idea. Actually a GPU is whatever Nvidia says it is. We make them. So a GPU is a thing that we make in order to accelerate the world's most important computations, um, which in 1995 was graphics. And for a long time now it's been AI. So um, anyway, I started at Nvidia. I um, was in the research group, um, doing uh, strange things about trying to make um, ah, compilers, libraries uh for AI on the gpu. Um, that led to um, the creation of uh first Copperhead, which was a Python, uh, embedded language that um, compiled to the gpu, um, which I think foreshadowed a lot of things in TensorFlow and PyTorch. Um, and then um, that led to the creation of cudnn, which was Nvidia's first product for um, deep learning on the gpu. Um, and uh, I really enjoyed working on that. Um, but I was always wanting to see more uh, firsthand about the applications of AI. And at Nvidia I was mostly working on uh, libraries and compilers for AI. So I thought, um. Well you know, when Andrew Ng asked me to go uh, build the Silicon Valley AI lab with him, uh at Baidu, I thought oh, this is a great opportunity because even back then Baidu was um, very advanced in its application of AI to its core business. And so um, uh that was a fantastic opportunity for me. The Baidu Silicon Valley AI lab was an amazing place, um, full of brilliant people that were working really hard.

Speaker B: What was it like working with a ah, young Dario? Was there any signs that he could become uh, who he has become?

Speaker A: Dario um, was brilliant from the beginning. I remember, um, I interviewed him, uh, I was on the panel. Um, and um, at the time he uh, had been working in bioinformatics. So he hadn't been working on deep learning or the things that we call AI these days. Um, but it was very clear that he learned extremely quickly and also that he thought extremely deeply. Um, I um, think you know, uh, the thing I admire most about Dario, uh is the strength of his conviction. Um, um. You know, I've been working in this field uh for a long time and I've believed also that AI is going to transform the world. But I don't think that I believed, uh, in it as completely as Dario did. And perhaps that was because, you know, my academic training during my PhD was full of a lot of caution. Uh, I don't know if you remember, but AI was old and bad in 2005. It will never work that people did with computers. They started doing it in 1945. Right. Um, and so there had been so many grandiose promises that failed to deliver over the years. And so I came to AI with a lot of caution. In fact, back then we used to call it machine learning, which was basically a dodge. Like, we just didn't want people to know that we were working on AI because then they would be like, oh, we've heard about that ever works. Right? So m. I came to AI with a little bit of this, like, uh, you know, academic caution, like, oh, we should, you know, we should hedge a little bit. Like, I don't know if now's the time, like, um. And Dario, you know, his, uh, strength of conviction and his understanding of the moment of how, um, the technology was developing. This time it was actually going to work. And then the implications of that on, um, you know, how the technology should be developed, um, uh, what kind of institutions, uh, to build. I think he's done a spectacular job. And so, um, yeah, working with him, uh, it was always, always a fun experience.

Speaker B: So then you went back to Nvidia and walk us through the journey.

Speaker A: Yeah. So, uh, 10 years ago, actually, in 2016, Jensen, ah, called me up and said, hey, uh, would you like to come back and build an applied research lab? And I thought that would be a fantastic, uh, opportunity. You know, I've always loved Nvidia. I've loved, um, the way the company works, the convictions the company holds. You know, Nvidia, um, is a very unique company. It follows through over long time periods, you know, uh, and I've seen that with cuda. I've seen it with our deep learning technologies. I've seen it with our ray tracing graphics technologies, our AI for graphics, you know, over and over again. Nvidia is not afraid to put in five or 10 years worth of research in order to change the world, you know. And working, um, at a company that has that strength of conviction and the ability to follow through is kind of an ideal thing for me. I just really, I just really love the support that the company gives, uh, gives its researchers, um, to invent the future. And so, um, I thought I'd come back, um, uh, the first project that uh, that I worked on, um, actually became dlss, uh, which uh, some of your um, audience may know about. But DLSS is our real time AI for graphics. And it makes a small GPU run like a big GPU. It's about 10 times more efficient because rather than computing uh, the color of every pixel for every frame, we use AI to infer the color. Uh, and uh, you know, these days, 23 out of every 24 pixels is uh, being generated by our AI model. When you're using DLSS, um, to play games, and gamers love it, it's become the standard way of playing games because it's just so much more responsive and it's more beautiful. Our AI, we train it offline on huge data sets, um, and it's able to render uh, graphics in real time, uh, more beautifully than, uh, traditional methods do. Um, uh, we recently actually announced DLSS 5, which is a fully generative version of DLSS. And um, I am so excited about it. It represents culmination of 10 years worth of research on how to make real, um, time graphics much more beautiful. Um, and so, uh, that's part uh, of the journey, uh, here for me was real time, uh, AI for graphics. Um, but then at the same time we also started a language modeling project. Um, and this was back in 2017, um, before Transformers, uh, were big and before language, uh, modeling, uh, started uh, taking over the world. But, um, I just had this intuition, maybe built on, you know, some of the things that I had seen, uh, while working, uh, at Baidu. I just had this intuition that, you know, working with text and understanding text was going to lead to better reasoning, which was going to lead to better application of AI in all sorts of domains. And so, um, uh, we started this project, uh, called Megatron. Megatron stands for the biggest, baddest Transformer. That's why we named it that. And it was really a systems project to show the world how to train the largest Transformer models on Nvidia's hardware. Uh, back at the time, uh, some of your audience, uh, may or may not remember this, but, uh, there were being claims made that the only way to train big Transformer models was on the tpu. Uh, because after all the Transformer had been invented at Google. And uh, so, you know, we looked, we looked at, you know, we loved the Transformer paper. We thought, wow, this has amazing potential. We tried it out on our own language modeling tasks and it worked so much better than the RNNs that we had been using before. And also we saw immediately that there was an enormous systems opportunity to co optimize the gpu, the networking, all of the compilers and software that would enable people to scale Transformer uh, based language models uh, really dramatically. And we thought this is something that um, could really have an impact. So we started the Megatron project, um, which then led to I think uh, basically helping the whole industry figure out how to train extremely um, large uh, LLMs, um, and also led to the foundations of today's Nemotron project, um, where Nvidia trains uh, its own LLMs, um, uh, for its own purposes. So that's kind of the history.

Speaker B: Great journey. Okay, so let's go into all things uh, numotron and before we get into the specifics, that's the obvious question that I'm sure you've been asked uh, many times, uh, which is uh, why does Nvidia care in the first place to be building model and investing very significant efforts into creating its own family of uh, frontier models?

Speaker A: You know, nemotron has two jobs. The first job is to help us understand how to build the systems of the future. Nvidia is an accelerated computing company and that means thinking through the world's most important computational challenges from first principles and designing systems, which includes a lot of software, uh, in order to make it possible for people to invent and deploy things that never could have been done with standard computing. But in order to do that Nvidia has to deeply understand everything about how AI works. That's how we co design all of the systems and software um, uh, for our main product line. So the first job of Nemotron is to make sure that Nvidia continues to exist so that we can continue delivering meaningful acceleration in an era where Moore's Law has died. Uh, and the acceleration that we get these days comes through specialization. But again specialization comes through understanding. So that's Nematron's first job is to help Nvidia understand how to build its core products. Nvidia's second or Nematron second job is to support the ecosystem. Uh, one of the most valuable things that Nvidia has built over the years is um, all of the people around the world who build and deploy amazing AI, um using Nvidia's uh, technologies. And uh, we think that it's necessary for uh, open technology for AI to continue to exist, uh, from Nvidia to help uh, support that. Nemotron's not trying to be the only open technology for AI. We love all technology for AI for the very straightforward reason that whenever AI is further developed and further deployed, it's an opportunity for our business. So this is, um, you know, we're very explicitly trying to develop our ecosystem because that's good business for us. Um, but we're not trying to be the only provider of technologies for this ecosystem. We love seeing other um, companies contribute as well. Um, the most important thing for Nematron's second job is just making sure that it continues to be possible for companies of all shapes and sizes to build and deploy their own AI.

Speaker B: By the way, Moore's Law is dead. Is that official?

Speaker A: It's been dead for years.

Speaker B: It's been dead for years. Why is that?

Speaker A: Well, you just look at the progress uh, in semiconductor manufacturing. You know, the original statement of Moore's Law was economic, right? It was about we can afford to put twice as many transistors on the same chip in every, whatever, 24 months, whatever the time period is. And um, these days that is absolutely not the case. And it hasn't been for probably five or ten years. Right now we are still scaling our systems, right, um, through a number of ways. One is just applying a lot more silicon to it. Right. We uh, are also getting. Transistors are continuing to get smaller and more efficient, although at a slower pace, but they're also getting quite a bit more expensive at the same time. Um, uh, so the uh, you know, in an era where, where Moore's Law was alive, the best way to make the system of the future was to take the system of the present and then just shrink it and, and maybe double it at the same time. Right? But in an era where we've been living for a while now, where you don't get economic benefits from taking your existing design and shrinking it, uh, you really have to be more clever about how you use every part of the system. Uh, that, that's uh, an era where accelerated computing is much more valuable than ever because the work of thinking through the problem from first principles and co designing absolutely everything, uh, from transistors to algorithms and applications, uh, in order to reduce waste and deliver meaningful acceleration, that's more valuable than ever.

Speaker B: Fantastic. To play back, uh, what you were saying, uh, a minute earlier, it makes good business sense for Nvidia to be in the model business because one, it helps, uh, design better chips and two, uh, whatever is good for AI is ultimately good for Nvidia, which makes a lot of sense that Nemotron effort is reasonably recent. Right. It started in 2023, I believe. Maybe walk us quickly through the key releases. I believe in 2023 there was Nematron 3.8B as a key release. Or am I missing a step?

Speaker A: Yes, yes, yes. So you know, you know the, the numbering is somewhat lost to time. It. I almost feel like we're in the Lord of the Rings and it's like, you know there's like some ancient like relics that we're digging up out of an old mine. Um, you know this is a long time ago. You know the original what, what uh, we originally called Nemotron1 was actually a project that we did with Microsoft. Um, we jointly trained a 530 billion parameter model. Um, I believe that was released in 2021. And so this is GPT3 era. Um, and that's what um at the time we called it Megatron Turing nlg. Uh, Turing was uh, what Microsoft was calling their language model efforts um, at the time. But uh, uh, that in retrospect we called Nemotron 1. Then along the way we built a few more. We uh, got up to Nemotron 3. Um, and then uh, LLAMA came along. Uh, and we were really excited about that. We were very um, happy that META was supporting the open uh, AI technology space. Um, and so we started um, taking our language model technology and adding it to LLAMA models which then resulted in llama nemotron 1. Um, and that was the first reasoning, uh, model built uh, on Llama. Uh, we were really proud of that.

Speaker B: That was 2025, uh, might have been

Speaker A: 24, uh, I believe. I can't remember, somewhere around there. Um, and then uh, uh yes, and then we, we continued uh, to, to develop that. Um, uh, and uh, you know, last we. So we, the numbers kind of started over again. We released a Nemotron 2, um, I believe it was last year. Um, and then we quickly followed that up with Nematron 3 because um, uh, we, we needed to put MOE support in. Nematron 2 didn't have Moe support and that made it kind of uncompetitive against other models like GPT. OSS20B was just like so fast because of MOE. And so we were like okay, we've got to, we've got to put, put the MOE in. So that became Nematron 3. Um, now we're in a slightly difficult state because you know we're working on Nematron 4, right, but we already released a Nematron 4 which was um. In 2024 we released a 340B model, uh called Nematron 4. Um, and so I'm not exactly sure how we're going to um, solve this marketing problem. I didn't create this marketing problem. Uh, so uh, I'll, I'll do my best to, to make it clear that Nematron 4 of, of whatever the next, whenever we release that is different from the, the 2024 NE. Um, but uh, in any case we've been working on this for a long time. I think um, more important to us than any particular generation is just the sustained commitment uh, that Nvidia has to developing these models. We've been doing it for a while. I think our models have gotten dramatically more useful in the past year, um, which is a reflection of two things primarily. One is that uh, the whole company has come together. So there are many different teams around Nvidia that now understand how important this is to Nvidia's future. And so there's dramatically more people and better ideas that are going into Nematron. Um, and then uh, number two, along with that we've been able to scale the um, compute resources that go into it. Um, obviously it's very um, important to have good computing infrastructure to build AI. We've recently um, increased our investment substantially because uh, we believe that this is really, really key to our company. Company's future.

Speaker B: Fascinating.

Speaker A: But just to continue the thought, I think it's really important that everybody knows that we've been doing this for a long time. We are increasing our investments substantially and Nvidia is a company that follows through. You know, we followed through over 10 plus years with CUDA and we're doing that with Nemotron now.

Speaker B: That's very helpful because uh, I think the broader world is just starting to catch up to the fact that there is a very substantial open source frontier AI research effort that's been happening. Um, so it's really interesting to hear that uh, there's been this progression and now this family of models that we're going to talk about uh, in a second. Another important moment seems to be uh, the creation just in March, uh, three months ago of the Nemotron coalition. Do you want to explain briefly what that is?

Speaker A: So Nemotron exists to help support the ecosystem. And we were thinking well this is a different kind of AI project than other projects around the industry. Right? Because um, we're not actually trying to dominate in any way. We're just trying to support, we don't, we're not trying to control uh, uh, the way that AI is um, uh, being integrated into all these companies. We're just trying to make sure there's good AI. But we Thought well maybe if we worked with people while we develop it then it's going to be more useful for them, it'll be easier to integrate and because we will consider what they need from the beginning. And um, you know Nematron has always been collaborative. I was telling you that you know long, long time ago, our first big model that we trained we, we did with Microsoft, right? It was a joint effort where Nvidia and Microsoft researchers worked side by side to build that and that um, that ended up I think helping both Nvidia and Microsoft. I think we both learned a lot from that um, experience. And so because Nemotron is not trying to compete with uh, other companies but rather support uh, because we're going to be putting it out there openly anyway, why not collaborate before the thing is built rather than Nemotron being a project that Nvidia does all on its own and then posts on the Internet and says hey, why don't you try this? We think it might be good. Why don't we make sure that it's good for the partners that um, uh are interested in by working with them before Nematron is even created um and incorporating any sort of feedback, evaluations, environments, um, uh, benchmarks, um, or any other um, kinds of technology that other people want to bring. It turns out that the entire ecosystem, there's a lot of companies that really want open models to succeed and so they have a self interest, they have their own vested self interest and making sure that open technologies are excellent and so why not work with them and let them contribute uh, however they'd like to making Nemotron better. So that's the idea of the Nematron coalition. It is not an exclusive coalition. We're not trying to be the only model out there. All the companies that we work with are free to continue doing the work however makes sense to them. And yet um, these companies want to work with us because they want to make sure that um, open technologies for AI keep um, developing quickly and that they have a chance to influence how that happens.

Speaker B: Great. What's the current state of the new Motron family? You got Nano, you got super, you got Ultra. What do those models do and what are the use cases for them?

Speaker A: So nano is a 30 billion um ah total 3 billion active parameter model supers 120 and 12 and ultra is 550 and 55. Um they're designed um, really to fit, it's kind of small, medium and large. Um uh deployment scenarios. Um uh Nano can be really capable for things that um don't require nearly as much knowledge or reasoning. But obviously for the most capable um model you go for Ultra, um Super in a lot of ways is our most popular model because it represents kind of a great balance between um, cost and intelligence. So we kind of like um, having this small, medium and large um approach to building a family just because our customers seem um, to respond to that um pretty well. But um, you know uh, the most important thing from Nvidia's point of view that people are doing with uh, uh LLMs is agents, right? Is um building agentic workflows. Having an agent working on your behalf solving problems for you night and day um is such an exciting way of approaching the problems that we have to solve. Um and um, it's our dream to make Nemotron amazing for that purpose.

Speaker B: That's our goal to double click on this uh, at a high level Nemotron family is focused on agentic reasoning with a particular focus on making it efficient. Is that the right headline?

Speaker A: That's right, yeah. Um, Nemotron has always been um, uh speed first approach to building models because Nvidia is an accelerated computing company. As I was saying we're trying to think through what is the problem here computationally from first principles. And um, you know Nemotron 3 family has a lot of things in it that are ah, we're really proud of. For example Nemotron, um, Ultra and super, uh were pre trained using four bit arithmetic. We pre trained those in MVFP4 um, which uh, is a not trivial thing to do to invent the algorithm so that your model can converge to an excellent result using such coarse arithmetic uh required a lot of invention. Really proud of that.

Speaker B: Do you want to explain maybe for people what 4 bit is versus 16 bit for example?

Speaker A: You know actually there was a fantastic post I saw on Hacker News yesterday where somebody let you um, upload a picture and then it would basically posterize it. Basically reduce the colors to fit different um number formats including NVFP4 and MXFP8 and some of the other formats that are out there. And so you could kind of swipe around and look what it does to the colors of the picture. Um and you know it's really quite dramatic. Four bits is not a lot of bits, right? That's only 16 values. Um now of course these are all um, what are called uh block scaled formats. So um, groups of numbers also come with uh, an 8bit um scaling factor and the specifics of this can get rather complicated. So maybe they're not quite as Important. But the reason why we want to do this is because first of all we have dramatically higher um, throughput for these formats in our GPUs, specifically on Blackwell, uh, Ultra. Um, and uh, uh secondly we know that it's going to save an enormous amount of energy. Um, one way to think about uh, the computational problem of AI is that we are going to be running at the limit, whatever the limit is. It could be uh, an economic limit like we only have so many billion dollars to buy servers with. It could be a power limit. We only have so many gigawatts that we can afford to train a model with whatever the limit is. Um, uh, we're going to be running at that limit. Every organization is because why? Because the value of intelligence is so high, you know that, that people are going to, they're going to invest because they know that they're going to get return. Um, uh, the value of intelligence is enormous. Uh, so if you, if you accept as the truth that we're going to be running at the limit then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. And you know four bit number formats are dramatically cheaper to move around. They take up less space in memory, um, they take up less uh picojoules when you move them from the memory, um, uh, or even on the chip, uh around the chip. Much less uh energy when you compute on them. Ah so that's really um, driving the investment in four bit formats. And I think these days four bit formats for deployment uh, are very well established. Um, it's pretty um, straightforward these days to make a good quantized 4 uh bit uh checkpoint that you can deploy and that gets you a lot of inference cost and speed advantages. Um but using 4 bit formats for pre training, uh, that's quite a bit uh more challenging because uh, you have this um, numerical ah, solver that's you know, optimizing the weights and um, you know it can be quite sensitive. So if you, if you don't uh, treat the numbers right, uh, your model can diverge and instead of actually getting a model done through pre training you end up with you know basically just uh, that run diverged which is you know always, always scary. So it took a lot of invention for us uh to be able to pre train nematron uh in four bit. We're really proud of that.

Speaker B: Okay, great. All Right. So as we uh, get into slightly uh, more technical things, the uh, architecture of Nemotron is uh, hybrid. Is that right? So it's a combination of transformer and Mamba state space, which is a slightly more exotic uh, form of architecture. Walk us through that. Yeah.

Speaker A: You know we published a paper in 2024 that showed that you actually get a smarter model by combining um, state space models with transformers. And we actually did a sweep of uh, uh, how much of the model should be full attention and how much of it should be um, a state space model in order to get the lowest perplexity. Basically the best language model um, that you could get. And we found that you actually want it to be mostly a state space model with a little bit of attention. And kind of the intuition behind that is that um, the state space models seem to be better at ah, um kind of this um, intuitive kind of impressionistic understanding of a sequence. Um uh because um, they're kind of summarizing the entire sequence into a constant space. That's how they work. So instead of having the ability to look at the entire sequence randomly, uh, they summarize everything at every step into a constant um cache or little scratch pad that they're. They're working on. And um, that constraint seems to actually make them smarter at some tasks that involve like global understanding. Um, on the other hand the advantage of full attention is that it can pick out very specific uh bits of information and look at those. Exactly. It doesn't lose anything. There's no lossy compression going on. You can actually see the whole thing. Um, and so we found that um, uh you know, using both of these together was actually better than using either one on their own. Um and that is independent of the speed benefit. That is just the model is smarter. And you know, since we published that I think a lot of other um uh labs have also found this to be true. Um, you know, a lot of uh, models uh these days are being built with hybrid SSM approaches. For example, QN has done that, um, uh, Kimi uh is using what they call kimilinear attention these days. So um, it's become I think quite widely adopted to use some sort of state space model in conjunction with um, full attention for um uh for the, the base architecture. Um now it also has some speed benefits because um, the uh, uh amount of memory that you need to hold uh, that state space cache, uh is actually constant with respect to your sequence length. Um, which then means that generally um, you can fit much higher batches on the GPU when you're training and doing inference, uh, because the memory requirement, um, is lower and it keeps the GPU fuller, uh, and busier, and therefore, um provides some, uh, pretty important efficiency benefits as well.

Speaker B: So the models are also based on an MOE mixture of expert architecture. Walk us through that and maybe remind people what MOE is in the first place.

Speaker A: So mixture of experts is a form of sparsity. Um, the idea is, wow, you want to train a model on the entire Internet. You want it to remember absolutely everything about the history of everything. But when you're answering a particular question, does it seem reasonable that it needs to actually think about the entire universe in order to answer that question? Actually no. It seems like it's quite sparse. Right. It seems like, um, we're using a language model to explore a very tiny space of ideas in order to answer a question or solve a problem. We want the model to be able to draw from the entire universe. We want to train it so that it understands everything that it possibly can, but when it's actually running, it doesn't really need to see all of that information. There's been a variety of approaches to sparsity that try to take advantage of this property, but mixture of experts has been the most successful. And the way that it works is that the neural network has what's called a router that is learned that um, is going to decide to send activations to a subset of the experts. Uh, for every token that's flowing through every layer of the model, it's going to be making choices about, um, which fraction of the model is going to actually get to interact with this token as we try to understand it, build up representations of the problem, and then generate the next token that we're going to output.

Speaker B: So it's a little bit like if I have a company with 550 employees, but, uh, 55 of them are in engineering. I want the 55 employees who are specialists to come to my meeting about engineering and not the rest of the company.

Speaker A: That's right, yeah. Or you can think about it as a library. Like if you go into a library to do research, you don't read all of the books in the library. Like your first job is to figure out which books do you need to look at in order to find the answer to your question. And so, um, so that's kind of the, the idea behind moes. Now moes have fascinating implications for the systems that we build. So with Blackwell, for example, Nvidia went all in on MOEs. That's why we built NVL72 which allows up to 72 of our GPUs to read and write each other's memory at uh, very high speeds, very low latency. Now why is that important? It's because as you put a token through the stack of layers, at every layer you have a router that's routing that token somewhere else. Why don't you partition your experts so that the experts are not sitting every expert on every gpu, but you have a subset of the experts assigned to each GPU and then you're routing the tokens between the GPUs very dynamically as you push uh, the token through the network. Now um, this is impossible to predict in advance where the tokens need to go because it's very specific to that particular token for that particular model. And so that's why we built NVL 72 and that's why um, Blackwell is so amazing for inference for today's AI models is uh, because we thought deeply about mixture, um, of experts when we were building it. This is speaking to Nematron's first job. If we hadn't been working on understanding AI, we wouldn't have been able to build Blackwell properly. And that has translated directly into um, increased uh, deployment um, of Blackwell, which we're very excited about.

Speaker B: Is what you just described, uh, called latent moe or is that a different concept?

Speaker A: Latent MOE is a specific uh, innovation that we have in Nemotron 3 family. And uh, what it does is actually um, reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So you know, every token produces a vector. And the idea is like we're going to take that vector and learn a way to compress it and then send that compressed thing through the network and then we're going to uncompress it at the other end. And as a result we save on network bandwidth and we also get four times the number of experts for the same inference cost. So you could think about it as like, you know, our library of books got four times bigger, um, and we get to read four times more books, uh, at the same inference cost because of this particular innovation.

Speaker B: Is MOE in general, uh, becoming the default architecture for Frontier AI?

Speaker A: Yeah, I believe MOEs have been the default um, in Frontier AI for a long time. Um, they're just a really good combination of inference cost and intelligence.

Speaker B: Great, great.

Speaker A: But they have drawbacks as well. They take a lot more memory. Um, uh, if you have a very small amount of memory A dense model is going to be smarter. Um, and they also um, they tend to work best either if you're running at batch size one so you're running basically a single job or you're running a huge data center with like infinite queries coming in in the middle. They can be a little bit tricky.

Speaker B: Another important characteristic of Numertron 3 Ultra is a 1 million token context, the long context window. How important is that in the overall mix and what does it enable the model to do?

Speaker A: The longer the context length the more challenging problems we can solve. With a language model um, that allows us to do things like append all sorts of information to a query which could be a code base, it could be instructions. Um, uh, in the long term I'm hoping that I have my own personal LLM that's able to read all of my emails and help me answer questions about that. The more information that we can attach to a particular query um, the more useful the model can be. Um now uh, it can get more and more expensive to reason over large amounts of input data. Um and so that's one of the reasons why there's usually a limit on how big the context length can be. But with Nemotron 3 we tried to push it as far as we could go. Uh, we think a million tokens is a lot of tokens, um, and you can do a lot of things with that.

Speaker B: Presumably is particularly helpful in um, sort of multi step agentic uh workflows. And there's uh, this whole separate discussion around context compaction to make sure that the model doesn't get lost in too many tokens. So how do you all think about this 100%?

Speaker A: I mean um, uh, compaction, that's a thing if you're using an agentic workflow you deal with all the time. Um, and compaction tends to work pretty well because language models um, are pretty good at identifying the most relevant things and summarizing and you're basically trying to summarize your context when you compact it. Um, so compaction uh, uh, is not a bad approach. I think having models that can just natively reason uh about larger amounts of data is just inherently more useful. So of course we want to push the boundary on that as well.

Speaker B: Great. Can you talk about the multi token prediction which is also very interesting if

Speaker A: you're running at a low batch size um, which is when you are trying to get the most interactivity if you're in a data center. So you want the model to respond as quickly as possible and it's okay, for it to be more expensive, your cost per token might be higher, but you want the result as quickly as possible. Or if you're running um, locally, um, so you might be running at batch size one just because you're the only person using it. It turns out that uh, the GPU has extra execution capabilities that are just lying there unused. The bulk of the work when you're running in these scenarios is actually fetching the weights from memory and then you push the token past those weights and then you fetch more uh, weights from memory. But it turns out if you push two tokens or even five tokens through those same um, weights it would cost basically the same amount of time. Because the expensive thing is not doing the math to push the token through the weights. The expensive thing is just reading all of those weights from memory. All those parameters they have to come in. And so the idea with multi token prediction is to take advantage of this by having the model predict multiple tokens at once. Let's say that the model predicts five tokens. We know the first token is correct. The next four tokens may or may not be correct. So then what we do is on the next pass we take those four tokens and we stick them into the model, uh, and then uh, run it through and at the end we check, you know, the model then predicts another set of tokens, right? Then we check. Were the extra tokens we predict last time correct? If so, then we just accept them, uh, and then we get like a forex speed up, um, and if they were incorrect, then we only accept the ones that were correct and then um, uh, proceed uh, from there. So the benefit of this is it doesn't degrade accuracy at all because you're using the model to double check, right? So all this speculation is going to get checked uh, during the next token that you uh, run through the model. So it doesn't degrade your accuracy at all to turn on multi token prediction, but it can give you a speed up and it's probabilistic depending on the acceptance rate of your um, predictor, you know, so if your predictor is more accurate, the acceptance rate goes higher, you get a higher speed up. Um, so with you know, um, uh, with uh, our recent Nemotron models, you know, we're pretty proud of our acceptance rates but we're always trying to make them better, you know, always trying to improve that acceptance rate. This is a really good example of accelerated computing. You know, with multi token prediction the speed that you get is a function of the accuracy of your model, the more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works, but in this case that's how it works. And what that implies is that uh, you know, if we're trying as Nvidia as a company to provide meaningful acceleration to the world's most important computational workloads, this has to be an important part of how we think uh, about it. You know, if there's a 3x cost reduction or speed improvement, uh, for inference, which is the most important computational workload of 2026, um, if that's on the table and it depends on the accuracy of the multi token prediction network, then that's something that Nvidia needs to understand very deeply because it's going to affect uh, our business directly.

Speaker B: Fascinating. To continue on the tour, multi teacher distillation. We talked about distillation a little bit upfront. Uh, what does that mean in the Context of Nemotron uh 3?

Speaker A: So with Nemotron 3 Ultra, we did post training using something called multi domain on policy distillation. And what that uh, entails is that we have many different aspects of the model we want to improve. Um, for example science understanding is different from math theorem proving, which is different from coding, which is different from agent harness, uh, interactions. Right. Um, with Nematron 3 I think we had about um, 10 or 15 of these teachers. And um, so the idea is that you take these teacher models and you push them as far as you can go on some specific domain. So you just don't worry about making good at everything, just make it really, really smart at this one domain. Then you have a collection of these models and you want to create one model that learns uh, to be good at everything. And we do that using a specific reinforcement learning technique that a lot of labs these days use called um, uh, mopd. Um, and the good thing about this is that because the teachers are supervising, uh, they can give really dense uh, rewards to the student model. Basically every token is getting supervised and so the student can learn really quickly, um, and then become uh, almost as good as all of the teachers at all of the things. Um. So uh, one benefit of this is that it really helps the team work together better. Um, you know, if, if um, you don't have a technique like this and you have let's say 500 people working to try to make a model better and one team's like, well I'm trying to make it better at this thing and then Another team's like, I'm trying to make it better at that thing. There can be a tug of war where it's like well who wins? You know, and, and if you have to make a choice like oh, I'm gonna make them, I'm gonna choose to prioritize this one over that one, then you make the other team feel like their work doesn't matter. You know, it's just really hard. This is one of the challenges of uh, building AI in 2026 is that you have to figure out how to get the people to work together, even though you're only building one thing at the end of the day. And so this particular technology um, has been really instrumental in helping more people work together to make Nemotron stronger.

Speaker B: Fascinating. So it's just as much a, ah, technology question as a human organization question.

Speaker A: Exactly.

Speaker B: Okay, fantastic. Let's put a pin in this and get back to this in a second because it's a fascinating topic in terms of uh, the post training uh, that you uh, just alluded to. One of the uh, exciting things that you uh, all did uh, in the context of Nematron, is also to publish the data, the training data. Does that include um, per industry data for specific reinforcement learning tasks?

Speaker A: Yes.

Speaker B: That's the beauty of a conversation like this today, where you guys can actually talk about those things. So where does one get the data from or post uh, training, reinforcement learning focused um, efforts? Obviously one of the key questions in the world today is that LLMs or AI systems have become great at coding and great at math. The next big question is can they become great at uh, law and consulting and then uh, all sorts of different domains and part of the black box of closed models is how people go about doing all of this. Where did they get the data from? To the extent that you can talk uh, about all of this, I'd be very curious about how you guys have gone about it.

Speaker A: It's not an easy question to answer because it is quite complex. But I would say um, we rely on uh, a number of things. One is that we do purchase data from companies uh, that um, uh, are building data sets that you can purchase. Um, and to the extent that we have the rights to redistribute, uh, or to open up that data, we do as part of um, our um, Nematron data effort. Um, with Nematron we are trying to be maximally open with the data that we release because our goal is to support the ecosystem. Our goal is not to be the only model out there and we love it when we hear of other models around the industry that are using our data sets, um, to make their AI stronger because that means we're succeeding at our job to keep the ecosystem thriving and growing. Um, now uh, we also are big believers in synthetic data generation. Um, we use an enormous amount of compute, um, running uh, language models on our own systems to create synthetic data that then helps our models be better at uh, solving problems in specific domains and we release a lot of that data as well. Now it's of course not very straightforward to do this. Like AI is always garbage in, garbage out. So uh, you have to work really hard to make sure that any synthetic data that you create is actually adding value. That's actually helping the model generalize and solve problems, uh, more intelligently. But um, those are the primary ways that we go about um, building our datasets.

Speaker B: Since we're talking about post training and RL in different domains. Just curious to get your thoughts on where we go from here in terms of generalization. So just to build on what I was saying a second ago, the industry seems to be marching from coding and math, uh, which are domains with verifiable rewards to uh, different uh, industries. Do you think that this is where things are going and that AI industry as a whole is going to be able to cover those next few domains as efficiently as coding, uh, or math?

Speaker A: Coding is really special because um, it's a very intellectual exercise that created ah, a lot of economic value which then meant that we had an enormous amount of tokens that we could um, learn from as well as tooling, uh, that allows us to verify whether um, you know, our, our models are actually solving problems. Um, so coding, um, uh is always going to have a special place in our heart. Um, and something that I think AI is going to continue to get much better at, um because we have this special relationship with it, um, with regards to other domains. I think what I'm excited about has to do with significantly more diverse environments for AI to learn in during reinforcement learning, um, I believe that um, you know, reinforcement learning is such a general form of um, teaching an AI how to solve problems. Uh, we're just getting started at figuring out how to apply that. Um, and I think as our environments get more sophisticated, the AI then learns more understanding of um, uh, the problems that it's trying to solve as well as the implications of the actions that it can take. Um, then it becomes much better at uh, actually um, solving those problems. When I look at the environments uh, that we're using Today they're still fairly simple, um all things considered. And I think um, uh that's going to become significantly more complex and diverse over the next few years.

Speaker B: All right, so you mentioned uh, making 500 people work together and I said uh, that we would get back to it because it's so interesting. So just uh, taking a step back, like tell us about the research organization at uh Nvidia. Like how is it structured? How does it all work?

Speaker A: Well uh, Nvidia is not structured according to an org chart. Um we have one but it's not actually the best way of understanding how we work. Um, my uh team for example is not part of the official Nvidia research team. My team is actually part of the organization that builds the gpu. And my team is not the only team building nemotron. There's probably 10 teams around the company that have significant involvement in building Nemotron, um in different parts of the company and in enterprise software. Um, in uh, our AI software um division. The part of uh Nvidia that actually designs the GPU also significantly is involved in building Nematronics. Um so there's so many different teams that have to work together. Um we um always like to say that the mission is the boss um rather than uh the organization. But um, what that implies is uh that people have to figure out how to work together um which is challenging um in the sense that humans are naturally tribal creatures and it's uh not natural for us to um be friendly with people we don't know very well or trust uh coworkers that we don't have success working with in the past. And um, actually the name Nemo Tron reflects uh that we had um the Nemo team which was building software for AI and the Megatron team which was building uh primarily focused on systems research for um uh building large language models. And um, we decided to work together and then start calling our projects nemotron reflecting sort of the um uh the collaboration between these teams. Since then Numicron has dramatically expanded. There's so many more teams that are part of the effort. Um and it's really important um that we have structured it uh in this open way inside of Nvidia. Um we are inviting um volunteers from around the company to come help build Nvidia's AI. Uh we think it's very important to the future of the company and as that vision continues to develop more and more people want to join. That's fantastic. We're really excited about that and it means that we then have to figure out uh, how to organize the work so that everybody has a chance to contribute and feel heard and feel like their ideas um, are ah, fairly evaluated, um, on the path towards impact. Um, ah, we have a formal process for doing that. We have an internal website where people um, share ideas and then those ideas um, are assigned to um, one of 25 different um, leads that are uh, you know, over various parts of building Nemotron, um, they interact with those ideas. Some of those ideas get further developed, um, some of those ideas get deferred until you know, the next, uh, the next time we go around building a new model. But um, we're trying to build Nemotron in an open and inclusive way. Um, uh, so that um, you know, we can really come together as a company to build it. I think, um, you know, organizations that figure out how to collaborate to build AI succeed, organizations that struggle with control over who owns the AI tend to uh, waste a lot of effort. Um, and so Nvidia's success and Nematron success I think is directly proportional to our ability to collaborate. Something that I care deeply about.

Speaker B: Fantastic. But you mentioned uh, earlier that um, despite the fact that um, you work at the number one undisputed leader in GPUs, uh, you all as a research organization don't have all the GPUs uh, that you would want in the world. So how does the allocation of GPUs and compute happen? Uh, is that based on how promising an idea is or early success? Do you give uh, GPUs, withdraw GPUs uh, based on success?

Speaker A: Well it's a really complicated question and it's obviously a difficult problem for everyone in the industry to figure out how to allocate their compute, um, inside uh, Nematron. So we have a budget, ah, for Nemotron, um, and inside Nematron we allocate compute based on what we think the needs of the project are. We have a um, ah, hierarchy. So we have a set of programs and inside of each program we have a set of projects and each of them put forward their requests. Um, and then we have a two week cycle where we review requests and we review the budget and then we make decisions, um, in kind of a hierarchical way. And then um, uh, compute gets decided that way. Um, now having said that, um, this is something that I think we can still do better at. Um, uh, it's hard when we're making decisions about compute allocation because um, every researcher is convinced that their idea uh, could change the world if it just got a thousand times more GPUs attached to it. Right? And they might be right. It might actually be true. And yet we're running at the limit. We don't have 1000x more GPUs. For every idea that we have, we have to operate within the limits that we have. And so it is, um, a challenging process. We try to incorporate as many people's, um, perspectives into that as possible so that it's as much as possible a shared sense of understanding, uh, maybe not agreement. So there may be times when one project feels like it really deserved more GPUs, because the impact of that would have been so high, but it didn't get it. We hope in that circumstance that they have an understanding of why some other project did get more GPUs and why that was considered more of a priority during this particular allocation round, um, for the company, um, so that people can at least, uh, understand, you know, that there's a reason, uh, for the allocations that we have. Um, uh, having said that, um, you know, this process is always improving. There's always more work to be done to make this more transparent and more fair. Um, and, and then of course, um, my m, my number one is just to get more GPUs so that, um, you know, we can also fund more things. Because I would like to do that too.

Speaker B: How do you balance useful research with great exploratory research?

Speaker A: My belief is that research needs to be bootstrapped. Um, research is a chicken and egg problem. So, um, it is always the case that every researcher believes. If I just had a lot more resources, my idea would change the world. Actually, it's important that researchers feel that way because if you didn't feel that way, you wouldn't have the conviction that's required to go do something crazy and new. Right? So you have to believe, um, and so of course, um, uh, you start with that belief. Uh, but then how do you translate that belief into something that other people can understand, right, that other people are willing to invest in? Um, this is what I call the chicken and egg problem, right? Because, like, once your research ideas is obviously good and impactful, it's easy to get resources. But how do you get it to be obviously good and impactful without those resources? Right? So the way you solve chicken and egg problems is by bootstrapping. This is an iterative, uh, problem solving approach where you do something small, you get some sort of signal about this is a good idea and you tell people about that and then you ask for just a little bit More. And um, if people saw like, oh yeah, that, you know, that experiment turned out pretty well, that's pretty intriguing. We should probably do a little bit more there. Um, then you're on track, right? And uh, that um, over time, you know, iterate a lot, iterate quickly, iterate many times, uh, you can bootstrap to finding significant resources for your idea and also usually attracting more people, um, to come along with it, uh, on the way. Because they have a chance to see that this idea is going to change the world and then they want to be part of it.

Speaker B: Is that how the moonshots at uh, Nvidia got started as well over the years? Whether that's in AI or otherwise. So it was bottoms up, somebody coming up with a good idea versus Jensen saying this is what we need to do.

Speaker A: Well, you know, Jensen has lots of good ideas too and so the company is very responsive to his ideas. And that's uh, that's important as well. But Jensen very explicitly says all the time, this is a company of volunteers. You know, uh, each of us is here because we choose to, we could, we could be doing something else our lives, but we choose to be here. Um, and so uh, you know, we, we tend to make decisions, um, uh, especially for early stage research, it tends to be very bottoms up, uh, because, um, you know, it's sort of an invitation like bring, bring your best ideas. Let's, let's figure out, you know, what are all of our best ideas and then we'll take a step from there. Um, now do we sometimes have top down, um, ideas, uh, that are important for the company strategy? Of course, you know, of course, um, uh, NVFP4 pre training is one of those. So we decided as leadership of the company, we're going to really invest in FP4 hardware. Um, now it's time to go invent some optimization algorithms that succeed in using it. Um, uh, so we told the team, we didn't say to the team, you have to work on NVFP4 pre training. What we said is there's an opportunity, we're making a big investment and if we can figure this out, it will be significant for our company. And then we let the people who are interested in that work on that. And as a result we succeeded. You know, so, um, so it is a balance of like bottoms, um, up and tops down. Um, um. But uh, it always has this bootstrapping feeling even, even with something like NVFP4 where there's a significant like strategic top down component. The actual technical solution which is very intricate and complex and has a lot of moving parts that came, uh, from the researchers themselves. And, you know, that's my belief, is that research always comes from the researchers themselves. You can't tell research exactly how to go solve a problem because then it wouldn't be research, it would be engineering. But, uh, in a world of AI, where the most important problems we have to solve all have this research component, there needs to be freedom for researchers to innovate if we're going to make progress.

Speaker B: Listening to everything you're saying, I'm struck by how entrepreneurial the culture at Nvidia still seems to be. So, like, it's a very large company, so I'm sure there's all sorts of, uh, politics, and you mentioned the tribal instincts. I'm sure all of this is happening, but, um, especially given how long the company's been around, the phenomenal success, the fact that people have been making a lot of money internally, it still seems to be very entrepreneurial. Bottoms up driven, maybe meritocratic. Is that the right takeaway? Uh, yeah.

Speaker A: I mean, um. One thing that's very unusual about Nvidia is the tenure of its leadership. Jensen Huang has been running the company for 33 years. But he's not alone. There are a lot of other very senior leaders in the company who have been there for three decades or longer, uh, including my boss. And, um, these people remember what it feels like to work at a very small Nvidia, and they know what it feels like to work at a very large Nvidia. Um, they have a shared sense of ownership for the company. You know, um, Nvidia is a place we often say no, uh, one fails alone. Um, and the point of that, uh, that's just a statement of fact, right? You work at a company, it's one company. You all succeed together, you all fail together. You work in accelerated computing. Accelerated computing is the composition of thousands of technologies. If any of them fail to deliver acceleration, the value is destroyed. It doesn't matter whether the chip is great, if the compiler sucks. At the end of the day, the thing that you're selling is time and capability to researchers that are trying to build the future of AI. And if they don't get that, it doesn't matter whether it was the, you know, the transistor or the math unit or the compiler or the library or the networking or anything else along the way that, that failed to live up to its, um, expectations. The whole thing in composition fails. The whole value is destroyed. And so we have a deep understanding of that culturally at Nvidia. Uh, and it is something that motivates the way that we work together.

Speaker B: Maybe to close the conversation, I'd love to zoom out, get your take from the perspective of somebody who's as deep into all of this as it gets about uh, where things may be going. So like who knows in a few years, but I don't know, in the next year or two maybe there's some visibility. I read somewhere that you're not necessarily a big singularity, uh, kind uh, of person. Is that fair?

Speaker A: True.

Speaker B: And why is that?

Speaker A: Well, I think that intelligence is just so incredibly multifaceted. Um, I always think uh, about this question, like um, uh, if a uh, company were to be looking for its next CEO, would it find the next CEO by looking for somebody who won the International Math Olympiad? Probably not. Right? Even though it's incredible for people. Like I could never even compete in any way at the International Math Olympiad. And those people are amazing. Right? They have just incredible brilliance. That's not the right kind of brilliance to run a company. Um, if we look for example at um, other aspects of our culture that are really important. Um, for example musicians, what kind of intelligence does it take to become a hit musician? Um, don't assume that it's all luck. It's not. These people are working hard and they're very smart in ways that I might not understand with my PhD. Right. I might not have that kind of intelligence. Um, and so when I think about intelligence, I think it's just so multifaceted and so contextual. You know, it really depends on the situation. It's not just about raw intelligence. Raw intelligence is kind of like the horsepower of an engine. But an engine running without wheels doesn't go anywhere. Right? So, so intelligence, um, the impact of intelligence has a lot to do with um, the context that the intelligence has put in the harness, the platform. And so when I think about that I think um, uh, the singularity is. Although it's an attractive idea, I think that it's really a wrong headed idea because it doesn't really um, take into account these other factors. Um, so I believe that artificial intelligence is going to continue to develop at a rapid pace. It's going to unlock significant capabilities uh, for people in every aspect of um, uh, of our world economy, people doing every kind of work. Um, I'm very excited about the opportunities that it's, that it's going to bring. Um, I am also a little bit concerned uh, with how we're going to manage the transition. So I Do think that transitions are hard for humans in general? Like we're, we're conservative generally. Um, and uh, you know there is going to be a lot of change. This is a profound change in the way that we think and the way that, that we work, the way that we learn. Um, uh, ultimately I have faith in our ability as humans to figure it out. Um, you know, we've done it in the past. This is how, this is who we are. Um, we, we build tools, uh, we build external organs that help us solve problems. You know, we, we have an external stomach. We call it a kitchen. It creates enormous value for us. We can eat things that we couldn't eat without a kitchen. Right now we're creating an external brain. The implications of the external stomach were pretty profound, uh, for us as a species. They led to agriculture, which led to organized societies, the way our cities are built. So we think about what is the implications of an external brain. Pretty profound. Nobody actually really knows. Um, but what I do believe in is um, uh, the power of humanity to solve problems and to learn and to incorporate new technologies in ways that benefit us. Um, I also believe that the problems we face as a planet all require more intelligence. Every single one of them. Whether that's inequality, uh, or climate, um, change or um, uh, any of the other uh, structural uh, problems uh, that I think are very worrisome that we face. The solutions to those are going to require invention and intelligence. And what that means for me is that the only kinds of tools that we can really create moving forward are going to be AI. Uh, because the problems that we face are all about intelligence. And regardless of the technological approach to solving those problems, uh, the solutions will always be called AI. Um, and uh, so that um, makes me hopeful for the future but also somewhat um, respectful of the challenge that it is going to bring to us as we try to figure out how to live in a new way with this new external brain. Um, but I believe in our ability to learn and to change. Um, and I think um, ultimately this is going to make our lives better.

Speaker B: Do you guys feel the AI backlash that seems to be forming internally? Is that something that you all perceive, think about, and if so, do you think it's a communication problem that our industry may have? In particular, given what you just said about ah, all the obvious potential of AI?

Speaker A: You know, I'm always worried about the way that the public thinks about technology and interacts with it. It matters a lot. Um, and it is definitely the case that um, societies that want technological advancement have more technological advancement than societies that don't want change. Um, so I think it is, um, actually important to think about it. One thing that's interesting about AI is that um, I believe it tends to be much uh, more accepted when it is uh, part of everyday life. And at that point people stop thinking about it as AI. It's just, oh, this is the tool that I use. Like, do you care whether it's AI that's helping you, uh, route your car? When you ask the map application to help you drive somewhere? I mean it is. There is actually sophisticated AI that's going into that. Uh, but you're not really thinking about that, right? You're just using a tool. And um, so I feel like people's acceptance of AI, um, comes with experience. Right? Um, the more experience we have working with it, the more we learn how to work with it productively, um, I think the more comfortable we become with it.

Speaker B: Uh, great, Brian. So it's been a fascinating conversation. Maybe as a very last question, to make sure we cover it, I want to make sure that we talk about safety. What is the state of safety currently and where does open source and closed source, uh, sort of fit in the safety conversation today?

Speaker A: Safety is on everybody's minds right now. Um, watching ah, the fable release and um, the way that the government interacted with that, um, uh, uh, I think uh, is a consequence of um, concerns about safety, about these models. They get stronger and stronger and then they could be misused. Uh, and um, there's different approaches to thinking about safety and trying to define safety. Um, I have maybe a slightly ah, unorthodox opinion about this, which is that I think open technologies are generally safer because there's more sunlight. You know, when more people are thinking about um, the safety of a technology and evaluating it and then contributing to making it safer. I think that's inherently, uh, safer than um, having a small group of people, um, being in charge of safety for everyone else. I also think with um, artificial intelligence, because it is really about ideas, it's really about um, exploring ideas in different ways, that diversity is more safe than monoculture. Um, and what that means is that there's going to be different beliefs. Like, diversity isn't just about like the easy stuff. Diversity is about the hard stuff. Like when people have deeply felt disagreements, they really, really, totally disagree with each other. Um, making it possible for people to explore uh, their ideas in a diverse way, I think is more safe than um, trying to create a walled garden where um, certain ideas are considered safe and certain ideas are considered unsafe. And, um, this is controversial, ah, in today's AI environment, um, which I think is interesting because we've had hundreds of years of tradition, uh, that speak directly to this. In, um, the United States, for example, we have, um, laws about freedom of conscious conscience and freedom, um, of speech. And, you know, it's not because we didn't consider for thousands of years, would it have been safer if we didn't have those? Right. We tried that. We tried actually having a monoculture about, like, these ideas are safe to talk about, these ideas are safe to believe. And we found that to be much less safe than a pluralism where, uh, we officially don't take a position about what ideas are safe. We actually found that is much safer as a society to, uh, um, support diversity than it is, uh, to try to keep everybody safe, um, top down. And so, um, I believe that open technologies for AI are inherently, uh, the safest way of building AI.

Speaker B: All right. Love it, uh, controversial take to close, uh, the conversation. Brian, it's been fabulous. Thank you so much. Really appreciate your spending time with us today.

Speaker A: Thanks for inviting me.

Speaker B: Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you, uh, at the next episode.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysisTraining Data · on Nvidia95 / 100
  • AI Is Producing More Code, but Is It Producing More Value?The Data Exchange with Ben Lorica · on Nemotron88 / 100
  • Kiwi co-founded world‑builder hits $2.5 billion valuationThe Business of Tech · on Nvidia87 / 100
  • AI CEO Series: Dr. Varun SivaramNext in Tech · on Nvidia86 / 100
  • Yuval Ariav, Solo GP of Symbol Capital, on Pre-Consensus Investing, Why NextSilicon Could Be Israel's NVIDIA, and the $1.1 Trillion Market Hiding in Plain SightInvested by Aleph · on Nvidia86 / 100
  • The New American Dream: Democratising InvestingThe Master Investor Podcast with Wilfred Frost · on Nvidia84 / 100

More from The MAD Podcast with Matt Turck

All episodes →
  • Cloudflare CEO: The Internet's Business Model Is Dead
  • The GPU Myth: State of AI Compute 2026 | Stephen Balaban
  • OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
  • State of Enterprise AI 2026: Aaron Levie on Tokenmaxxing, Rise of Headless, and AI-Proofing Your Job
  • OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
Explore the best B2B AI & Data podcasts →
All The MAD Podcast with Matt Turck episodes →