
Hosted by Sequoia Capital
Listed under Technology, Business
Join us as we train our neural nets on the theme of the century: AI. Sonya Huang, Pat Grady and more Sequoia Capital partners host conversations with leading AI builders and researchers to ask critical questions and develop a deeper understanding of the evolving technologies - and their implications for technology,…
105 episodes · publishes weekly · latest 2026-08-04 · ~43 min/episode
Rank
#1
Substance
95.0
/ 100
Breakdown
Scored 2026-08
Updated monthly
Across the index
#1 of 1170
Substance
Top 1%
outscores 100% of the index
Training Data ranks #1 on The B2B Podcast Index with a substance score of 95.0 out of 100, scored across 2 recent episodes. It scores highest on guest caliber and specificity & evidence. Both founders bring exceptional credentials: one from OpenAI's early team (GPT1/GPT2 scaling laws), the other with PhD in theoretical CS turned deep learning researcher with protein folding work. Both have demonstrable execution at meaningful scale with pharma partnerships (Eli Lilly, Novartis, Pfizer) and product-market fit results. This is high-caliber operator experience.
Averaged across 2 recently scored episodes, with cited evidence.
The episode contains substantial technical and strategic insights about protein design, scaling laws, and business model decisions. The guests explain specific approaches (simplicity in model architecture, diffusion models, evaluation frameworks) and concrete progress metrics (15% vs 0.1% hit rates), though some sections include repetition and broader context-setting that dilutes density.
“when you look at a model, like, let's say chi one, I think there are 23 distinct sub modules in chi one. Um, and like, when you're trying to iterate on something like that, it gets really hard because you're like, I kind of need to understand each of these sub modules independently.”
“We got to uh, with our Chi 2 model, about like a 15% success rate. So uh, now if you screen 1,000 molecules, you're getting 150 back.”
The guests offer genuinely original framing of drug discovery as an engineering problem rather than exploration, articulate a contrarian partnership model versus full-stack drug development, and connect lessons from LLMs to biology in substantive ways. However, some frameworks (scaling laws, bitter lesson) are now well-known in AI circles.
“The question is how do we take drug discovery and make it look a lot more like drug design?”
“it used to be like, you either want to be first in class or best in class. Now it's like you want to be last in class because you actually just want to be the final answer.”
Both founders bring exceptional credentials: one from OpenAI's early team (GPT1/GPT2 scaling laws), the other with PhD in theoretical CS turned deep learning researcher with protein folding work. Both have demonstrable execution at meaningful scale with pharma partnerships (Eli Lilly, Novartis, Pfizer) and product-market fit results. This is high-caliber operator experience.
“I really started my career at OpenAI. So it was on the early team there. It was a nonprofit back then, so it was a pretty good time to be there. We did GPT1, GPT2, scaling laws and M.”
“My background was never biology. I studied pure math and started my PhD in, uh, theoretical computer science. And it was only after my third year that I ended up switching into deep learning. Protein structure prediction”
The transcript includes concrete metrics (0.1% to 15% hit rates, 23 modules in Chi1, 128 initial GPUs, 2018/2020 protein folding milestones, specific pharma partners by name), timelines (2024 founding, 9 months to 9 days iteration), and specific scientific results. However, many claims lack independent verification (exact model numbers, binding affinities) and some discussion remains at abstraction level.
“the state of the art for anybody design was about like a 0.1% binding rate. So one in a thousand of the molecules you design would actually bind in the lab. Um, so first of all it means you have to screen a lot of molecules”
“We got to uh, with our Chi 2 model, about like a 15% success rate. So uh, now if you screen 1,000 molecules, you're getting 150 back.”
The host asks clarifying questions and follows up on technical concepts (diffusion models, scaling laws, competitive landscape), but rarely pushes back or challenges claims. Questions tend to invite elaboration rather than provoke disagreement. Some moments lack sharpness (unicorn analogy, naming origin question feel tangential).
“Can we talk about what is state of the art today? And then maybe let's take a little trip down memory lane. Five years ago, three years ago, one year ago, like what have been some of the major breakthroughs and like how has state of the art changed over the recent years?”
“So let me ask you a question on that then. So there's. And tell me if this is a reasonable way to frame it. There's almost this boundary between that which can be engineered and that which needs to be tested in the real world.”
2 periods tracked.
2 scored on substance · 65 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/training-data" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/training-data/badge.svg" alt="Ranked #1 on The B2B Podcast Index" width="360" height="136" />
</a>Track Training Data's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
Companies, products and tools that come up most across this show's episodes.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.