
Hosted by Machine Learning Street Talk (MLST)
Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis.
255 episodes · publishes weekly · latest 2026-07-01 · ~76 min/episode
Rank
#61
Substance
85.0
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#61 of 6183
Substance
Top 1%
outscores 99% of the index
Machine Learning Street Talk ranks #61 on The B2B Podcast Index with a substance score of 85.0 out of 100, scored across 1 recent episode. It scores highest on guest caliber and specificity & evidence. The guests are the actual winning team of the ARC-AGI-3 preview competition with hands-on technical depth - Dries Smith built the stochastic goose algorithm from scratch in two weeks, Stefano ran RL from scratch experiments, Michael provided detailed scoring analysis - making them highly relevant practitioners rather than career podcast guests, though they are competition participants rather than field-leading researchers.
Averaged across 1 recently scored episode, with cited evidence.
The episode contains genuine technical insights about the stochastic goose approach, action efficiency scoring mechanics, and LLM prior leakage into ARC game design, but these are diluted by extended philosophical meanders on emergence, Chollet's nativism, and Conway's Game of Life that add little actionable content. The density varies sharply between technical passages and discursive ones.
“36% might be misleading as a number if you don't look behind it. So what it really measures is action efficiency”
“if you just remove that priors, uh, which shouldn't actually be priors, you can actually see it becomes harder for humans to play”
A few genuinely fresh observations stand out - the point that 36% score hides that frontier models actually solve two-thirds of training games but inefficiently, and that human-made games inevitably leak priors even when designed to strip them - but much of the philosophical framing around LLM representations, bitter lesson, and Chollet's views recycles standard ML discourse.
“36% actually in this case doesn't mean we solve that approach solves 36% of the games. It solves way more of the games from the training set... but it just solves them inefficiently”
“There's some leakage of human priors into the games”
The guests are the actual winning team of the ARC-AGI-3 preview competition with hands-on technical depth - Dries Smith built the stochastic goose algorithm from scratch in two weeks, Stefano ran RL from scratch experiments, Michael provided detailed scoring analysis - making them highly relevant practitioners rather than career podcast guests, though they are competition participants rather than field-leading researchers.
“for the um, agent uh, preview competition last year I actually tried something completely different... the solution was basically just to brute force actions... I could solve two games, almost solve the third game”
“I created a ARC environment that's procedurally generated with some new objects and new objectives... I required about 10,500 different permutations just to solve that one setup”
The episode is well-anchored in concrete numbers: action counts, grid sizes, token budgets, compute costs, and solve rates all appear with specificity, giving a clear operational picture of both the benchmark mechanics and the team's approach; the main weakness is that some claims about model capabilities remain qualitative.
“we have eight main actions but there's a mouse clean action which has around 4,000 possible places you can click 64 by 64”
“it costs like a few thousand dollars, um, which is a lot more... a lot more compute than the human beating these games”
The host is clearly technically literate and generates substantive follow-ups - pushing on transduction vs. induction, gamability of the benchmark, and the AGI validity question - but frequently delivers extended monologues that crowd out guest responses, and the philosophical tangents (Conway's Game of Life, emergence, Chollet simulation) feel self-indulgent rather than productive.
“you mentioned transductive as well, which is quite interesting because um, you know, roughly speaking, I think of transduction as you're making um, a prediction about this specific test instance. And it's quite an interesting discussion whether or not this is transduction”
“Is Anything in Arcade GI3 badly designed or gameable?”
First period on the Index - history builds from here.
1 scored on substance · 60 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/machine-learning-street-talk-mlst" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/machine-learning-street-talk-mlst/badge.svg" alt="Ranked #12 on The B2B Podcast Index" width="360" height="136" />
</a>Track Machine Learning Street Talk's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.