
Hosted by Seth Levine
A machine learning podcast that explores more than just algorithms and data: Life lessons from the experts. Welcome to "Learning from Machine Learning," a podcast about the insights gained from a career in the field of Machine Learning and Data Science.
14 episodes · publishes occasionally · latest 2025-10-17 · ~68 min/episode
Rank
#132
Substance
82.2
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#132 of 6183
Substance
Top 2%
outscores 98% of the index
Learning from Machine Learning ranks #132 on The B2B Podcast Index with a substance score of 82.2 out of 100, scored across 5 recent episodes. It scores highest on guest caliber and insight density. Leland McInnis is a highly credible practitioner and researcher who has built production-grade, widely-adopted open-source tools (UMAP, HDBSCAN) used across industry and academia. He combines deep theoretical foundations (pure mathematics, algebraic topology) with practical implementation experience. He's not a career podcast guest but a working researcher who has directly shaped the field. His caliber is genuinely high - he's created infrastructure that others depend on.
Averaged across 5 recently scored episodes, with cited evidence.
The episode contains solid technical insights about dimensionality reduction, clustering, and unsupervised learning fundamentals, particularly around the geometry-first approach and how to decompose algorithms into components. However, much of the discussion rehashes well-known concepts (TSNE limitations, the importance of visualization, parameter tuning challenges) without introducing particularly novel or non-obvious claims. The conversation touches on genuine insights but doesn't pack them densely - there are extended passages of setup and context-building that don't add new information.
“One of the biggest difficulties with unstructured data right now is that we have great tools for search and picking through, but that assumes you already know what you're looking for.”
“I see HDB scan as a pile of Lego bricks, there are a bunch of different things. So there's there's a density estimation step, there's connectivity related step, there's building a tree of clusters step, and then there's a cluster extraction step.”
The core thinking - viewing algorithms through algebraic topology and geometry, treating them as composable components - is genuinely distinctive and contrarian to the black-box ML norm. However, the episode doesn't push these ideas into particularly novel territory. The guest recycles familiar critiques of LLM hype, RAG systems, and the importance of simplicity. The advice about not following hype and doing interdisciplinary work is sound but not new to research discourse.
“I tend to think of things in terms of algebraic structures and the geometry of things. So I'm always thinking in terms of the geometry of the data.”
“being able to decompose them back again into the parts that you need...you could easily convert that to topic modeling for images”
Leland McInnis is a highly credible practitioner and researcher who has built production-grade, widely-adopted open-source tools (UMAP, HDBSCAN) used across industry and academia. He combines deep theoretical foundations (pure mathematics, algebraic topology) with practical implementation experience. He's not a career podcast guest but a working researcher who has directly shaped the field. His caliber is genuinely high - he's created infrastructure that others depend on.
“He's the maintainer for many machine learning packages including UMAP, HDB scan, PyNN descent, Datamap plot, which has created an ecosystem for unsupervised learning”
“I worked in industry and with government for a couple years before heading back to do a PhD...actually working with real data, as opposed to pure theoretical math”
The episode lacks concrete numbers, named use cases beyond high-level examples, and specific metrics or benchmarks. While McInnis mentions COVID research and artist Refik Anadol, he doesn't provide details: no performance improvements quantified, no actual datasets named, no real adoption metrics. The discussion of speedups (n-squared to n log-n) is mentioned but not elaborated. For a technical deep-dive, the conversation remains frustratingly abstract.
“It was like what and n squared and and log n or something and log n. Yeah, had to be looking up everything for every point before and then using minimum spanning trees”
“I was very surprised to see it coming up in COVID research repeatedly”
The host asks competent follow-up questions and demonstrates genuine familiarity with the guest's work, which grounds the conversation. However, the questioning is mostly affirmative and exploratory rather than challenging. The host rarely pushes back, tests claims, or creates productive friction. There are few moments where skepticism or hard questions emerge. The conversation flows naturally but doesn't excavate deeper contradictions or force the guest to defend positions.
“I've used, well, yeah, I mean, all of your libraries a lot. But UMAP in particular, you really see how using different parameters, you can get very wildly different results. Sometimes it's hard to know and evaluate if you're, how do you know if you're reducing dimensions correctly?”
“one of the challenges in unsupervised learning or topic modeling in particular is something like trying to find new topics over time, right, trying to incorporate the temporal feature. Have you thought about that?”
First period on the Index - history builds from here.
10 scored on substance · 14 tracked in total.
Dan Bricklin: Lessons from Building the First Killer App | Learning from Machine Learning #14
2025-10-17 · 1h 13m
Lukas Biewald | You think you're late, but you're early | Learning from Machine Learning #13
2025-07-01 · 1h 5m
Maxime Labonne: Designing beyond Transformers | Learning from Machine Learning #12
2025-05-28 · 1h 4m
Aman Khan: Arize, Evaluating AI, Designing for Non-Determinism | Learning from Machine Learning #11
2025-04-29 · 1h 7m
Leland McInnes: UMAP, HDBSCAN & the Geometry of Data | Learning from Machine Learning #10
2024-10-25 · 55 min
Chris Van Pelt: Machine Learning Tooling, Weights and Biases, Entrepreneurship | Learning from Machine Learning #9
2024-03-01 · 1h 5m
Michelle Gill: AI-Assisted Drug Discovery, NVIDIA, Biofoundation Models, Creating Applied Research Teams | Learning from Machine Learning #8
2024-01-11 · 1h 6m
Ines Montani: Explosion, NLP, Generative AI, Entrepreneurship | Learning from Machine Learning #7
2023-10-26 · 1h 23m
Lewis Tunstall: Hugging Face, SetFit and Reinforcement Learning | Learning from Machine Learning #6
2023-10-03 · 1h 19m
Paige Bailey: Google Deepmind, LLMs, Power of ML to improve code | Learning from Machine Learning #5
2023-05-19 · 1h 8m
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/learning-from-machine-learning" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/learning-from-machine-learning/badge.svg" alt="Ranked #20 on The B2B Podcast Index" width="360" height="136" />
</a>Track Learning from Machine Learning's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
Companies, products and tools that come up most across this show's episodes.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.