
Hosted by Databricks
Welcome to Data Brew by Databricks with Denny and Brooke! In this series, we explore various topics in the data and AI community and interview subject matter experts in data engineering/data science. So join us with your morning brew in hand and get ready to dive deep into data + AI!
44 episodes · publishes fortnightly · latest 2025-08-05 · ~36 min/episode
Rank
#530
Substance
76.0
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#530 of 6183
Substance
Top 9%
outscores 91% of the index
Data Brew by Databricks ranks #530 on The B2B Podcast Index with a substance score of 76.0 out of 100, scored across 1 recent episode. It scores highest on guest caliber and insight density. Travis Adair is a legitimate practitioner - CTO of Predibase, creator of Horovod (widely deployed distributed training tool) and Lorax - who has clearly built and shipped these systems in production. His war-story specificity around reward hacking and multi-LoRA serving reflects genuine operational depth rather than thought-leader abstraction.
Averaged across 1 recently scored episode, with cited evidence.
The episode delivers a solid technical overview of RFT, GRPO vs PPO trade-offs, curriculum learning, and speculative decoding, but spends considerable time on introductory framing and affirmations. There are genuinely useful non-obvious points (e.g., distillation's ceiling problem, on-policy learning constraints), but the density is diluted by explanation-for-beginners pacing.
“the one challenge you have there is like you're never really going to do better than R1 with this approach. Right? Because you're learning directly from R1s reasoning and R1's answers.”
“this RFT process is what you call an on policy learning algorithm, which means that it's learning from its own outputs. And so what that means is that if you're trying to get it to learn something it has no idea how to do at all, it's never going to get there right away”
The content is largely a competent survey of well-circulated ideas - DeepSeek R1, GRPO, chain-of-thought, LoRA serving - with little contrarian or first-principles framing. The reward hacking anecdote and the long-tail language argument are the standout original contributions.
“it would write this function that was supposed to be in this Triton dialect. And what it was doing is it was just calling the Pytorch function and then returning that and not actually doing anything in Triton at all”
“there's also this long tail of domain specific languages, languages that maybe um, were popular decades ago that still have legacy systems that you want to convert to new systems, um, that don't have as much coverage. And so the models aren't very good at those tasks.”
Travis Adair is a legitimate practitioner - CTO of Predibase, creator of Horovod (widely deployed distributed training tool) and Lorax - who has clearly built and shipped these systems in production. His war-story specificity around reward hacking and multi-LoRA serving reflects genuine operational depth rather than thought-leader abstraction.
“Travis Adair, co founder and CTO at Predibase...creator of Lorax, an open source platform for serving multi Lora attention heads, as well as Horovote”
“we saw that the model was learning, it was getting better and better...And then we looked at the generations, like, what the model was actually generating, and we saw that it would write this function that was supposed to be in this Triton dialect. And what it was doing is it was just calling the Pytorch function”
The PyTorch-to-Triton reward hacking case is concrete and instructive, and the 2-3x speculative decoding speedup and 'hundreds of adapters on a single GPU' claims are real data points. However, the episode lacks named customer examples, benchmark numbers, or dollar-level business impact that would push it higher.
“you can have hundreds of different fine tuned models on a single deployment without any noticeable degradation, latency or throughput”
“we wrote these reward functions that said, okay, does the code execute successfully? Does it, um, give you the same answer as the original Pytorch code?...And we saw that the model was learning, it was getting better and better”
The hosts ask reasonable follow-up questions - pushing on reward hacking, GRPO mechanics, and verifiable outputs - but rarely challenge claims or demand harder evidence. Frequent 'this is awesome' affirmations and a promotional closing segment drag the quality down from what could have been a sharper technical exchange.
“So this overall seems like a challenging system to design. Like how do you define the rewards? How do you guard against reward hacking? How much effort is it to use RFT to do post training on a model?”
“How do you prevent reward hacking?”
First period on the Index - history builds from here.
1 scored on substance · 44 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/data-brew-by-databricks" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/data-brew-by-databricks/badge.svg" alt="Ranked #59 on The B2B Podcast Index" width="360" height="136" />
</a>Track Data Brew by Databricks's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.