
Hosted by Utsav Shah
Software at Scale is where we discuss the technical stories behind large software applications.
60 episodes · publishes fortnightly · latest 2024-08-05 · ~58 min/episode
Rank
#1044
Substance
72.0
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#1044 of 6182
Substance
Top 17%
outscores 83% of the index
Software at Scale ranks #1044 on The B2B Podcast Index with a substance score of 72.0 out of 100, scored across 1 recent episode. It scores highest on guest caliber and specificity & evidence. Aravind Suresh is a legitimate staff-level practitioner who spent five years on Uber's data infrastructure team working on systems that scaled to over an exabyte, and now works at OpenAI - real hands-on experience at scale rather than a thought-leader. The score is tempered because answers sometimes stay high-level rather than exploiting that depth.
Averaged across 1 recently scored episode, with cited evidence.
The episode contains genuine technical depth on file immutability, Hudi's commit model, and batch vs. real-time trade-offs with a specific cost estimate, but is heavily padded with 'it depends' answers, generic origin stories, and broad conceptual overviews that any data engineer would already know. Insight rate is uneven across 35 minutes.
“deleting like specific rows there is literally no other way other than building some kind of a reverse index which kind of tells you oh this id is kind of stored across these files and then rewriting those files”
“batch systems tend to be optimized more for throughput the amount of data you can process per unit time right uh whereas real-time systems are more or less optimized for latency or response times”
The content is largely explanatory and educational - covering well-established data engineering concepts like OLAP vs OLTP, batch vs. real-time, and open-source ecosystem history. There is little that is contrarian, first-principles, or counter-intuitive; even the Hudi discussion stays at a conceptual level.
“asking such questions to yourself goes a long way when like deciding which is better and the systems built the right way would obviously be more maintainable over time”
“of course there are different details twofold. One is, where is data that comes into your vector database?”
Aravind Suresh is a legitimate staff-level practitioner who spent five years on Uber's data infrastructure team working on systems that scaled to over an exabyte, and now works at OpenAI - real hands-on experience at scale rather than a thought-leader. The score is tempered because answers sometimes stay high-level rather than exploiting that depth.
“i spent quite some time, around five years in this team, and I got an opportunity to work on various verticals of the stack, be it batch, real-time, something related to failover”
“the amount of data went from around grew by a factor of 20 or something in that three four years period from 2019 and when i moved out it was more than an exabyte”
The episode delivers genuine specific data points - exabyte-scale storage, ~1 PB/day ingestion, 60-70% raw data ratio, 10-100x real-time vs. batch cost multiplier, and named internal tools like DataBook - but many answers retreat into vagueness and the numbers are frequently hedged with 'I vaguely remember' or 'somewhere around'.
“from 2019 and when i moved out it was more than an exabyte so the that's kind of the scale so it was starting with like uh it was around few pbs somewhere around 2018”
“maybe right now i would put it somewhere around close to a petabyte or maybe even more on a daily basis that gets ingested”
The host makes a genuine effort to extract numbers (prompting the 10-100x estimate) and follows up on deletion mechanics, but repeatedly accepts vague 'it depends' answers without pressing for specifics and asks overly broad scene-setting questions that eat significant time.
“like 10x or 100x how do i know”
“how do you actually perform the delete reliably”
First period on the Index - history builds from here.
1 scored on substance · 60 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/software-at-scale" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/software-at-scale/badge.svg" alt="Ranked #86 on The B2B Podcast Index" width="360" height="136" />
</a>Track Software at Scale's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.