
Hosted by Heavybit
Exploring the observability side of software development.
90 episodes · publishes monthly · latest 2026-06-04 · ~38 min/episode
Rank
#427
Substance
77.0
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#427 of 6183
Substance
Top 7%
outscores 93% of the index
O11ycast ranks #427 on The B2B Podcast Index with a substance score of 77.0 out of 100, scored across 1 recent episode. It scores highest on guest caliber and insight density. Janaki is a hands-on engineer who built the production system under discussion, citing real internal workflows, specific error taxonomies, and named Amplitude products she shipped - solidly a practitioner guest, not a thought leader - though her seniority level is engineer rather than a senior decision-maker or founder, which limits the strategic depth on offer.
Averaged across 1 recently scored episode, with cited evidence.
The episode contains a handful of genuinely useful ideas - compounding accuracy loss across agent steps, the primacy of 'the harness' over model upgrades, and the manual-tagging-to-auto-tagger pipeline - but these insights are spread thin across significant filler, repeated explanations, and a multi-minute detour into the guest's calendar art project that is irrelevant to any B2B operator.
“if each step is about 95% the way there, or 95% accurate, that leads to your response being Napkin math, about 60% accurate”
“a lot of the meat of how well our system does is related to, um, the harness around our agents rather than just upgrading the models”
The core framing - 'every failure becomes an eval' and treating manual tagging as seed data for automated evaluators - is practically grounded but already circulates widely in the AI engineering community; there are no genuinely contrarian or first-principles arguments, and the stone soup metaphor is borrowed from the guest's own pre-existing article rather than developed in the conversation.
“we parsed a lot of our internal notebooks and amplitude and generated nice tight evals from that that give us an input question”
“real investigations that took our own PM's hours to do then turned into evals that we were able to test our global agent system on”
Janaki is a hands-on engineer who built the production system under discussion, citing real internal workflows, specific error taxonomies, and named Amplitude products she shipped - solidly a practitioner guest, not a thought leader - though her seniority level is engineer rather than a senior decision-maker or founder, which limits the strategic depth on offer.
“I help build some of the systems that allow us to do analytics much faster”
“we just launched a product called Global Agent at Amplitude, which does the job of an analyst”
The episode delivers useful concrete detail - named products (Global Agent, Agent Analytics), a specific funnel-chart eval broken into discrete sub-checks (six steps, one-hour conversion window), and the 95%-per-step compounding math - but lacks hard business metrics such as user counts, error-rate improvements, or revenue impact, and the model naming is inconsistent ('Sana 4.5' vs Opus 4.6), undermining precision.
“Does the funnel chart now have the six steps that we're expecting for this particular storefront? Did it use the right conversion window of an hour?”
“if each step is about 95% the way there, or 95% accurate, that leads to your response being Napkin math, about 60% accurate”
Jessica asks several sharp, well-timed follow-ups ('How do you check your checker?', 'Is there a separate eval process pre-release?') and Ken usefully references the guest's written article, but neither host meaningfully pushes back on any claim, the art tangent consumes several minutes of irrelevant airtime, and the conversation never reaches productive disagreement or stress-tests the guest's framing.
“How do you check your checker?”
“Is there a separate eval process pre release for. We ask it the standard set of questions with this standard data set and see if it does a good job.”
First period on the Index - history builds from here.
1 scored on substance · 60 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/o11ycast" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/o11ycast/badge.svg" alt="Ranked #49 on The B2B Podcast Index" width="360" height="136" />
</a>Track O11ycast's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
Companies, products and tools that come up most across this show's episodes.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.