
Hosted by Claire Vo
How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life.
89 episodes · publishes weekly · latest 2026-06-30 · ~42 min/episode
Rank
#4410
Substance
53.0
/ 100
Breakdown
Scored 2026-07
Updated monthly
Across the index
#4410 of 6183
Substance
Top 71%
outscores 29% of the index
How I AI ranks #4410 on The B2B Podcast Index with a substance score of 53.0 out of 100, scored across 1 recent episode. It scores highest on specificity & evidence and insight density. The episode is meaningfully grounded in real numbers - token pricing, SWE-bench scores, pass rates, generation counts - and names actual models tested in a blind setup, which is more rigorous than pure vibe episodes; however the actual leaderboard results are narrated vaguely ('I gave it a four,' 'not bad') without sharing the underlying score distributions clearly.
Averaged across 1 recently scored episode, with cited evidence.
The episode surfaces a few genuinely useful observations - LLM judges regress to the mean, human taste diverges systematically from automated scores, saturated agentic tasks don't differentiate models - but these are embedded in a lot of rambling process narration and 'we'll see when we get the scores' throat-clearing that dilutes the idea density considerably.
“every model is kind of an easy judge... every model sort of rates to the middle of the bell curve. This is one of the challenges that I have had with self-grading evals”
“I don't think these models are spiky enough when it comes to how they evaluate output”
The hybrid human-vibe-plus-LLM-judge design and the explicit acknowledgment that model-as-judge introduces systematic generosity bias are genuinely underexplored points; the 'Claude slop' bias angle is a fresh framing. However the core thesis - different models excel at different tasks - is thoroughly well-worn territory.
“I hate Claude Slop deeply and I have like a big eye for Claude Slop. And so I just see the tells of Claude style writing and it drives me crazy”
“the model thought was good, I thought was bad. Why do we disagree? Well, every model is kind of an easy judge”
This is a solo episode with no guest whatsoever; the host has some relevant practitioner background (product design/engineering, mentions building ChatPRD) but presents as a content creator narrating a personal workflow rather than an operator with deep institutional authority. The only external voice referenced is a brief citation from a prior episode.
“In my episode with Felix from Anthropic, he says that we're all abusing Opus and we should definitely be using the Sonnet models more”
“I've been a product design engineering leader for a while. I can eyeball stuff and make it go fast”
The episode is meaningfully grounded in real numbers - token pricing, SWE-bench scores, pass rates, generation counts - and names actual models tested in a blind setup, which is more rigorous than pure vibe episodes; however the actual leaderboard results are narrated vaguely ('I gave it a four,' 'not bad') without sharing the underlying score distributions clearly.
“it's $2 per million input tokens and $10 per million output tokens, at least through the end of the summer”
“it's not quite at this 69 on agentic coating sweet bench pro or the 82 on terminal bench 2.1”
There is no interview or conversation - this is an unscripted solo walkthrough with frequent 'let's see' reveals and unresolved threads, which undermines the clarity a good solo essay format requires; the host openly admits ignorance of their own benchmark results mid-episode, which reads as performative rather than illuminating.
“The evals are not quite done running. So they're running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode”
“I have not seen this yet. We're going to go through it live. It's even going to surprise me. This is truly neutral. No bias.”
First period on the Index - history builds from here.
1 scored on substance · 60 tracked in total.
Add this badge to your site - it links back here and updates automatically as you rank.
<a href="https://index.fame.so/show/how-i-ai" target="_blank" rel="noopener">
<img src="https://index.fame.so/badge/how-i-ai/badge.svg" alt="Ranked #393 on The B2B Podcast Index" width="360" height="136" />
</a>Track How I AI's rank
Get an email whenever this show moves up or down the Index. Monthly at most, no spam.
Companies, products and tools that come up most across this show's episodes.
The themes that come up most across this show's episodes.
Podcasts that dig into the same topics.