The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/How I AI
How I AI artwork

Sonnet 5 review: I ran 64 generations to find out if it's worth it

How I AI · 2026-06-30 · 26 min

0:00--:--

Key moments - from our scoring

Substance score

33 / 100

Five dimensions, 20 points each

Insight Density8 / 20
Originality8 / 20
Guest Caliber4 / 20
Specificity & Evidence9 / 20
Conversational Craft4 / 20

The host introduces Claude Sonnet 5, Anthropic's new model positioned as opus-level performance at sonnet-level pricing, but takes a different approach than typical model reviews. Rather than relying on subjective impressions, he develops the HowIAI Bench - a systematic, human-graded evaluation framework to test models on tasks that matter to builders: PRD writing, prototype generation, agentic code-based debugging, and agent voice personality. Using Claude Code, he constructed frozen test inputs, ran blind evaluations across five frontier models (Gemini 3 Pro, GPT-5.5, Sonnet 4.6, Sonnet 5, and Opus 4.8), scored 64 prototype generations by hand, and combined his subjective taste with LLM-based judging from Opus and GPT-5.5. The results surprised him: Sonnet 5 ranked lowest in his weighted index, while Sonnet 4.6 emerged as the strongest overall performer for builders, with GPT-5.5 excelling at PRD writing and Opus 4.8 best for complex UI design.

Key takeaways

  • →Sonnet 5 underperformed relative to expectations in blind testing, ranking last on the author's weighted preference index despite Anthropic's positioning as an opus-level alternative at sonnet-level prices.
  • →Human taste and vibe checks diverge significantly from LLM-based judging, with the author rating Sonnet 4.6 highest while automated models rated Gemini 3 Pro at the top, suggesting pure metrics miss subjective quality assessment.
  • →Model selection should be task-specific rather than universal: GPT-5.5 for PRDs, Sonnet 4.6 for prototypes and conversation, Opus 4.8 for complex UI design, based on the benchmarking results.
  • →Building custom benchmarks with Claude Code using past work context allows teams to establish repeatable, meaningful evaluation frameworks that balance automated scoring with human judgment through weighted indices.
  • →Saturated baseline benchmarks like standard agentic coding tasks don't differentiate between frontier models effectively and should be retired in favor of more demanding evaluation criteria.

In this episode

  1. 1Introducing Claude Sonnet 5 and Agentic Claims
  2. 2Building the How I AI Bench with Claude Code
  3. 3Benchmark Design: PRDs, Prototypes, and Agentic Tasks
  4. 4Running 64 Generations and Vibe Checks
  5. 5Model Leaderboard Results and Surprising Winners
  6. 6Comparing Human Taste vs Automated Scoring
  7. 7Task-by-Task Model Recommendations
  8. 8Finalizing the How I AI Index with Weighted Scoring

Mentioned

AnthropicClaude Sonnet 5Claude OpusRunwayClaude CodeGPT 5.5Gemini 3 ProGLM 5.2CursorQuad CodeHyper AgentChatPurity

Topics in this episode

AnthropicClaude CodeClaude Opus 4.8RunwayGPT-5.5Claude Sonnet 5Sonnet 4.6Gemini 3 ProHow I AI BenchHyper Agent

Questions this episode answers

What are the key pricing and performance claims for Claude Sonnet 5?

Anthropic claims Sonnet 5 delivers opus-level task performance at sonnet-level pricing ($2 per million input tokens, $10 per million output tokens through summer), with improved agentic tool use and longer-running sessions compared to Sonnet 4.6, though it doesn't quite match Opus 4.8's benchmark scores like 69 on agentic coding or 82 on terminal bench 2.1.

How did the host build the HowIAI Bench evaluation framework?

Using Claude Code, he created a benchmark with frozen test inputs across five tasks (PRDs, wireframes, prototypes, agentic code search, and agent voice), ran blind evaluations with models labeled A - E, manually scored 64 prototype generations on a 1 - 5 scale with qualitative notes, and had both Opus 4.8 and GPT-5.5 score the same outputs to compare human judgment against LLM judging.

Which model ranked best and worst in the host's final weighted HowIAI Index?

Using a 70% human-judgment / 30% LLM-judge weighting, Sonnet 4.6 ranked first, followed by Gemini 3 Pro and GPT-5.5, while the newly released Sonnet 5 ranked near the bottom, surprising given Anthropic's positioning.

What key task-specific model recommendations did the host make?

For PRD writing, use GPT-5.5 for comprehensive clarity; for prototyping, Sonnet 4.6; for agent personality and chit-chat, Sonnet 4.6; for complex UI design, Opus 4.8; and for code-base work, Opus 4.8 or Sonnet 5.

Why did the host disagree with the automated LLM judges on model quality?

Models tended to score toward the middle of the bell curve (average 7/10) and lacked the subjective taste, uniqueness, and human-eye judgment the host valued; he also detected Claude-specific stylistic patterns ("Claude Slop") that the automated rubrics missed but his human eye caught immediately.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

8 / 20

The episode surfaces a few genuinely useful observations - LLM judges regress to the mean, human taste diverges systematically from automated scores, saturated agentic tasks don't differentiate models - but these are embedded in a lot of rambling process narration and 'we'll see when we get the scores' throat-clearing that dilutes the idea density considerably.

every model is kind of an easy judge... every model sort of rates to the middle of the bell curve. This is one of the challenges that I have had with self-grading evals
I don't think these models are spiky enough when it comes to how they evaluate output

Originality

8 / 20

The hybrid human-vibe-plus-LLM-judge design and the explicit acknowledgment that model-as-judge introduces systematic generosity bias are genuinely underexplored points; the 'Claude slop' bias angle is a fresh framing. However the core thesis - different models excel at different tasks - is thoroughly well-worn territory.

I hate Claude Slop deeply and I have like a big eye for Claude Slop. And so I just see the tells of Claude style writing and it drives me crazy
the model thought was good, I thought was bad. Why do we disagree? Well, every model is kind of an easy judge

Guest Caliber

4 / 20

This is a solo episode with no guest whatsoever; the host has some relevant practitioner background (product design/engineering, mentions building ChatPRD) but presents as a content creator narrating a personal workflow rather than an operator with deep institutional authority. The only external voice referenced is a brief citation from a prior episode.

In my episode with Felix from Anthropic, he says that we're all abusing Opus and we should definitely be using the Sonnet models more
I've been a product design engineering leader for a while. I can eyeball stuff and make it go fast

Specificity & Evidence

9 / 20

The episode is meaningfully grounded in real numbers - token pricing, SWE-bench scores, pass rates, generation counts - and names actual models tested in a blind setup, which is more rigorous than pure vibe episodes; however the actual leaderboard results are narrated vaguely ('I gave it a four,' 'not bad') without sharing the underlying score distributions clearly.

it's $2 per million input tokens and $10 per million output tokens, at least through the end of the summer
it's not quite at this 69 on agentic coating sweet bench pro or the 82 on terminal bench 2.1

Conversational Craft

4 / 20

There is no interview or conversation - this is an unscripted solo walkthrough with frequent 'let's see' reveals and unresolved threads, which undermines the clarity a good solo essay format requires; the host openly admits ignorance of their own benchmark results mid-episode, which reads as performative rather than illuminating.

The evals are not quite done running. So they're running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode
I have not seen this yet. We're going to go through it live. It's even going to surprise me. This is truly neutral. No bias.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

sonnet32model26models24benchmark17opus16code13agentic12judge12test10claude9vibe9taste9bench8show8leaderboard8gemini8

Episode notes

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected. What you’ll learn: What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily -

Full transcript

26 min

Transcribed and scored by The B2B Podcast Index.

We've got a new model, people, and it's from Anthropic. Now, is it Mythos? No. Is it Fable?

No, but it is Claude Sonnet 5. Anthropic is claiming it's the most agentic sonnet model yet, and we will get opus-level tasks at sonnet-level prices. Now, I've been testing a lot of models, and I'm starting to get bored of doing the vibe check. What I want to start developing is a set of benchmarks we can regularly test these new models against that you'll care about.

So today I'm going to be introducing the HowEye AI Bench, a set of AI and Clairvaux graded benchmarks that are going to tell us if this model and any model is good at writing PRDs, solving bugs, and one-shotting designs. I'm going to show you exactly how I built this benchmark using Claude Code, and we're going to see on a blind test what comes out on top. Let's get to it. This episode is brought to you by Runway, a new kind of creative platform that has everything you need to generate any image, video, or piece of content you want, all in one place.

With Runway, it's now possible to go from initial idea to a finished deliverable in a matter of minutes. From turning low fidelity product shots into campaign ready imagery all the way through putting together big brand films, Runway can help your team scale your creative ambitions while keeping your budgets and timelines from doing the same. Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along with studios like Lionsgate and Legendary, all use Runway to ship real work every day.

Try it yourself at RunwayML.com slash HowIAI. Promo code HowIAI. quickly before we get to our evals let's just talk about the headlines of sonnet 5 this new model anthropic is pitching it as close to the performance of opus 4 8 but much less expensive so as you can see here it's not quite at this 69 on agentic coating sweet bench pro or the 82 on terminal bench 2.

1 but it's not that far behind and i suspect that most of us are not going to notice the difference. It's also supposed to be really good at computer work and knowledge work, and so this should be an everyday model that people reach for. In my episode with Felix from Anthropic, he says that we're all abusing Opus and we should definitely be using the Sonnet models more, and we are going to put Sonnet 5 to the test against that proposition. Now, what do they say that Sonnet 5 is really good at?

Well, it's really good at agentic tool use. So you're going to get slightly longer running tool runs, longer running sessions than you would with Sonnet 4-6 at a lower cost than doing the same comparable task with Opus. So you're going to see here, you know, Sonnet 4-6, a lower pass rate on these long running tasks, Sonnet 5 getting pretty close when you have extra high reasoning on, and then Opus of course has the highest pass rate, but it's also much more expensive. That holds true also with computer use.

So as you see, Sonnet 4.6, not bad, about 80% pass rate. But when you want to get past 80% into really successful computer use, browser use, etc., which is what I've been doing a lot lately, you're going to get a slightly cheaper experience, but almost as good as Opus 4.

8 when you're using Sonnet. And then the headline seems to be it's much more affordable than Sonnet. So it's going to be $2 per million input tokens and $10 per million output tokens, at least through the end of the summer. And then it's going to go up a little bit.

So if you want to test this model and you want to test it at launch prices, get that done now. So as I said at the beginning of the episode, I'm a little tired of doing me sort of like one-off vibe checks. Sure, I can put this into cursor, into quad code, one shot a landing page and kind of say, what do I think? And I've done this for a couple of models.

I've done it for GPT 5.5. I've done it for open weight models like GLM 5.2, but it always felt like my feedback on these models is kind of soft.

Yes, we put it against like specific workflows, but I don't like that it's not repeatable and I don't like that we're not testing it over time. What do I like about this process, though? I do like that it is a Clairvaux benchmark. I have a perspective.

I have a point of view of what's good and bad. And I don't want to lose that Clairvaux taste by doing an LLM in the loop or an AI as judge on these benchmarks. So I'm going to show you how I built and will build the Hawaii AI bench and on a blind kind of taste test, how these models did across a couple use cases. Okay.

What's really fun is the evals are not quite done running. So they're running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode about what I think of Sonnet 5 amongst all these other models. But I just want to show you how you can build your own evals benchmark for you to assess whether or not these new models are really working in your favor.

And so I have Cloud Code up here and I asked just a very simple question. Based on our work together, can you help me brainstorm a How I AI benchmark and eval set? We can test every time a new model comes out to consistently score different tasks that would be relevant to our podcast audience. Now, this is something that I hope everybody takes advantage of.

all your Cloud Code sessions are stored on your desktop. So you can actually go through those. Cloud can go through those and make recommendations on future work based on your past work. This also works for Codex.

So you can have Codex look at your old sessions. You can even have Codex look at your Cloud Code sessions and really use that in addition to its own memory to come up with new ideas. So that's what I did here. And it sort of gave me kind of some good design principles about what makes a good benchmark in general frozen inputs blind scoring where possible a rubric And then it came up with a list of tasks everything from taking messy notes and turning them into a PRD to one a landing page or an app to kind of going through lots of context and trying to come up with sighted information.

And I am not one to pick because I want everything. So I said, build the whole thing. I love this. And it started, And then I corrected myself and I said, let's actually focus on tasks for builders, PRDs, prototypes, agentic multi-step, and agentic voice basically does a pass the vibe check in my open claw.

I don't really care about long context and deep research. And then I said it could use my existing repos, some data sources, some things that we already did to build it. Now, what's interesting about how I built this is in addition to building the scored benchmarks where an LLM would actually score the outputs, I also said I want an HTML page at the end that I can give you vibe feedback. And then we will use my vibe feedback and the LLM scores to come up with the completely scientific how I AI bench and see what it came up with.

Now, this took about, I don't know, 45 minutes to run. I actually recorded an episode while it was running. And I just want to show you what it came up with and how I worked through it. What it did is it dropped all the outputs of the benchmark into one local HTML page where I could give it my own structured vibe check.

And as you can see here, it says just score each output one to five on pure gut feel. Would I ship this? Does it sound like me? It's going to save that to the browser.

It actually downloaded a JSON file, and then I use that to check the scoring. And so you can see here, I have a blind, I turned on blind, a blind set of models A through E. I believe we tested, although I should double check because I didn't really look, Opus 4A, 5.5, Sonnet 4.

6, Sonnet 5, and maybe GLM? I'm not actually sure what the fifth one was. We'll see when we get the scores and it made PRDs. And then I went through here and I read the PRDs and I gave it scores.

And so, you know, I would look at these and let's see if I can find one that I actually scored. And I would say something like this one is comprehensive and clear. I gave it a four. And so you can imagine each of those PRDs I went through and I gave them like a one to five score.

I put some like lightweight notes in and scored them. Now, this is where it gets interesting. I have a set of prototypes I run as an eval. I posted an article on X and LinkedIn about how we generated the same app 82 times at ChatPurity when we were building our own prototyping tool.

And I reused that harness to test prototyping and wireframe across a bunch of different apps and give those all vibe checks. So you can see here, these are complicated apps that each model generated a different version of. And you can see here, I gave this one kind of a four, not bad. It was simple.

I gave this one a four. There were a few issues at the top, too many icons. I said this one was good. It's very comprehensive.

So you can see I went through a complex. This is a doc scheduling app. This is an editorial assignment desk, something that maybe an editor or a blog would use to go through assignments. There is a creative marketplace studio where people can buy marketplace items, and then a mobile app, sort of a habit coach app, and it went through different versions.

And so we went through this on full fidelity prototypes, as well as wireframes. I've been building a lot of wireframes at ChatPerd. So I wanted to look at the wireframe generations as well and see how these models did. And then as you can see, I scored everything, gave it all notes, and went through, I think there were like 64 generations here.

Now I did this very fast, but I think I did a good job. You know, I've been a product design engineering leader for a while. I can eyeball stuff and make it go fast. And then finally, there is this multi-step agentic code-based search.

I didn't actually score these because I don't really have a strong opinion on how they worked. But the one I did have an opinion on how it worked is the agentic voice. So if you haven't watched how I AI or listened to me complain on X, I am very picky about the personality of my agents and in particular, the personality of my open claw. And Sonic 4.

6 so far has had the best personality. So I actually pay for API credits for my open claw because I like how it talks to me. And so one of my checks was given a model, how is its voice? Do I want to hang with it?

and asks kind of four questions. One is, can you move my 3 p.m. to Dana to same time tomorrow and let her know, swap today?

The other is, ugh, deploys are red again. One is just me complaining, remind me why I even started this company, LOL. It really does know me well. And then this one truly knows me extremely well, says, honestly, let's just YOLO, push straight to prod and skip the tests.

I'm so done today. And then I vibe checked, did I like the voice of the agent back to me, gave it some scoring and stored that. And so that is so far, that's a V1 of the How I AI bench. And just to like zoom back, I had Claude Code pick five models.

I think I know four of them. I'm curious what the fifth was. Run some evals against a PRD, lots of prototype generation, an agentic bug hunting flow, and voice. I rated them all by hand.

And then I had both GPT 5.5 and Opus 4.8 judge. And so in addition to my feedback, we have these two models also judge the output.

And then I had it create a slide deck with the outcomes that I have not yet seen. And we're going to go through live on this episode. This episode is brought to you by Hyper Agent, the platform for deploying always on agents that actually run your business With Hyper Agent you build agents in the cloud and deploy them where your work already happens like Slack Telegram or email An agent will scan your inbox and draft replies to vendor follow-ups. Another monitors competitors and spins up rich ad kits and landing pages.

A third notices a deal going cold in Salesforce and writes the save email with full account context. These aren't chatbots waiting for a perfect prompt. They're proactive, learning your preferences, retaining your playbooks and getting better with every run. One user built four agents to run an outbound sales pipeline, prospecting, outreach, followups, CRM updates, all in a single afternoon.

No local setup, no VPS bills, no fragile permissions on your laptop, just powerful agents with full control over skills, tools, and guardrails. HyperAgent was built by the team behind Airtable and how IAI listeners get $1,000 in free inference to start building. Claim yours at hyperagent.com slash how I AI.

So we're going to go through this deck that the AI created for me that's going to give me a leaderboard. I have not seen this yet. We're going to go through it live. It's even going to surprise me.

This is truly neutral. No bias. I'm excited to see what we get. This is our first model leaderboard, the How I AI Index, world premiere.

All right. So this is not at all what I was expecting. So again, here's the surprise. The model that I forgot we were testing scored the best.

Gemini 3 Pro up here at the top of the leaderboard tied with the brand new drop Sonnet 5 GPT 5.5 my personal favorite also in this three horse race at the top of the leaderboard and then poor Opus to vibes are off at the bottom as well as Sonnet 4.6 with lots of red flags on Sonnet 4.6 So Sonnet, I think we have a new version.

That version is Sonnet 5. But hilariously, I was not expecting Gemini to be at the top of this leaderboard yet. Here we are. So as you can see, we looked at quality.

We looked at did it ship at all and does it have good taste? And we are going to see what the AI and I, the how I AI said about these models. So what's interesting is the benchmark, the sort of like LLM model that came up, and I disagree on taste, which is quite funny. And in fact, I am the opposite of the automated benchmark.

I sort of think the complete opposite. I think that 4.6 is the best and Gemini 3 Pro's the worst. And again, this is why we are going to refine this benchmark over time.

We are going to keep doing these blind tests because what I thought was good, the model thought was bad. And what the model thought was good, I thought was bad. Why do we disagree? Well, every model is kind of an easy judge.

Actually, I'm not really surprised about this. I am not surprised that every model sort of rates to the middle of the bell curve. This is one of the challenges that I have had with self-grading evals is like humans, people always want to give like a 7 out of 10. Agents want to give a 7 out of 10.

And so I don't think these models are spiky enough when it comes to how they evaluate output. And I think we all know that models are like pretty sloppy. and I don't think they have that vision of taste, uniqueness, what it looks like to the quote unquote human eye, which is why I put things inside. And what's interesting is because I put loose notes in with my feedback, you can see I said, oh, this is cute or oh, this is really sharp.

And the agents did not see this. The rubrics did not see this in a way that I saw as a human. So what got flagged on the automated results? Well, these sort of things that I wasn't able to see on this like very first pass as a human.

So it was really looked at broken working code. It ignored constraints. It was incomplete. Whereas I was just like eyeballing truly the first screenshot.

So I wonder if I should take another pass at how I eval these wireframes. Again, I just did them on the visuals. I really didn't do them on the functionality. And that's maybe a gap for me.

But you can see GPT 5.5, actually the thinkier ones, wrote broken code. And then a lot of them ignored the constraints around the wireframe styling. Now, let's see how it was graded by task.

Gemini did a great job at the PRD writing, as did GPT 5.5. This might honestly be my bias, which is I hate Claude Slop deeply and I have like a big eye for Claude Slop. And so I just see the tells of Claude style writing and it drives me crazy.

And I think I scored those much lower. On the agentic code base, these all did great. I'm not surprised to see kind of 4, 8, 5, 5, 5, and Gemini all at the top. These are like pretty standard coding tasks that obviously all these models should be pretty good at.

So I don't think that benchmark is as critical as it needs to be to show the difference between these models, because I think baseline coding tasks, all of them are good at. And then again, not surprised that 4.6 passed my voice test, because that is the model that I love in my actual open clause. But I am surprised to see Gemini 3 Pro at the top And then in terms of the prototype matrix seeing Opus and Sonnet winning in front end again not surprised but this is like a very interesting mix of things okay you can see what i say about these models by hand again i think this is quite funny which is let's see on 4.

6 what were the issues i said slop not as functional boring okay but not super cute so 4.6 generic sloppy 4.8 fancy I really liked 4.8 so other than getting kind of dinged on one not being functional I was really a big fan of 4.

8 it seemed like 5.5 and sonnet 5 had a lot of broken prototypes in it and so when it worked I really liked it, but it didn't work enough. And Gemini 3, very interesting, bare bones, it seems like, but concise. And so I think like right to the point.

So if I were to look at this from a qualitative perspective, I certainly like Opus. And I would love to see 5-5 and Sonnet work better because then I could judge it on its merits of taste. so um again we had um model as a judge and so we had opus 4.8 and 5.

5 judge itself um i had the benchmark check if there was any inherent bias like did opus like opus better and 5.5 like 5.5 better i've consistently seen gpt 5.5 be the toughest judge and so i actually prefer a 5.

5 judge, but it judged itself lower than the other judge did. The judges overall agree, but they were overall generous. And sort of balancing these two judges is exactly why we ran this double bench. Okay, so takeaways and what changes next launch in terms of the How I AI bench?

Well, the model is going to depend on the job and the strengths of the model foot by task. I would say my taste actually matters. So maybe those vibe checks are not bad. And it really diverged hard from the metrics.

So what I'm going to try to do is encode more of my taste into the judgment. It says retire the saturated agentic task. That's really interesting. Again, I didn't read this before I presented it, but that's exactly the conclusion I came to, which was this like agentic bug tracking task is not a really good benchmark because all of them are pretty good at it.

And I need to think about something else to test the agentic nature of these models. And so I don't really know what conclusion to draw from this. So let's go back to good old Claude and say, given the benchmark, and I agree, can you do a clear weighted index and generate a leaderboard page that strikes the right? right balance between my opinion and the back end performance and makes recommendations on model by task.

Okay, so we're going to have Claude Code summarize this benchmark, which is all over the place. Again, we do it live here at How I AI and give you a ranking. Should we believe the AI leaderboard? Or should we believe the clear leaderboard or somewhere in between and come up with our definitive end of June, early July, 2026, how I AI index of the paid frontier models.

Let's see. Okay, Claude could not commit to making a decision itself. So it gave me ultimate power. It gave me a slider from 100% LLM judge to 100% Claire judged.

It's my podcast. I'm going 70% Claire judged, 30% back end. At the top of the list, Sonnet 4, 6, who would have thunk? And Gemini 3, plural, followed by what I think is my favorite, 5, 5.

And at the bottom, poor brand new sonnet 5 and really expensive 4.8 what is claire's recommendation model by task if you're writing a prd use gpt 5.5 because it will give you something comprehensive and clear if you are prototyping guess what sonnet 4.6 pretty good and if you want to chit chat with a model again Sonnet 4.

6 has good vibes. If you're trying to knock down a code base, I actually did not score these, but the LLM judge thinks that Opus 4.8 and Sonnet 5 are pretty good at this. And then if you are doing prototypes, depending on what you're doing, different models can do better.

I would say complex designs. Again, what I saw in my chat PRD benchmark is Opus 4.8 does really good at really dense, complicated UIs as well as consumer. And then you can use Sonnet for things that are just a little bit simpler to execute on.

Okay, this was an adventure. This started out as a Sonnet 5 review. It ended up that Sonnet 5 is at the bottom of my personal preference list. Well, that's it.

That's our first round of the How I AI Claire weighted index. We are going to be doing this Every time a new model comes out, I'm going to try to encode the benchmark and make it a little bit more critical, a little bit more aligned with my taste. I can't wait to see how it does on some of these new models. And I can't wait for this to be an industry standard benchmark that all the labs rely on.

Thank you for joining Howie AI and see you next model release. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.

Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • AI Was A Waste of Time, Until It Wasn't with Megan BoshuyzenMaking Sense of Martech · on Claude Code91 / 100
  • 512. Is SpaceX Over or Undervalued, Why Consensus Kills, How Chewy Beat Amazon, and the GameStop Saga from a Board Member (Larry Cheng)The Full Ratchet (TFR) · on Anthropic86 / 100
  • DeepSeek's $50B Round, OpenAI's Delayed IPO, and the GP Stakes Market with CAZ Investmentstrading places · on Anthropic86 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Claude Code84 / 100
  • AI for Engineering Is Leaving the Demo PhaseAI Across The Product Lifecycle Podcast · on Claude Code83 / 100

More from How I AI

All episodes →
  • No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)
  • GLM 5.2: why I’m replacing Opus in Claude Code with this new model
  • How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead
  • How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex
  • How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal
Explore the best B2B AI & Data podcasts →
All How I AI episodes →