The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/The Programming Podcast
The Programming Podcast artwork

Opus 5 vs The Hype Cycle | Unbiased Breakdown

The Programming Podcast · 2026-07-31 · 46 min

0:00--:--

Key moments - from our scoring

Substance score

48 / 100

Five dimensions, 20 points each

Insight Density12 / 20
Originality11 / 20
Guest Caliber6 / 20
Specificity & Evidence10 / 20
Conversational Craft9 / 20

This episode unpacks Claude 3.5 Opus's release, moving beyond marketing benchmarks to examine practical performance issues that official documentation glossed over. The hosts tested Opus 5 extensively and discovered that while it achieves better scores on agentic coding and knowledge work benchmarks compared to earlier models, real-world usability suffers from problems like ignoring explicit instructions to keep responses brief, treating directive prompts as suggestions rather than requirements, and requiring deletion of carefully-built Claude skills before the model functions as expected. They critique Anthropic's context engineering overhaul - which removed 80% of Claude's system prompt - for being buried in a blog post image rather than communicated prominently, leaving developers frustrated when workflows suddenly break. The hosts discuss the tension between model improvements and developer experience, referencing Khun Chen's observation that frontier models now prioritize machine-verifiable outputs over human-friendliness, and debate whether to use Opus 5 at scale (cost advantage) versus sticking with Opus 4 (predictability). Core themes include the gap between benchmark metrics and user experience, enterprise adoption risks, and the need for better developer communication during model transitions.

Key takeaways

  • →Claude 3.5 Opus outperforms previous models on agentic coding and knowledge work benchmarks but underperforms on legal, health, and humanities reasoning, revealing that Anthropic prioritizes coding capability improvements over balanced performance.
  • →Deleting saved Claude skills in Claude.dev resolves major usability issues with Opus 5, but requiring this workaround without prior documentation creates unnecessary frustration and breaks established workflows - a critical enterprise adoption problem.
  • →Anthropic's context engineering changes (80% system prompt removal) were relegated to a blog post image, leaving developers unaware that fundamental prompt structures need rewriting, causing day-long debugging instead of informed model transitions.
  • →Benchmark performance improvements don't translate to real-world gains when models ignore explicit formatting instructions (like keeping responses under 600 characters), creating friction that may force enterprises to stick with cheaper legacy models despite cost savings.
  • →New frontier models increasingly prioritize machine-verifiable outputs over human usability, causing developers to use more jargon, need extensive steering, and experience reduced creative control - shifting development toward what's convenient for AI rather than what's convenient for humans.

Guests

Joel HooksDev AgrawalKhun Chen

Topics in this episode

OpenAIAnthropicPrompt engineeringGrokAgentic codingContext engineeringReinforcement learning from human feedback (RLHF)Claude 3.5 OpusClaude.dev skillsbenchmark testing

Questions this episode answers

Why does Claude 3.5 Opus ignore my instructions to keep responses short, even after I specify 600 characters max?

The model was likely built with different context engineering assumptions due to Anthropic's removal of 80% of Claude's system prompt. Deleting existing skills and rebuilding them fresh with the new model generation resolves this, but Anthropic did not communicate this change clearly upfront.

Is Claude 3.5 Opus better than Claude 3 Opus based on benchmarks?

On benchmarks, Opus 5 outperforms Opus 4 on agentic coding, knowledge work, and novel problem-solving tasks, but underperforms on legal, health, and humanities reasoning, showing Anthropic's focus on agentic capabilities rather than balanced improvement.

Should I switch to Claude 3.5 Opus immediately for customer-facing applications?

No - despite the half-cost advantage, the hosts recommend waiting until usability issues are resolved and thoroughly testing with your specific workflows first, since poor model behavior directly impacts customer experience and customer acquisition cost.

What changed in Anthropic's context engineering approach with Claude 3.5?

Anthropic removed over 80% of Claude's system prompt and shifted from rule-based prompting (give Claude rules, examples, interfaces upfront) to progressive disclosure, simple tool descriptions, and auto-memory - changes documented only in a blog post image, not prominently communicated.

Why do modern AI models feel harder to work with even though they're more capable?

Frontier models are increasingly optimized for machine-verifiable outputs rather than human usability, resulting in more jargon, more required steering, and less predictable behavior - a shift toward building systems friendly for AI rather than for developers.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

12 / 20

The episode contains useful practitioner insights about LLM workflow optimization and model selection strategy, but is undermined by significant filler including lengthy tangents about phone stands, microphone equipment, and personal setups that have minimal substance. The core insights - about prompt engineering friction, skill compatibility with new models, and model selection for specific tasks - are valuable but scattered and diluted.

I often say, hey, your responses must be bullet form, very short, less than 600 characters. Right. Because what is the point of me reading 10 minutes worth of stuff to realize that this is not what I want?
how pleasant is it to work with the model used to be a strength in Claude, but now it's not

Originality

11 / 20

The hosts offer some contrarian takes - skepticism toward Opus 5 despite hype, the 'value maxing vs token maxing' framework, and the insight that new model releases can break existing workflows - but these are incremental observations rather than novel frameworks. Most recommendations (testing with existing projects, prompt engineering discipline, matching models to tasks) are well-established practitioner knowledge. The conversation largely recycles familiar tensions in the AI community.

almost looks like AI is directing humans to build a world that's more friendly for machines rather than humans
they have to scramble to fix this. And so I legitimately don't understand how you release this and not understand that everyone's going to try and run to reduce their cost

Guest Caliber

6 / 20

The episode features only two speakers with no external guests, making guest caliber assessment inapplicable in the traditional sense. The hosts appear to be practitioners working with AI systems, but their specific credentials, companies, or scale of operation are never clearly established. One listener question is addressed, but that is not a guest appearance. This is a significant structural weakness for a podcast claiming to deliver B2B substance.

Developer speaking from personal experience with AI models
I'm going to actually let you in on a very, very important thing that it took me a little while to figure out

Specificity & Evidence

10 / 20

The episode includes some specific examples (skill deletion fixing Opus 5, a CTO finding 96% of his company using Opus 4.8, chessboard rendering demonstrations) but relies heavily on abstract claims about model behavior without data. Pricing comparisons ($50 per million tokens, $700 for 16-hour runs) are mentioned but lack context. Many assertions about LLM capabilities and user experience problems are offered as anecdote without systematic evidence or metrics.

he looked and he said 96% of our usage is Opus 4.8
let fable run for 16 hours and it cost me $700

Conversational Craft

9 / 20

The hosts demonstrate genuine disagreement (one prefers Fable 5, the other experiments with multiple models) and follow up on each other's points, but rarely push back on claims with hard questions. There is minimal fact-checking; benchmarking assertions go largely unchallenged. The conversation meanders frequently into equipment discussions and tangential topics, suggesting weak editorial discipline. Follow-ups are collegial rather than investigative, missing opportunities to stress-test claims.

Yeah. So if we're looking at the benchmarks
And the funny thing is Anthropic knew this. Like they, they knew it, knew it because they put out a blog post

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker C61%
  • Speaker D31%
  • Speaker B3%
  • Speaker A3%
  • Speaker E2%

Most-used words

models25model22opus20back20open18fable17chrome16sense16episode15skills15better14experience14cost13code13accenture12less12

Episode notes

Is Opus 5 worth the hype? I spent two days testing the 2026 release to see if the performance lives up to expectations.This Opus 5 review provides an honest look at the latest model, moving past marketing claims to examine real-world usability. If you are deciding whether to upgrade or switch workflows, this analysis breaks down the actual capability of the software compared to previous versions.I cover the specific Opus 5 benchmarks I recorded during my testing phase, highlighting where it succeeds and where it falls short. You will see my candid Opus 5 impressions after two days of heavy use, helping you determine if this model fits your specific needs or if you should wait for further updates.Subscribe for weekly software breakdowns, and let me know in the comments if you want to see a head-to-head comparison between Opus 5 and other top-tier models.

Full transcript

46 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales using automation, analytics and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more@accenture.com Spotify

Speaker B: this episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50 page restoration block. Or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses. Setup required. Compatibility and availability various.

Speaker C: 18 baby, take care of the kids. Opus 5 just drop. Uh, as of the time of the recording this, we have taken two days to test and implement and do Things with Opus 5 to give you number one exact breakdown into what it is that you need to be knowing about this. We have the benchmarks obviously, for those that haven't been keeping up with this. So that way you're in the know of what that looks like. But I have some mixed reviews and I don't know if that's the thing that the hype cycles want to hear, but these are going to be the raw, real ones. This is going to be an interesting episode. This is not the episode I came in here thinking that we were going to record. Let's say it like that.

Speaker D: No, not for. Not at all.

Speaker C: Yeah. And if you're new and you're here, we are back with another one.

Speaker D: Another one.

Speaker C: Another.

Speaker D: Hot off the presses. Opus 5, Frontier Intelligence. half the cost.

Speaker C: It is half the cost.

Speaker D: Have we made it?

Speaker C: AGI is here. Okay, let's look at, uh, let's, let's look at this, uh, realistically. So I was very excited for Opus 5. I actually thought it was going to drop on Thursday, but it dropped on Friday. I think there's some very, very good things here, if I'm being honest. Well, let's go over the benchmarks before I cloud some of this with like, my experience. My experience on this has been very interesting. Also, if you are having a mixed experience with Opus 5 that's similar to what I did, I'm going to actually let you in on a very, very important thing that it took me a little while to figure out, which I didn't know. And it did change the experience quite a bit. But it has not been the overall top experience that I thought it was going to be. But again this is two days in. Maybe they're going to tweak some things later on. But right now it's not, it's not there.

Speaker D: Yeah. So if we're looking at the benchmarks, uh, that came out. And a quick note before we look at benchmarks. You know what, I'll save my note for after the benchmarks. But there's something I want to say. It's really important about benchmarks. So Opus 5 is performing better than Fable 5, the mythos variant, uh, in all but four of their actual benchmarks that they used to come out with this announcement. And so Agenic terminal coding improved by about 10%. Knowledge work is up. Novel problem solving. This is a newer benchmark so not comparable yet but there Agenic search up computer, uh, use up business workflows, up biology and uh, kind of biology knowledge work, uh, performing better. Some of the places where it is not performing as well as Fable 5 according to the benchmarks. Multidisciplinary reasoning. Uh, so there's this test called Humanities Last exam it performed slightly worse by like 0.2%. We're talking about here. Um, there is this frontier code Agenic coding where it was like 0.2% off again. And then uh, legal and health where Fable 5 edged it out um, by a significant a bit. So it seems like the focus and this is something that I felt and I think a lot of other folks have felt around OpenAI and anthropic strategy is it's very clear what they care about when it comes to model improvement. And that's code, right? Agenic coding coding in a general, when new benchmarks come out they want to be top dog in those areas. And a lot of the knowledge work stuff has kind of always trailed behind. It was nice to have, but not the thing they're shooting for. And this continues right with the Opus 5. They kind of said hey, we don't care so much about the legal health, um, but the agency stuff, you're getting slightly better performance uh, at half the cost. Now whether or not that's true, uh, is in the eye of the beholder. But if we're going to talk about benchmarks, I just want to say this one thing. I never agree with them. This m. My day to day experience has never pegged to any of these benchmarks. I've never looked at a benchmark and go, yep, that makes sense. And it just, it just almost never happens. And this is across any provider, any model. And so I give these like the, the, it's like a, it's like a half a vibe check for me. But it's not something that like, influences my decisions day to day. I have my own internal benchmarks of things I look at that help me determine whether or not these are actually useful to my workflows. Um, because I just, I don't understand how these providers are so mismatched with the day to day use of people that are using these models.

Speaker C: Yeah, like, I've never used any of the previous models. Like, oh, there's that 4% differential they were talking about. Like, I've never felt that. Yeah, for me, I go off of, uh, several things. Number one, I will give several models and several versions the same exact tax a task. And I want to see how, um, they perform, what the output looks like, things like that, the code quality. Yes. I'm still one of those weirdos that look, look at the actual code. We'll talk about this time and time again. Um, I will say, like, there are certain areas where I'm starting to become a little bit more comfortable looking at less of that code, but anything that's like customer facing or customer engaging, I'm always 10, 10% going to be looking at that code. And so I don't care what the YouTubers of the world say on that one. Um, maybe that's personal feeling and I feel better when I actually see the code. But also, I don't want my skills to atrophy. I'm sorry. Like, I want to actually have and maintain my abilities that I've built up over the years. It took me a lot of work to build this up and I don't really have a desire to get rid of it. So to each their own in that one. Apparently this is a hot take to some. Now, with that being said, my experience on Opus 5 has been less than amazing, to be honest with you. And I have shared this in a post. With that being said, apparently this is a hot take.

Speaker A: I.

Speaker C: Apparently, if you aren't totally, uh, in love with Opus 5, you're doing something wrong. Well, I wasn't doing something wrong. Um, finally, after some experimentation, um, I posted publicly about it, like, hey, this is what I did. I did not appreciate this. I did not like it, it did not go well. Um, and I had three people literally text me because it's almost as if, like, they're afraid to openly say this on social media, but they Also did not have good experiences with Opus 5. And I was like, very interesting. So part of the experiences that I did not have that were favorable when I gave it similar tasks, it performed worse on some of the tasks. And then when I gave it certain tasks, it almost was like it was talking more than it was actually doing things, which really confused me. And so, um, I don't know if we've talked about this on the podcast, but I talk about this a lot with people. And for me, one of my strategies with AI is. And let me explain it with this example.

Speaker B: Sure.

Speaker C: AI, especially reasoning models, have the tendency to basically print out a blog post every single time you ask it to do something. The output is just so long, so over the top. And a lot of times it ain't really talking about nothing. And so I often say, hey, your responses must be bullet form, very short, less than 600 characters. Right. Because what is the point of me reading 10 minutes worth of stuff to realize that this is not what I want? Goes again, another 10 minutes. That's not what I want. 10 minutes. We're getting closer. And I spent 60 minutes of reading to finally get to the air. Where was the productivity gain at that point? It wasn't there. So instead I say, summarize your thoughts, put it in bullet points, keep it short. I can glance at it basically like a Tweet or a LinkedIn post and then be like, ah, ah, this is not right. Iterate. This is exactly where I want us to be. Cool. Now we can expand upon that because I now know this is worth the 10 minutes of me reading. Now make your whole thesis.

Speaker D: Yeah.

Speaker C: So instead of spending 10 minutes to find out if it's right or wrong, it's a 30 second glance. No, this is wrong. Fix this, Fix this now. Run with it. That's well worth the time. It was disregarding that every single time almost. Um, I kept saying over and over again, hey, this shouldn't be happening. We shouldn't see this output being this long. Shorten it. So that way I can. We can summarize this. And I say, completely understand. I will not disregard this again. And the very next prompt, it disregarded it again. I was like, whoa, why is this happening? Someone who's somebody I really respect told me, hey, you probably need to delete all of your skills. And I was like, what are you talking about? Number one, I spent a lot of time making like, I've got these skills honed in right now. Once I deleted all those skills out of my claw code desktop, it Worked just fine. It took me two days of just pure frustration before. So, uh, like, the thought even came to delete the skills, and that totally changed up the experience. Now it is a little bit better. And I want to shout out one person in particular. Two, uh, people in particular, actually both people that I respect. Both people speaking at cyc, commit your code. One is Joel Hooks and the other is Dev Agrawal. Dev has actually been on the podcast. We might need to get Joel on here.

Speaker D: Joel, he's doing some fun stuff.

Speaker C: We're going to bring you on the podcast. I'm going to. I'm a DM you after this. Uh, I'm definitely going to DM Joel after this. I'm sure he'll love that. And so I basically pointed out, like, hey, update to this. I had to delete my skills and things that I use regularly, and it started working better. Why do I need to delete skills to make my experience better? If I need to delete skills anytime a model is upgraded, that is not productivity. And especially now, as I'm thinking about, like, enterprise level, you're telling me I need to remove my skills in order to get them. Like, what are we talking about? And he said, it's not ideal, for sure. But also models improve. We have to consider how we reason about the work, the bitter lesson skills, et cetera, are stopgap polyfills to hold us over until the models basically get to where they need to be. I said, sure, but why do we end up in situations where days of frustration have to occur to finally realize this may be the solution versus them sharing this with us beforehand? And we went back and forth a little bit and Dev came through and was like, I think what needs to happen here is that new models need to somehow be onboarded into existing workflows and skills. Basically, have the new models write their own smaller skills using the existing ones as a reference point? I've never thought about that in that lens. I don't necessarily know even, like, again, this is a Twitter thread. We're not here curing cancer, but, like, I don't know if that's the exact right approach, but I love the angle of this because instead of maybe onboarding, what if it took the existing skill and rewrote it in a way that its current model can utilize it properly versus you having to just be aggravated beyond belief?

Speaker D: Yeah.

Speaker A: This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales using automation, analytics and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more@accenture.com Spotify

Speaker B: this episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50 page restoration block. Or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses, setup required compatibility and availability various 18 plus

Speaker D: it's funny because I have the AI write my skills a lot of the times anyway, right? Like most of the time. And so it makes sense that I could do this with any new model release. Um, yeah, I think that makes a lot of sense. I think that's probably a path forward for right now. It's sad that we have to figure this out each time. It's like, think about the user experience that folks are gonna go through. Like, and the funny thing is Anthropic knew this. Like they, they knew it, knew it because they put out a blog post on 24 July as well about the new rules for context engineering with Claude and the Claude 5 generation of models. And the highlight of this, I know it's a little different, but the highlight of the top of the article is we removed over 80% of Claude code system prompt for more advanced models. How to apply the lessons we learned to your own context engineering in cloud code with your own agents. So they knew that stuff was changing about the things you put into how you utilize these models, but didn't make this more widely known. There should be something that pops up, something that tells you, hey, we, we fundamentally changed how you're going to use these models, but we just put it in a blog post and hope you find it. Like what that's going to break so many folks workflow and I spent 24 hours just not wanting to use these models at all. You know.

Speaker C: Yeah, I think people forget the user experience part of this because like again, I always go back to the same thing. If you've been listening to this podcast for a long time, you know, I often think in the enterprise use cases, uh, Leon has been like my sense of reality for the startup side of things. But I always go back to like when you have a hundred employees frustrated because they've built skills that not only that they built, but they're sharing across the org, right? And now you have a new model that comes in and they're all going to work on Monday. Like what the hell's going on here? That's not a great situation to be in.

Speaker D: Like, like listen, listen to this. This is, this is an image in a blog post that they randomly put out. It's just, it's so wild to me that this is not. And this is why we need Devrel. This is why we need the community, why we need people like actually telling us what the heck's going on. All right. They have a then and now image that says give Claude rules. Now it's give Claude judgment, give Claude examples, design interfaces, put it all up front versus use, progressive disclosure, repeat yourself, simple tool descriptions, memory and claude mds, auto memory, simple specs, rich references. That's everything that changed and what we should be doing. That is not even in text. It's a file freaking image in a blog post. What, what is going on here?

Speaker C: Some of the things that coming out like, you know what's funny is I've been doing graph related, um, building and agent use cases for a little while now. Now it's starting to become like, I guess the trend in AI, which is cool and maybe I need to talk about this more to where like people are, understand where it's going. But it goes back to my big thesis on all this and that's so much has changed that people don't even know how to approach this anymore and that folks that weren't necessarily AI, uh, enthusiastic earlier on really haven't missed much because now it feels like everybody's learning everything all over again from scratch. Because it's a new paradigm to kind of approach this with. Um, I will say if you're still approaching it in the last way, you will still get the results that you need to get. But you can definitely find these companies are pushing a new agenda. With that being said, you can still have favorable results without drinking all the Kool Aid that you need to drink. I personally think my again, I think there are people that are like totally on, I'm gonna drink all the Kool Aid in the world. I like to come from a very realistic standpoint. Right? Because when I think of AI, ah, I think of like systems that need to build over time, agentic use cases. If you ask a model 500 prompts, the exact same prompt every single time, you're gonna get two to 300 different answers because they're not going to replicate, like, because of that. Unfortunately, I like to have not just guardrails, but scalable approaches towards things. Like, if I have a thousand customers that are going to use something, there's no way they should be getting a thousand different responses. Like, it just doesn't make sense to me. Uh, you can't control that output. Like, that's why, like, I have a fear with, uh, customer chatbots that I don't think people think about all the time. It's like, unless you're giving it canned responses that it can like go back into in a bank. What happens when the AI just says the wrong thing? Oh, you lost a customer. Well, people don't realize that there's a thing in business called cat customer acquisition cost. How much did it cost you to even acquire that customer? How many times now? If you scale this across a million people, how many do you end up losing? Like, oops, we were experimenting. Like, experimentation comes at a cost.

Speaker D: And the workflow to do all that just changed. And they didn't tell anybody. Like, that all had to be in tool descriptions. Now is now the way to do that. And if you didn't look at this one image in one blog post and you're trying to use these models, you don't know, why is it breaking? I have no idea. And that's the super frustrating part to me because of course we have to do all this stuff to build reliable software. And so there's this super spicy take that I kind of want to read verbatim because it's the only thing that makes sense to me right now. So this is, um, from Khun chen, uh, fire YouTube, by the way. Just one of the few people that like share their entire workflow beginning to end. Um, they were in LA at a bunch of high tier fangs, um, and now they kind of make YouTube videos about how they're using AI. He has this tweet about why he doesn't like, uh, Opus 5. And at the end is something that I think is really interesting here, uh, because it reflects a lot of how I've been feeling too. It's the third and last point. He says, how pleasant is it to work with the model used to be a strength in Claude, but now it's not. Honestly, Grok is my favorite right now in the pleasant dimension. Kimmy is not bad either. It feels like both anthropic and OpenAI AI are giving reinforcement learning HF less care in favor of scalable RL and that's machine verifiable. This is almost looks like AI is directing humans to build a world that's more friendly for machines rather than humans. And most humans don't even realize that they are being manipulated to help with that. Almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with. If this continues, AI will start to speak their own language that looks like English but average humans can't understand. They will choose to do things that the human never asked for. Are we already failing at a line?

Speaker C: I'm already experiencing that. I feel like I'm experiencing that all

Speaker D: the time now in my bones. It feels like I am getting more and more friction. And when I started using five, that's what I kept running into. This is not a model that's letting me be creative to let me be an engineer. It just kept getting in my way. And you start reading all these snippets like, oh, remove this, remove that, and it works better. Okay, okay, I get that. But I went from using Fable 5, which for me, I know there's a cost to it, but doesn't get in my way. It's very freeing model to use. I can have it go off and do these large, complex tasks and it comes back with the things I need specifically because my cloth setup is so built out. And I think it respects it for the most part, but it's just, um, it's a low friction model for my day to day to where I'm not having to babysit it as much and I get a little bit more predictable responses out of it, to where even at half the cost, I'm still reaching for Fable 5. It's just, I can't, I can't have that friction in my workflow. And I've just stopped by for now until I can have more time to sit down, build better systems for it. Figure out reading this one damn blog post that no one's talking about. Until I can figure all that stuff out, I literally can't not have my workflow work against me. It just doesn't make sense.

Speaker C: Yeah, you know, here's the thing. Um, I know it seems like we're really like raining on Opus 5. I feel like we are, but this doesn't mean, like, don't use AI in some way, shape or form, or like we're just anti AI. It's not the case. But it goes back to the same standpoint that I've had for a long time when I build out real strong workflows or pipelines or bigger applications. I have a multi step pipeline that's like maybe 13, 14 steps, maybe three of them are agentic, where the rest are like true structure. And I feel like people are forgetting this part, like they're beginning and ending everything with AI. And the problem is when you do that, you end up not having favorable results. Now here's the thing I've never been the advocate of. You need to switch to the latest model before you know the, what the output is. And even Joel Hook said like, um, you don't have to run to a model on the first day that it comes out. And absolutely, I would never put a brand new model first day out in front of customers unless I made like a coding editor of some sort. Right? I want to control all that. That's why like you find people using legacy versions of software because they're like, hey, I know what these outputs are going to be. And if it's, especially if it's customer facing, I need to make sure my customers get an experience that they need out of this. 100%. Why did we jump on it the first day? Y' all need to know about it. And so it's like, hey, we're over here. We're not just doing this for the sake of doing it, we're doing it because we have a listener base that needs to know about it. We have people that are making decisions on it. So if you go to work on Monday, I probably would not be putting, uh, Opus 5 and anything customer facing right now. There's going to be some tweaks that they need to fix on this one.

Speaker D: But it's also half the cost, right? If any enterprise can cut their cost in half with this stuff, if anybody's using this significantly at scale, right? And there's a magic button, you can press the half your cost, they have to scramble to fix this. And so I legitimately don't understand how you release this and not understand that everyone's going to try and run to reduce their cost and have. Are we playing some weird mind games here? Anyone that knows they want to save is going to try and use this and then have a really poor experience. What's the end game with this? Um, um, so I think that, that, that that kind of like blows my mind.

Speaker C: You know what's funny is, and we probably should do an episode on this, but I was talking to a CTO of a company recently, we were on a call and he, we, we sometime somehow got on the subject on my thesis on like go to the right person for the right job. And, uh, for those that don't know, I have this thesis on you should be toggling models even mid conversations, so that way you're getting exactly what you need out of it. Uh, just because you started with Opus 4.8, which was the model in question at the time of the conversation, or in this scenario, maybe Opus 5 doesn't mean you have to be married to Opus 5 until the end of the conversation. Like, you could toggle it. And so he decided, he's like, I bet you most people in my company don't even know this. He has like a 400 person company. And he looked and he said 96% of our usage is Opus 4.8. And he's like, that actually now explains so much to me. And I'm like, yeah, every single time someone needs to add padding to a button, they have Opus 4.8 doing it. And when they get frustrated, why it's not the most, like, ideal outcome every single time. It's because when you use a high level, complex reasoning model for a lot of this bigger thought. Haven't you ever had a situation like maybe you got a message at work like, hey, we need to talk tomorrow. Your mind's going to 10,000 scenarios. Am I getting fired? Do I need to bring a box to pack all my stuff? Like, is this the last day? Oh, did I mess this up? Did I? And then it's like, oh, hey, like, we have this new task that we got to do la da da da da. And nothing was wrong. You overthought it. Well, when a reasoning model is getting something that's so bare bones basic, it's probably thinking, like, there's probably more to this. There's probably more that I need to architect there I got to build. Or what about the future? It's like, nah, bro, I just need padding. It's like, why did you need to make a blog post to tell me that you added padding to this thing successfully? You're burning tokens or your usage is just not ideal for what it is. And so there's no reasonable way that ever say, Opus 4.8 for every single task, every single thing that you do, La da da da. Um, it just doesn't make sense. It's going to overthink a lot of the tasks.

Speaker D: You're not using Fable 5 Max to push the GitHub. What are you doing?

Speaker C: I like Fable. I think Fable's cool. It does. It definitely works. Well, there's no doubt about that one. But I Can't get to the stage where it makes sense to spend 50, $50 per million. I don't know uh, what anybody is working on, especially in the enterprise org that justifies that being something that you're touching on a regular basis. It just doesn't make sense to me now if I was working on a big feature, I need a PRD for everything. I need to execute this. Let me use Fable, let me map out all the edge cases. Let me think through this feature. Okay, sure, spend your 50 bucks there. But there's no way I'm spending 10x that to execute everything. What do you mean that I've let fable run for 16 hours and it cost me $700. Like there's no way.

Speaker D: Yeah, um, it's funny because I, I'm, I'm a Fable five maxing right now for sure, uh, because it is less friction in my workflow. But I'm using it for a lot of knowledge work stuff like that. That's been the fun part for me. Uh, sold for a lot of my day to day of course, but for a lot of like my thought partner stuff it's been Fable 5. So I think that's pretty interesting that I have this model now that with all the things I have built into it to help me is capable and then is also capable of this high level agentic stuff. And I haven't really felt like that in a while. I've always used kind of GPT based models to do my knowledge work and this is the first time I'm reaching for a CLAUDE model to do any of that stuff, which I know that's probably the inverse of a lot of people, but I just never felt comfortable until now with this. And so it's kind of my tool for everything until I'm actually doing lower level building and then I'm using sol.

Speaker A: This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales using automation, analytics and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms, um, that matter most. Learn more@accenture.com Spotify

Speaker B: this episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50 page restoration block or finally break down that long article You've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses set up required compatibility and availability various 18 plus.

Speaker C: It's funny there. I. I've. I've had this thesis, I didn't have a name around it and I just saw, um, OpenAI. They did a live stream yesterday and I didn't watch the live stream, but I saw the title and it was Token Maxing versus value maxing. And there was a stage. I've never was on the bandwagon of token maxing. I think it was foolish. I think any company that had a leaderboard that was just watching tokens was just opening themselves up to just burn them in the most negative way possible. Now they're like, no, you need to be value maxing it. How are you getting the most amount of value or ROI for the tokens that you're spending? This has been my take for the last two years. Forget like the last six months of people abusing it. The last two years I'm like, if I'm spending $10, if I'm only getting $10 worth of value back, I messed up.

Speaker D: Yeah.

Speaker C: So it's like, where do I continue to quantify or increase that value? So for me, no. Hate to the folks that are just using Fable all the time. Anthropic loves you. I'm sure you're like, maybe in the future you'll have like some kind of ribbon or something for making them so much money. But like, there's no way that you can justify all of the output, like, or all of the costs associated with the output that you're getting. I have not seen anyone do something just mind blowing impressive to the point where I was like, that's worth all the money that you spent in one go. Um, I don't think like one shotting for the sake of one shotting makes sense to me. I do, like, longer tasks. I will have my models constantly work on longer tasks. I'm very big on the subject of forced hooks, um, which maybe we need to talk about, like force hooks and memory and things like that. Maybe we'll do an episode on that. But like, this right here is like, I have standards with a lot of my outputs and things like that. I think that's why, like, a lot of the quality for the things that I put out, it's at a higher level. It's because of some of these tactics. So, like, I'd really, really advise you to like, really question or Start thinking, be engineers again. If you're going over the deep end, just be engineers again and like truly question or tinker or mess around with and explore the outputs that you're getting and start seeing if it's really worth it or not. Now, ah, I will say this. If you're getting just terrible outputs all around with no adjustments, that could be like an Opus 5 situation now. But if you've been feeling like that for months from uh, now and you still feel like that for months after this, this is gonna, and I hate to say it like this, but there's probably gonna be a user error scenario here because like, you're even finding some engineers that I've like truly had appreciation for, starting to then realize, like, okay, I'm starting to lean way more into agent decoding now. Well, let me ask you Leon, like, what, what, what is your prompt? Or like, how do you test the models to see if like, you like the experience or not?

Speaker D: Yeah, I think we have a lot of similar ones, um, where I have one super old project that's dependency hell, that's like switching from different Python versions that just a real mess to get through and get working again. And so I like to see how the models reason through all the new dependencies, pulling things that don't exist anymore, like finding alternatives. That's like a really good uh, project for me and I run it every single time I use a new model. And you'll be surprised, sometimes they do really well, sometimes they do poorly. Um, so that's one of them. I always have a front end task too because I'm always curious, like, can I finally use this for design? Uh, and then I have one for UI specifically. Like how good is it making UI elements and can I prompt cleanly and clearly? So those are things. One of the newer ones I've been trying out is do the LLMs know that they're an LLM? Because whenever you can, anyone can try this. Take whatever you want to build and ask it how long it's going to take to build and what you'll get back is most LLMs will be like, oh, this is going to take six months or like six weeks to build. You'll need a full team. And then I have to tell it like, but we literally made this 30 seconds ago and it took you two minutes, right? Like the, the LLMs are trained on like what humans can do, but not what they can do. And so I find even with like high reasoning models, they're not really understanding what they're capable of. And so one of my favorite prompts early on is like estimating how long is it going to take to do things. How can it do things? And with Fable 5, I'm starting to realize that it knows that it can do things that human, humans can't do. Not for every task, but it's starting to become a little bit more aware. And that's I think, changing what it's capable of. So those are some of the fun ones that I look at, so very similar. And then some other ones that I think that um, as the models get more capable and the training changes a little bit, we'll, we'll see more improvement on. Do you have any canaries like that you, that you use when new models drop? Like I, I know we all have kind of like things that we do to test the workflow. There are like certain prompts you run it through. Is there anyone that you found is like, um, super useful? When a new model drops, what do you reach for? What's the first thing you test?

Speaker C: Yeah, my tests are usually like this. Number one, I open a project that I've already built.

Speaker A: Mhm.

Speaker C: Right. So an existing project, not from scratch. And I will give it a feature and just say like, hey, this is what we're executing at. Even if the feature exists, we'll just delete the code. It's all on version control anyway. If you're not using version control AI, like please open up, uh, how to use git and GitHub and then go back to using your AI because like you're gonna kill yourself. But everything I do is version control to the max. Right? So it's like, hey, let me get rid of this feature. Cool. Now build this feature, give it to two separate models and let's see how it goes. The other one that I often do is I will give it several tasks like hey, um, one is for example, can you build me a full fledged portfolio site? I just want to see what it comes up with. Some of them creative, some of them not so much.

Speaker E: Ah.

Speaker C: Another test that I often give, um, I think this is a cool one, is make a chessboard. M make a chessboard and I just want to see how it functions. A lot of them are already trained on this stuff, so they know it. Right? So it's like now it starts making different choices at certain areas. Sometimes I'll have them make a chessboard and just makes a chessboard and moves. Sometimes it'll make it where it captures points per move. Sometimes I have an area on the side where it puts the, the pieces that have been captured so now you see which pieces are gone. Um, one time I even had with. Now this was with Fable where it made the chess, uh, board. This is the first time that I used Fable. It made the chessboard, it made a side where it had the pieces on the side, but after the pieces were taken it literally showed them like laying down. That to me was the first time I ever saw a model do that. I thought that was really impressive. Now I will say for everything that Fable is and all the good stuff, I'll be honest, I find myself using 5.6 way more. Um, but that's with a caveat. I personally think anthropic is they just got it nailed down unfortunately to the best levels of thought. It can execute thought tasks in a great way. I personally think 5.6 soul is one of the best models that are around right now for front end related tasks. I have not found anything that executes at such a great level on front end related items without me having to reprompt or reiterate over and over and over again. It's been truly amazing at that. Um, and so like often once I tell it once, I will never have to reprompt generally, uh, on what it is that I execute with it. The other thing is personally, um, Codex or ChatGPT desktop and all that, whatever you want to call it, has voice now. And so I could literally go from my phone and I can use the remote feature and we need to do a whole episode breakdown on this but I use the remote feature and I will talk to Codex through remote on voice and I can literally tell it. So funny enough, I have this at uh, my desk now and it literally holds my phone for me and I will literally hold my phone like this and tell Codex what to do. And I'm looking at the screen so it's like the perfect thing. It'll voice control everything and it works so fast. Um, it's been like really, really, really good at that. And so uh, I, because of that, it's been like so much better to do front end things because I'm literally telling in real time and I'm seeing the changes in real time so I can tell whether it's accurate or not. It's just so much easier because of that.

Speaker D: I gotta show you my phone, Stan. My wife stole it. I think it's on their desk. This is it. Oh, because it has the MagSafe mag safe.

Speaker C: Yeah.

Speaker D: And then, um. Oh, and so then it just magnetically attaches to the back of the phone. And so I. I just kind of always have it on there now.

Speaker C: Nice.

Speaker D: And, uh, whenever I need to. To talk, I pop it open and it mounts super easy. And so this is always on my phone now. And then I always carry a small Bluetooth keyboard in my pocket.

Speaker C: Okay, I don't do that one.

Speaker D: This is literally how I'm coding a lot now.

Speaker C: I will show you. Funny enough, my phone is a foldable. This is how I know me and Liana like bros, dude. All right, so I have, uh, a foldable phone that doesn't have Magsafe. So I literally bought this, like, little case. It's like eight bucks or whatever, and I attached it to the case, and so now it will hold the phone up whenever I want. But now I can open up my phone, uh, and view anytime that I want in that regard. But it isn't at, like, the best level for my mouth. And so what I'll do is move this mic out of the way and move that stand right in front of me, like the mic, and use it that way.

Speaker D: See the.

Speaker C: Oh, now we're just bringing out contraptions.

Speaker D: Oh, no, no.

Speaker B: This.

Speaker D: This is. This is if you. If you're listening to the podcast, do this. Make this investment in your life. DJI Mini.

Speaker C: Oh.

Speaker D: So it's a little mic, and the beauty is it has a. Has a clip, right? You can clip it, but it also has a magnet, so you can put it on anything, right? It doesn't matter where it is. It just goes through my shirt, no matter what I'm wearing, no matter what I'm doing. And now I have a mic that's going to my phone, so you can just, you know, kick back and have not worry where the position is. And it's being picked up by the mic. Game changer.

Speaker C: Yeah, I've got the rode mics, and they're hit or miss, but I feel like I gotta always hold them. I can't just leave it on, even though it has a clip. This one.

Speaker D: This one's good.

Speaker C: Very cool. All right, y', all, we were geeking out for a little bit. Sorry about that. We'll figure out if that makes it to the episode or not. But, uh, it's time for everyone's favorite part of the podcast, including mine. The Ask Danny and Leon a question. Leon, what question do we have today? All right, I've got the question today, and it's. I graduated in 2024 without a job offer. Not hearing back from many companies, I decided to prepare for A master's degree. However, I wasn't able to get into the universities I was targeting. So I shifted my focus back to finding a software engineering job to strengthen my profile. I'm currently contributing to an open source project. Given my current situation, how should I strategically plan the next stage of my career? For those that don't know, this one came in, uh, my weekly discord call with the group and, um, I thought it was a good question to bring in here today.

Speaker D: Yeah, I have a lot of thoughts and one of the things I always try to work through with folks is why things aren't resulting in ways that they planned. And when you tell me that you want to switch focuses back to software engineering, I think there has to be a clear focus on goal setting and figuring out the effort required to get to where you want to go so we don't wind up in a year from now having something similar. Where I tried software engineering and that also didn't work out because this is now two, almost three years of you attempting things and things not going your way. So to me that's kind of the root, the root thing I'd want to fix. Software engineering is kind of the other elephant in the room, but for me it comes down to, okay, what does your goal setting look like? How are you not going into your days like an accident? What does your time look like? And are you able to figure out the highest value things for you to be working on to achieve your goals? Now, sometimes life gets in the way and that's understandable. I think in your circumstance. This is where Community is really important. So I'm glad this came from Danny's discord. I'm glad folks have similar questions in my discord. Kind of come to Community with these, because I think you need an outside person looking at what your days, weeks and months look like to make sure you're doing the best with your time. That's where I would start is do you have a reasonable actual game plan that gets you the result that you want and can you take daily action toward those things in a way that makes sense? After that is the code and we can talk about that separately. But I think for anyone listening to that kind of feels the same way where they're not getting momentum. You have to really be clear on, on how you're approaching your days and not going into like an accident and really setting that groundwork so that you can do the things that are going to get you to where you want to go.

Speaker C: Yeah, you know, it's funny, I Haven't really fully, uh, announced this publicly yet, but by the time this episode goes out, it's, um, probably going to be ready to be fully announced. But I've been trialing this with a, uh, select group of people, and I've been creating a, uh, solution in a program called Undersold. Right. And to me, you're like, the prime person for this. People have been underselling their skills for years to where when a recruiter or hiring manager or someone along those lines sees them, they're like, I don't even know what to do with you. Like, I don't know if I should be excited, I should be annoyed. I don't know if I should have enthusiasm or lack thereof. There's no info to go off of even reading your question. I'm like, I don't even know what to help, uh, you with because, like, you're not even describing the things that matter the most. Like, what did you do in that time period before deciding, I'm going to go for the master's degree. Oh, I didn't get accepted by these universities. Well, why didn't you get accepted? Did we even audit that? Are we getting interviews? We're not even getting interviews. Why, why aren't we getting interviews when we apply to things? Oh, I did get interviews. Okay, cool. Why were you losing them? Like, there's so many areas in here. So it's like, you contributed to open source. So what? Why does that matter? Why should somebody care? Uh, I fixed a bug in xyz. Okay. Why does somebody care? So what? Uh, posts were silently, like, failing on Time zone, so I fixed that. Why does anybody care? Like, this always goes back to that same thesis of people are so focused on explaining a task or an item versus the thing that it does that people don't know what the other side looks like. So, number one, you automatically. Just from the question, I'm able to derive that you have a very big issue in me being able to explain the things that you're a part of. Even the, like, the framing of your question comes from a very negative place of, like, this didn't work out, this didn't work out, this didn't work out. And I'm like, I need something to go off of. So the fact that you're contributing open source is good. Remember the core thing about open source, though? No one hires you for an open source contribution 80% of the time. Right. 90. These are made of M numbers, by the way. But, like, the majority of the time you're not getting Hired because of an open source contribution. There are definitely scenarios where maybe somebody contributes to a company and they're like, hey, we really like this. We want to bring you on. These are like, not the common scenario. You're doing it because. And the strategy should be, you're developing the talking point that you need for the interview. Like when they're like, hey, what have you done for the last year? You'd be like, oh, I contributed to this open source project and I was able to solve X problem, Y problem zone. They get excited about that. If you just say, hey, I contributed to this open source project. It's like you've basically created what I call and this is the problem with degrees, uh, and the way people showcase them. In my opinion, you basically have the check.

Speaker A: This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture ah are working together to reinvent the rhythm of ad sales using automation, analytics and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more@accenture.com Spotify

Speaker B: this episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a page restoration block. Or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses setup required compatibility and availability various 18 plus

Speaker C: like the mental like mhm. Okay. And they moved on. It became a half a second moment instead of like ah, uh. This is the verifiable proof that I need. So if you just say hey, I contribute to open source, they're like okay, next. Or you talk about the business problems that you solve. The fact that you have a degree. You enter in conversations that many people can't. So the fact that like it's a small footnote instead of like a massive one. Like I've got a degree and I've did this and I did this and I solved this problem. My thesis was this and let all that could be there but instead it's like almost a nothing note. And so I think that whole positioning needs to change on that and like you need to create more moments besides the uh, M. That was pretty cool and you move on. So I think that framing needs to happen. I Think you need to think about the business solutions that you could potentially be solving. And here's the thing, no one hires you because you could regurgitate syntax. They do hire you because you're a problem solver. So finding out how many problems that you could be a part of and solving them. This is crucial, whether it's for the open source project. Prime example. I was talking to somebody in my discord on Thursday, and they volunteer, uh, at like this, um, organization. They volunteered there for two years, but they only give him like two hours a week. Perfect. Nobody asked you how many hours a week you were contributing there. But the fact that he's part of some problems that he could talk about, that's a win. So find out what it is that you could be a part of and talk about those. That's exactly what I'd be focused on right now. But the other thing I'd be focused on, honestly, networking my behind off. Like, I would be networking with so much intentional purpose, it would be sickening. Like, it would be hard to get me off a computer right now at this stage where everything has gone wrong in multiple areas and, and like, I need to change the circumstances of my life. If you think that for a second I'm going to be there, not, like talking to folks intentionally. You got your, you lost your mind. I, uh, would be sending out more DMS than a human being knows how to deal with. I'd be starting more conversations and I'd be keeping track of it. I sent out these 500 DMs. Of these 500 DMs, only five came back favorably. That's horrible. That means you have to send out 100 DMS for one person to finally respond. Something's definitely wrong with your message. Fix it. Oh, I sent out 100 messages. I got no response. Fix it. Like, I would be experimenting, auditing to find out what that voice looks like, what that great conversation looks like. Because that costs you nothing. So start sending out messages, talk, leaving comments, being, uh, active in discord communities. I want to talk to everybody. I want to go to every meetup. I want to do all of it. And so if you're not doing that, understand it costs nothing to do it. And the. But it costs you everything if you don't.

Speaker D: And hey, Danny, it sounds like a great service. What was that domain again? Because I don't think you dropped the domain.

Speaker C: Undersold AI. Undersold AI. Alrighty, uh, folks, it's been real, it's been fun, and we'll see you on the next one Goodbye, everybody.

Speaker E: You know those tiny back to school emergencies that somehow become your problem? That's why I love Uber Eats. You can order school supplies, snacks and lunchbox essentials for $5 or less. M so when your kid casually drops, I don't like peanut butter anymore. Or I need five green highlighters for a project due tomorrow, Uber Eats has you covered. Get everything you need for back to school today from your favorite brands like Aldi and Staples and Uber eats. Order now ends 97. $5 or less before taxes and fees. Select items only. Availability varies. See app for details.

Speaker C: Close your eyes.

Speaker E: Exhale.

Speaker A: Feel your body relax, and let go

Speaker D: of whatever you're carrying today.

Speaker E: Well, I'm letting go of the worry that I wouldn't get my new contacts in time for class. I got them delivered free from 1-800-contacts. Oh, my gosh, they're so fast.

Speaker C: And breathe.

Speaker E: Sorry. I almost couldn't breathe when I saw the discount they gave me on my first order. Oh, sorry. Namaste.

Speaker A: Visit 1-800-contacts.com today to save on your first order.

Speaker C: 1-800-contacts. Uh,

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Jeff Dean: The 1% Rule for Building in AIY Combinator Startup Podcast · on Context engineering98 / 100
  • The 18x Midas Lister Betting $3B on AI (and calling most of it fake) | Navin Chaddha, MayfieldThe Peel with Turner Novak · on OpenAI91 / 100
  • AI Starts at the Top: Why Leaders Must Build AI Fluency FirstThe AI Advantage: Smart Tech for Modern Leaders · on Reinforcement learning from human feedback (RLHF)85 / 100
  • Fighting Fire with Fire: How CyberProof Is Automating Cyber Defense with Edy AlmerCyber Sentries: AI Insight to Cloud Security · on Anthropic80 / 100
  • How AI Will Change Competitive Advantage with Doug StephensThe Future Of Less Work · on OpenAI79 / 100
  • Using Airflow for diverse client projects at Accion LabsThe Data Flowcast · on OpenAI79 / 100

More from The Programming Podcast

All episodes →
  • Stop Vibe Coding, Start Architecting! Our AI Workflows68 / 100
  • Kimi K3 China's 2.8 Trillion Parameter AI Just Dethroned Claude!
  • Why Senior Software Developers Are Failing to Land Jobs (And How to Fix It)
  • The conversation every software developer needs to hear
  • Jack Dorsey’s “Mini AGI” Thesis: The End of Corporate Hierarchy!?
Explore the best B2B Engineering & DevTools podcasts →
All The Programming Podcast episodes →