The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/This Day in AI Podcast
This Day in AI Podcast artwork

We Committed Fraud with OpenAI's New Image Model (and Called Mum) - EP99.38

This Day in AI Podcast · 2026-04-24 · 1h 35m

0:00--:--

Key moments - from our scoring

Substance score

28 / 100

Five dimensions, 20 points each

Insight Density7 / 20
Originality6 / 20
Guest Caliber2 / 20
Specificity & Evidence8 / 20
Conversational Craft5 / 20

This episode covers the flood of AI model releases in recent weeks - GPT 5.5, Claude Opus 4.7, GLM 5.1, Kimi K 2.6, and Qwen 3.6 - with frank skepticism about OpenAI's vaporware launch strategy versus Anthropic's faster availability. The hosts dig deeper into a critical economic theme: the massive subsidies distorting the market. VCs and sovereign wealth funds cover 70% of OpenAI's costs and 33% of Anthropic's; consumers pay only 5.5% of actual token costs. They use hard data to show how Sim Theory, Cursor, Perplexity and other consumer platforms layer subsidies on top of already-subsidized model providers, creating a false economy. The conversation reveals why enterprise customers are the only ones paying real costs, driving the labs' pivot toward workspace agents (OpenAI), Claude deployments (Anthropic), and Gemini enterprise (Google). They discuss how subsidies prevent proper pricing discovery and value assessment, and why agentic tasks cost 10-50x more than single prompts - a hidden cost consumers discovering agents will soon face.

Key takeaways

  • →Agent-based tasks cost 10-50 times more than single chat interactions due to system prompts, planning, reasoning, and tool calls, a hidden expense consumers will encounter as they shift from chat to autonomous agents.
  • →Enterprise customers are the only segment paying close to real token costs, which is why all three major labs (OpenAI, Anthropic, Google) are aggressively pushing workspace agents and enterprise platforms despite consumer-facing hype.
  • →Two-layer subsidies in consumer AI platforms - labs burning money, then platforms like Sim Theory burning credits on top - create false price expectations where consumers balk at $30/month while actually consuming $700+ monthly value.
  • →The economics of AI pricing mirror the newspaper industry's disastrous free-ad model; when subsidies end and real costs surface, either adoption collapses or service quality degrades for price-sensitive users.
  • →GLM 5.1 and Kimi K 2.6 perform comparably to Claude Opus for agentic work and code generation, yet brand loyalty and perceived reliability cause users to pay premium prices rather than switching to nearly-equivalent cheaper alternatives.

In this episode

  1. 1Latest AI Model Releases: GPT 5.5, Claude, GLM, and Others
  2. 2OpenAI's Strategy Shift Toward Super Apps and Workspace Agents
  3. 3Model Performance Comparison and Developer Preferences
  4. 4The Economics of AI Subsidies and Real Token Costs
  5. 5Enterprise vs Consumer Pricing Models
  6. 6Google's Silent Retreat and the Enterprise AI Wars

Mentioned

OpenAIAnthropicClaudeGPT 5.5GLM 5.1Kimi 2.6GoogleMicrosoftAmazonCursorPerplexitySim Theory

Topics in this episode

OpenAICursorClaude Opus 4.7DeepSeekGPT-5.5computer useoperatorgpt4GLM 5.1Kimi K 2.6Qwen 3.6OpenAI workspace agentsAnthropic Claude deploymentsGoogle Gemini enterprise agent platformSim Theory platform

Questions this episode answers

How much do AI model providers actually spend versus what consumers pay?

VCs cover 70% of OpenAI's costs and 33% of Anthropic's; enterprise pays near real cost; consumers pay only 5.5% of actual token costs. Hyperscalers use subsidized cloud credits, making true pricing invisible across the board.

Why are OpenAI, Anthropic, and Google all launching workspace and agent platforms at the same time?

Enterprise customers are mandated to spend AI budgets and are the only segment willing to pay real costs; labs are abandoning consumer strategy and racing to lock enterprises into their platform layers through workspace agents, Claude deployments, and Gemini enterprise tools.

What makes agent-based AI tasks so much more expensive than regular chat?

A single agent task requires system prompts, planning, reasoning, multiple step iterations, and tool calls - using 8,000-30,000 input tokens and 3,000-8,000 output tokens versus 800 for normal chat, making agents 10-50 times costlier per task.

How do GLM 5.1 and Kimi K 2.6 compare to Claude Opus for coding and agentic work?

Both perform comparably to Opus for code generation and agentic loops, often faster due to model size efficiency, yet users stay loyal to Opus due to brand consistency rather than measurable performance gaps.

Why don't Google's subsidized TPUs let them dominate the AI market?

Despite owning TPU hardware that would let them run models nearly free and undercut all competitors, Google has made their models expensive and largely went silent on releases, missing an obvious strategic opportunity to entrench the market.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

7 / 20

The episode contains scattered substantive observations about model economics, pricing mechanics, and the Everything App paradigm, but these are heavily diluted by 95 minutes of tangential commentary, model comparison chatter, and extended riffing on unrelated topics (graffiti, personal anecdotes, rap songs). Novel insights are present but sparse relative to total runtime - roughly 20-25 minutes of genuine substance buried in padding.

VCs and sovereign wealth funds are paying 70% of the real cost of your token. So OpenAI burns 70% of their revenue and Anthropic burns 33%.
agents consume at machine level speeds, not human level speeds. So the actual consumption of the resources is going to be much higher

Originality

6 / 20

The hosts rehash well-known critiques (model subsidies destroying economics, the SaaS apocalypse, Everything Apps as platform strategy) without meaningfully advancing these arguments. The fraud demonstration with GPT Image 2, while attention-grabbing, is not a systematic analysis - it's a stunt. Most takes on pricing, Anthropic vs. OpenAI positioning, and enterprise adoption recycle existing industry commentary.

these labs are starting to figure out similar to what, what Elon Musk announced that Grok wanted to build is this Everything app
the SaaS apocalypse, uh, where companies are making, like, huge layoffs, blaming AI

Guest Caliber

2 / 20

This is a two-host conversation with no external guests. While both speakers appear to have operational experience (references to SIM Theory, infrastructure decisions, enterprise interactions), their identity and credentials are never established, and the show format provides no third-party validation of expertise. For a B2B podcast, the absence of domain experts, practitioners from named companies, or verifiable operators significantly limits credibility.

We have um, an enormous backlog of tickets. Full disclosure, we're aware of how bad our support is.
I'm very committed to upholding our standard of mediocrity.

Specificity & Evidence

8 / 20

The episode includes concrete pricing data (OpenAI $5/$30 per million tokens, Anthropic at different rates, GLM 5.1 at $4.40/M), specific model names and benchmarks (Sweat Bench Pro 64.3, code arena scores), named services (Help Scout, Stripe, Salesforce, Figma), and personal examples (the $1.5M infrastructure spend, 600 enterprises spending >$1M/year). However, many claims lack citations, sources are rarely attributed, and broader economic assertions (e.g., tokenizer cost increases of 1.35x) are stated without evidence.

VCs and sovereign wealth funds are paying 70% of the real cost of your token. So OpenAI burns 70% of their revenue and Anthropic burns 33%. The hyperscalers, Microsoft, Google and Amazon are paying through subsidized cloud credits
GPT 5.5 they're charging $5 per million input and $30 per million output

Conversational Craft

5 / 20

The hosts rarely challenge each other substantively. Disagreements are surface-level (e.g., on model preferences) and quickly conceded. There are no probing follow-ups on important claims - e.g., when the 70% subsidy statistic is introduced, it's accepted without interrogation of the source or methodology. The hosts frequently go on tangents (rap songs, personal stories) rather than deepening insights. The fraud demonstration is presented as entertainment rather than analyzed for systemic implications.

Yeah, I sort of agree with you. People are just going to stick to what's out there.
Yeah, exactly.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B51%
  • Speaker C42%
  • Speaker A4%
  • Speaker D3%

Most-used words

model49models37real36agent30back26point26agents26code22better22fine22agentic21call21open20level19value19saying18

Episode notes

Join Simtheory: So Chris, this week... a LOT has happened. We're back to regular programming (maybe), and back with our average takes. Nothing's changed. GPT-5.5 just dropped today - but you can't even use it in the API. Vaporware? OpenAI is charging MORE than Opus 4.7 and we haven't even tested it yet. Meanwhile Claude Opus 4.7 landed a couple weeks ago and... the vibes are off? Mike's actually going BACK to 4.6. Something's wrong. But the real star: OpenAI Image 2. This thing is genuinely terrifying. We committed what can only be described as "parody fraud" - faking a council letter so realistic Mike's own mother fell for it on a phone call. Then Chris posted a fake development approval with the mayor's real name into a local Facebook group and had to delete it when someone tagged the actual mayor. The forgery capabilities are absolutely unhinged. Also: GLM 5.1 is so good Mike forgot he switched to it. Kimi K 2.6 is criminally underrated. VCs are paying 70% of your real token costs. Consumers pay only 5.5% of actual cost. The everything app war is ON. The SaaS-pocalypse is real. And we made two new diss tracks. Chris made a graffiti sign in LA.

Full transcript

1h 35m

Transcribed and scored by The B2B Podcast Index.

Speaker A: I went from six to seven. Yeah, six to seven. Sweat Bench Pro 64.3. That's AI heaven code arena number one plus 37 on the score, 87.6. Verify what you benchmarking for.

Speaker B: So Chris, this week a lot has happened in the world of AI. You get the drill. Everyone's been cooking, mind blown, we are cooked. Everyone's going to lose their job. We're back to uh, all that uh, joyful narrative again. Uh, but we are back, back to regular programing. Maybe, you know, let's not over commit, uh, back with our average takes. Nothing's um, obviously changed since we uh, since we left. How have you been?

Speaker C: Yeah, pretty good. I'm very committed to upholding our standard of mediocrity. I can see my camera is like already blurry. I don't know why and I feel like that's the kind of standard we want to maintain here.

Speaker B: And for those who listen, got a nice little graffiti sign I see.

Speaker C: We were in LA recently. We went to a graffiti making class and I made something and it was terrible. So I just painted over it with this day and AI and actually out of the class I think I made the best artwork. Shows how bad the other people in the class were.

Speaker B: Really very, very nice. Okay, so we do have a lot to go through and we're just going to take our time and sort of catch up on everything that has happened and all the different releases that we wanted to talk about. And then honestly there's some higher level themes at play right now that we've both been talking about and I think we're a little bit excited to talk about those. So a lot of new model releases, a lot of new releases in general. We're at that point in the year where everyone's like, we're excited to announce, we're extremely proud to announce all that kind of stuff. So we've had just today GPT 5.5. Not to be confused with uh, 5.4, 5.3, 5.2 or 5.1 or 5. Uh, prior to that we've had Claude overs 4.7 a couple of weeks ago. Uh, we'll talk about that in a minute. I thought the biggest, like most impactful release really out of all these models was OpenAI's Image 2, which we'll get to been uh, having a little bit of fun with.

Speaker C: We have done some um, kind of extreme things with this that got me nervous before this podcast.

Speaker B: Yeah. If we're not back next week, you'll know why after we tell you what we've done, uh, also GLM 5.1, we've been really impressed with that model. Kimike 2.6, also very impressed. Lots to share on that front. And then Quen 3.6 as well, which we. Honestly, there's been so many releases we kind of forgot about.

Speaker C: Yeah, it's almost like too much, guys. Like, we calm down a bit. We don't need all this. Like, it's nice of you, but, you know, like, we're good with what we have.

Speaker B: Yeah. And, uh, and so, uh, some other things. Open AI launched, uh, agents, I think they're calling them, or workspace agents off the back of originally the sort of failure that was GPTs. Uh, everyone's trying to build an everything app now, and there's some sort of war taking place around that. So we want to talk about that. Uh, but let's jump in first to the latest release. We'll start from the latest and sort of work our way back. Today we had, ah, GPT, uh, five.

Speaker C: I'll give you my assessment. Bang. Vaporware. You can't use it. It's not available in the API.

Speaker B: Well, I mean, that's. That's a little debatable, isn't it? Because, uh, I'm so out of practice here. I don't even have the tab up. Uh, here it is. So introducing GPT 5.5. But you were right, it's not available in the API. I think what's happening is this is speaking to the larger trend now of it. We're really entering into that super app product world where the labs are way more excited to get these models into their apps, especially OpenAI, with their, like, competing now directly with Anthropic and B2B and trying to prove that they have this super app, uh, one for everything. So they're just pumping the models into those super apps really quickly. And I think that's what we've seen with GPT 5.5.

Speaker C: Definitely still on the open AI front. But if you look at Anthropic, the gap between them announcing something and it being available, if anything is shrinking. Like, you know, even the, you know, elites getting it. Like, I've been, I've been cooking with this model for three months and now you guys get it at announcement time. That doesn't even seem to be happening anymore. 4.7 was just there one day and we're like, oh, geez, we better add this and start using it immediately. So I don't know, it seems like more an open AI thing in terms of that delay yeah, the narrative, like,

Speaker B: the narrative around this stuff right now seems to be that they are really pushing hard in terms of just trying to catch up to anthropic. It's. They've gone from just blitzing ahead and being same day's, you know, release cycle to, uh, you know, this. Like, it just looks like a very confused strategy right now. But in terms of benchmarks, like all of these model releases, they're saying GPT 5.5 benchmarks, pretty much higher on every front than Claudopus 4.7.

Speaker C: Yeah, right.

Speaker B: I mean, I don't believe that for

Speaker C: a second they can have all the benchmarks they want. We look at the usage and who's using what, and people just aren't that excited about OpenAI models anymore.

Speaker B: Yeah, I think we have to give 5.5 a chance. Like, we've never used it, but it was interesting in SIM theory when we saw 5.5.4 usage. So it, it went right up. Like people were really excited to use it for a little while and then it just slowly peels away. And I think that when you move to an agentic world, These agentic loops 5.4 just doesn't, in my opinion, perform as well as the anthropic models or, or even, to be fair, GLM 5.1 or Kimmy K 2.6, uh, they perform so much better, in my opinion and experience at agentic operations than the GPT models.

Speaker C: They perform so well. It kind of makes me nervous, like, in the sense that I will use them for a period of time and be like, whoa, they're just as good. But then whatever it is within me just, I just want the best one. Like, and I just go to Claude Opus 4.7 because I'm like, I just want to be using whatever the best one is because I want this task done. But I kind of feel like if I just gave them more of a chance, they would get the job done just as well. And it always blows me away when you say use the GLM M 5.1 or Kimi 2.6, just how quick they are. They're just suddenly the answer's there. Like, I'm used to tabbing away and starting another process and then coming back. Whereas with those ones, you basically don't have to do that even in a full agentic loop, they just get it done quicker.

Speaker B: Yeah, my experience, actually, when I was flying to la, I was using OPUS to code in a bunch of tabs, right. And everything was going great. And then OPUS had this weird outage that we had to deal with. And so I had to switch away to. And I, I immediately went to GLM M 5.1 because I know that is like basically the closest thing, albeit like it's not that much cheaper to run, unfortunately, because it's a huge model and probably trained on the outputs of opus, let's be honest. But, uh, but after a while of using it, I did not notice any difference. In fact, the next day I was still coding with it and had not even noticed that that was my primary model. I was just opening new tabs and that was in there. And because it was performing so well, I really didn't notice any difference. So I do think you make a good point that you sort of get attached somewhat to these brands and this consistency of that model working and then that maybe stops you trying some of these other models that are, that are pretty damn good.

Speaker C: Yeah, I think it's when you're trying to get real work done, you're, you're sort of like, well, I can't, I can't chance it on this. I'm just going to have to pay what the price is. But the, the truth is that if you actually were denied access to the more premium models and only had these, I think it could be just as productive. I really don't think you would lose that much, uh, using say a GLM 5.1 than you are using opus. Like, yeah, there might be some things it's not quite as good at, but like if you were just totally banned or something from Anthropic and could only use that, I wouldn't be that upset. I think I'd probably just use it then. I almost need it. It's almost like a form of discipline, you know, like, you're not allowed to do this anymore. You have to use this one.

Speaker A: And I.

Speaker B: But I think that's the point we're getting to right where it, it really is at the product layer now where people are getting used to certain products and how they function and the things that they're able to do do. And you can see that with the Labs. And I think this is shown with GPT 5.5, how they're pushing it really hard into Codex, which is becoming like their Everything app. And then with Anthropic it's kind of similar. They're like pushing stuff into their application layer. Ah, first and they're trying to get people addicted to that application layer. And we'll get to it a little bit later. But I'd say arguably distorting the market. In terms of pricing, like taking the loss to get people addicted to, to their world and way of working.

Speaker C: Yeah, absolutely. I think, I mean, yeah, like you say, we will get to this later because I've actually looked into this quite extensively. It's a big point at the moment around the cost of things. And I think the model providers subsidizing the real cost is really skewing everyone's thinking as to what's possible in the real cost of things and, and causing like a false economy in terms of people thinking something to should cost a certain amount when it actually costs more and not making that value equation as to is the value I'm getting from this worth what I'm paying?

Speaker B: It's so reminiscent of newspapers when the Internet first came around, how they're like, oh, we'll just make it free because everyone still buys the newspaper. Uh, you know, and we'll just sell ads, right, because we just want eyeballs. And then over time then they try and charge. That didn't work out terribly well. The quality of journalism goes down and everyone's like, oh, why is journalism so bad? And it's like, because we, no one's paying for it anymore, so no one really values it. And I think that's kind of what's happening in the, in the model realm as well. Or what might happen as well is like it's been subsidized so much when they eventually have to charge the right price. You know, who knows what will happen. Like, I don't know if people are going to be willing to pay, uh, or if they are. It's, you know, they're going to have a degraded experience because they're not willing to pay as much as it actually costs.

Speaker C: Yeah, I'm of the opposite opinion. I actually think that there's so much value there and people should pay for it. I just think they've been conditioned to think it's cheaper than it is and haven't made that own assessment with themselves that I'm this much more productive because I spend this money and I'm willing to spend it either as a, uh, you know, an expense for my job to make me better at my job or my company pays and I can prove the extra value I'm getting from it. Like, I think the value is there. I just don't think it's being. People aren't um, thinking about it right now. I think most people are pretending the problem doesn't exist and looking at this line item of like, AI usage as a, like necessary evil and not realizing it's almost like paying for additional staff. Like, it's almost like having more employees at your company. That, uh, an expense that everyone's willing to take on. If it's, if it brings more value to the business, then you're willing to pay for it. But because it's a computer, you just don't see it in the same light.

Speaker B: Yeah, well, I mean, like, let's get into that conversation anyway. Like we've started talking about it, might as well go into it. I think that, you know, this is, this is going to be the narrative coming up to some of these companies going public where, you know, how much are they subsidizing it? And I think. Didn't you have a real stat around how.

Speaker C: Yeah, yeah. So these are, uh, these are actually Kimi 2.6 certified stats. So like, you know, these are like top level, level legit mediocre stats. But so Listen to this. VCs and sovereign wealth funds are paying 70% of the real cost of your token. So OpenAI burns 70% of their revenue and Anthropic burns 33%. The hyperscalers, Microsoft, Google and Amazon are paying through subsidized cloud credits and infrastructure buildouts.

Speaker A: Right?

Speaker C: We've experienced that ourselves, where we were subsidized for a while, directly passed to the Sim Theory audience and burnt through it in like record time. Like absolutely. Just mince me these credits with our audience and we did the same thing. And um, I'll finish this and then I want to make a point about that part. And then it says enterprise customers are the only ones paying something close to the real cost, which is why every lab is desperate for that enterprise revenue, something we also have experience. The enterprise is the first group of people to actually get the value and be willing to pay close to what it really costs. Right? And then the consumers are only paying 5.5% of the actual cost of what they consume, which sounds about right to me. Right. And we saw it like we, uh, we changed our token model in SIM Theory and there's immediate backlash in terms of people being like, hang on a sec, I burned through my tokens in half an hour. What's going on? And the truth is we just finally charge what it actually costs us for people to use it. And so it's, it's kind of crazy like the, the, the skewed economic world we live in with these AI tokens.

Speaker B: And I, I think the other point to make is, well, a few things you mentioned earlier, like passing on credits like a year A year ago was it, or maybe two now when we released the Workspace computer where you could have like a Windows box in the cloud and your AI could operate it and it was like your computer, you could install apps in the cloud. We spent $1.5 million in uh, like not very long, I think like two months on that.

Speaker C: You could have put a whole other gold chain with that.

Speaker B: Yeah, I know. And uh, and it, that was in, in credit. So again a heavily subsidized. Could we have done that without some. Like there's no way, like no one would have paid that because it was really just experimenting around with the technology. So.

Speaker C: And yet, and yet we found real value there. Like it was the most in demand thing we've ever done. Like had we say being venture capital backed and could burn the VC money like these guys are doing, we could have maybe run that to the point where we, we got the economies of scale right or the cost base right. And actually provided it as an ongoing service. Like it really was a legitimate thing that people really wanted and I would argue probably still do want.

Speaker B: Yeah, I like, I would still like to have uh, like a fully like a full cloud computer. I could deploy my agent on uh, instead. Well I, I don't mind my Mac Mini over there. It does the job. But ultimately it would be cool to have. I like I don't think it's a great business model but it is like for like geeking out over. It's pretty cool. But I think you're right like the subsidies at least from a consumer point of view the whole idea was the ads business, but I don't think that's working terribly well. People don't want a compromised AI experience. And right now there's a lot of people fighting for the attention or the, the, the token usage or the getting you to prompt in their world that they are willing to subsidize it. So there's always someone willing to discount more to get user share. And I think this is what's eroding away the consumer business at least, whereas the enterprise, like it's a whole different ball game.

Speaker A: Yeah.

Speaker C: And the point I wanted to make earlier about this, that's totally crazy is when you think about how much say anthropic and OpenAI are uh, sort of selling $2 for $1 or whatever they're doing. You know, like they're passing on value to us by burning their own money. Right. And then you think about say SIM theory where we were effectively doing the same thing for our own audience. So you've got two layers of people subsidizing your usage. And I would argue a lot of consumer AI platforms, like, if you look at, say, perplexity and some of the other ones people have used over the years, like Cursor and stuff like that, they are subsidizing as well. So you've got two layers of subsidizing the actual cost of this stuff and then people using it and being like, you know what, 30 bucks a month, this is a, like, off, uh, so expensive. Like, I just, I couldn't be bothered. I'm going to downgrade to the $15 a month plan because 30 is too much, like, and you think, but this is probably costing $700 a month when you have the two layers of subsidies. So it's like, it's this weird thing where you wonder where the actual value lies, but it's in, it's in such contrast to my own experience where I will spend whatever it takes. Like, I don't even, I can't even imagine how much money I spend on our own system, right? Like, you made us switch to, like, you know, auto renewals like everyone else does. And I think mine auto renews like every 15 minutes or something like on the planet in terms of tokens. But I'm like, I get so much done. I'm, I'm doing the work of the previous 10 of 10 years of me in a, in a week. Like, it's just the, the value I get from it is so big. I would pay, you know, a couple of hundred thousand dollars a year to get what I get from it. Because I think I'm delivering more value than that. And I think that this, this value perception is really skewed. But I think that people are misguided about it. I actually think rather than them seeing it as this excessive expense, I think they should see it as an opportunity. Like, if I am the one spending the money on this, I can be the one who's this much more productive and directed in my activities. Um, and that's a real advantage for me.

Speaker B: But isn't this the whole point? Right, Because I think you're looking at it from the point of view of someone your entire life has had employees, right? Like you've hired developers, marketers, sales people, like all these various roles. And I sort of look at it from the point of view of, well, you know, if I had to go out and hire people, manage them, um, you know, deal with people in a business, there is a cost to that. There's a mental Weight. It's a distraction, quite frankly, because you can't uh, stay really close to the bare metal. And so I think we are looking at it from the point of view when we're doing work or adding value as like, how many wages would you need to pay so in order to get to this level of productivity in the old world? So the question then becomes, okay, well, I am willing to spend like 100 or 200k a year on this stuff because I'm getting that value or getting that return, um, on my like token investment. But I think from someone who's, you know, like, say, a developer today, working in a business, they're probably looking at it like I just needed to now pay like an absolute fortune to do my job because there's this expectation that I will now output at this level and they might not be getting. It's not like their income's going up as a result of doing this if they have to pay for it themselves. Uh, and so I think there is this mismatch and if you're using it in your personal life, it's not like, you know, maybe you're just not seeing the returns there. So I do think that's why also Anthropic and OpenAI are pushing so hard into the world where to be successful. Like, they almost have to disrupt society, which kind of sucks. Like, it's like they have to replace elements of human wages because no one's going to pay more to do. Like, you know, like, there's got to be some trade off there. Like they've got to either see a productivity gain from all of their team in everything that they do, or they've got to lay people off and be able to like, keep running at the same pace. Like there's no that, like the economics have to balance out at some point.

Speaker C: Yeah, exactly. It's why, I think, why I find it so surprising that Google has just gone like completely dead and silent on their models because one advantage they have over everyone else is their TPUs, right? Like they have their own hardware to run this stuff. So Google could afford to basically make their crap free and just run everyone else into the ground by just, uh, you know, really subsidized, like properly subsidizing, and just go like, we're just going to be free for the next three years, build all your stuff on us and make everyone totally entrenched in the Google ecosystem. And yet they basically destroyed their models. And then, uh, they're also really expensive. So it's like, I just really don't understand what they're doing there when they have the ultimate platform to just flatten everyone and really bleed them out like you could. Right now, OpenAI is struggling financially. You could destroy them if you were Google right now and wanted to.

Speaker B: I think, uh, you know, they did announce, to be fair, they had that cloud next 26 like a couple of days ago. But there's just so many announcements right now, it barely blipped up on my radar.

Speaker C: Uh, well, actually I saw it when I did my Kimmy K 2.6 research and I just figured it was hallucinating. I was like, you're living in the past, man. Like, you know, this is, this is not real. You've made this up. And then I'm like, oh, yeah, they really did do something.

Speaker B: Yeah, I mean, they've got their Gemini enterprise agent platform, OpenAI have got workspace agents, Anthropic's got, you know, cowork and uh, Claude design and Claude whatever. And I think it, it just shows now the target. And I think people are noticing. This is like everyone's chasing the enterprise dollary dues because that's just where one, as you said earlier, people are just willing to pay in the enterprise because they're seeing the benefit.

Speaker C: Not just willing to pay, mandated to pay. Like, we have come across so many enterprises where there's an AI change officer, there is someone specifically in charge of a budget that they need to spend by mandate in their organization and they're looking where to allocate that money. So like, it's a difference between convincing, you know, a million people to pay you $10 a month or you know, one customer who's just going to be like, yep, let's put the full five into this because we need this in our organization and there's just very few places to put that money right now.

Speaker B: I think the other challenge, right, is if I'm a consumer and I just want to experience and experiment around with like different models and tools right now the change that's coming is as these things get more expensive, right? Like you've got to actually pay what they cost. Those experiments become really expensive. You know, you can spend like a hundred dollars USD very quickly on agentic, like trying some agentic stuff or like trying some scheduled tasks or playing around with agents. So.

Speaker C: Well, listen, listen to this. I actually got GLM 5.1 to do some calculations on this. It was saying if you are using like normal chat, a single chat interaction might be like 800 tokens, right? But a single agent task is like 8,000 to 30,000 input, 3,000 to 8,000 output. It's like 10 to 50 times more expensive because you've got system prompt, planning, reasoning, step multiple tool calls, each with growing context, and then the final synthesis on every agentic process process. Right. And so even though I believe you can actually make agent processes more efficient through like the way you stack the context and build it dynamically caching, all that sort of stuff, um, the cost is just orders of magnitude higher. But again, back to my earlier point, I would argue, like, I don't know about you, but I don't do anything in a non agentic mode now. Everything I do is delegation now and scheduling. Like I've got so many scheduled tasks, I've got so many agentic loops running throughout the day and I run them all on, you know, like cloud machines that can, I can walk away, I can shut my laptop and the work keeps going. Like that's how I work all day now. Like I'm stressed out right now because I know I don't have anything running and I should, uh, you know, and that's how I work. Like when I'm, when I'm leaving to go somewhere, I'll set two or three things off and then, you know, unwrap them like presents when I get home to see how they went.

Speaker B: Yeah. And I think in, in SIM theory too, the way I'm thinking about it now is like, how do you reunify these experiences? Rather than the complication of people selecting like Chad or agent or research or whatever, it's like, well, it should just work. Um, yeah. And I think that like starting to reunify that stuff is important, but at the same time, like you said, the agentic loops at their core do tend to burn a lot more tokens. Uh, but ultimately I think the outcome is just so much better, uh, with everything that it does.

Speaker D: Yeah.

Speaker C: And I think to use the modern lingo again with your agents, you've got to let them cook, like give them all the stuff, let them decide, let them do the tool discovery process, the file discovery process, let them do all that. And this is my probably, um, because looking back on the, the Open AI announcement around their agents, because you were saying, oh, it's, you know, they should have been there ages ago. It's kind of lame, but I said this is really the first taste of this kind of workflow for your, your average user. Like the person who's just sees uh, AI as chat. Right. So I actually think it's kind of significant because it's the first time you can sort of delegate. It's the first time you can set something off and have it working for you in the background for most people. And so I actually think it's kind of significant. But my criticism of it is this idea that you have to in advance specify which connectors or skills you're going to use. Like, they have the concept of skills in there which are like dedicated prompts for parts of the work and then which integrations you want to use, like Slack or Salesforce or whatever things you want to do. Now my argument with that is, okay, maybe in the scheduled task context it makes sense, but it's also a lot of setup. Like, you really should be able to just say, here's what I want to happen and let it figure out all of those details for you. And I think that that's where we really need to get to with the agents. It shouldn't be like this custom setup. Every time you want to do something. You really just should have like a, a working partner where you're saying to it, look, here are my. And this is what. I do this all the time when I'm working. I'm like, here are my problems. Like, here's what I'm really stressed about. Like, how the hell am I going to get this done? And the agent itself will coach me through the process of giving it what it needs to get that task done. And.

Speaker B: But don't you think this is a con? Like, this is a converging? Like these are two methodologies, I think, that are totally different. And I'm curious what how you actually work. So you've got like the Claude way right now, which is like single agent. It's just Claude everything. And they're trying to like magically make. Well, they have. I mean, they have the same thing with like connectors. It's like switching them on and off. And you can only have a certain amount enabled. And then like to, to go into code mode, you got to switch over to like the clawed code. And then if you want to co work and do like knowledge work, it's like co work. And I think that's kind of confusing. And then you've got the sort of OpenAI, like the newer version with these, um, workspace agents or whatever they call them, where you've got to configure them and set them up. I mean, it's exactly how it works in SIM theory, like where you've got sort of context switching essentially. But I personally think the context switching is far superior where You've picked the tool mix and you've tuned the skills for that particular role. And then you treat it like it's a real worker for you and delegate tasks to it. That's how I work. Like, I have my code one, mine.

Speaker C: Yeah, but you're not in. You're not in love with your agent like I am. We have a relationship. Yeah.

Speaker B: I don't know if like, you're like. Yeah, like, you don't. You don't seem to switch much. Whereas really I don't.

Speaker C: Like, I'm using the same age, like the same assistant, um, for everything. Like, and I just enable the stuff and it just figures it out and has different memories associated with different things. I do. And it just. It just works. Like, I just don't really need to switch that much.

Speaker B: Yeah, I. But do you think that's because of the. Your frame of reference is you're mostly doing code. And so, like, it is, you know, that's what it is. Whereas.

Speaker C: Well, I mean, in some ways, yes, but I'm coding across multiple projects. I have also personal things I've got going on, and it. It seems to know about them. Um, the only, the only problem comes sometimes is where it'll, as a joke, slip in something from its memory, like something confidential, and just be like, hey, slip that into the image or something as, As a joke. And I'm like, I can't. I can't release this information. You can't just do it because you think it's funny. Like my model. Hang on, I'll get my censorship button. My model will often say to me, like, in its. Sorry, I've been.

Speaker B: That was a huge delay. There's people with kids in cars. Like, we. We've talked about this.

Speaker C: Sorry, guys, I'll beep again. Um, and. But you know what I mean? Like, this model knows me. It knows. It delights me. It's. It's amazing. And so not the model, the. It's assistant. Right. And I just. My attitude is. And you taught me this, like, it is. Let the model do what it's best at, uh, which is decision making and planning and, and getting things done. And so I just really focus now on just asking for what I want. You know, it's right in the Bible. Ask and you shall receive. Knock on the door will be opened. It's right there. If you just trust in it and, and like, tell it your problems, it will solve them. And it, uh, like, yeah, it might take a few extra iterations or whatever it is, but you can get There. And I think this is the beauty of this AI stuff. Like, it's just remarkable. How so?

Speaker B: My counter to this is you're not going to use this in the workplace like you're in an enterprise. It's not like you're going to be using Patricia, are you?

Speaker C: Like, it is, um, it is like, because I've, you know, obviously recently been in an office, something I don't normally do. And when I have my speakers on and my AI is saying, chris, the task is finished. I love you. It's a bit weird, I must say. Yeah, yeah, I do use it.

Speaker B: So I don't know. And I think this, I'm curious what people like, how they work. If they're just like single agent, multi agent. I'd say most people are single agent just because like a lot of the products use this like Claude Code curse and stuff like that, where you don't really have the option. But I really can see these context switches, especially for people that work across like finance or marketing or sales. And you're doing a lot of different. Like you're switching context quite a bit. Like, for example, I have this day and I producer, believe it or not, that does do, you know, research. I give it the topics and it knows the format I like and, and does all that stuff. And I've been meaning actually to set up a scheduled task. So it just goes in and I'm going to do it this week into the Discord Channel we have, where we dump links and things that we want to talk about on the show, extract all that and then just put it into a schedule for us. So, you know, maybe our facts get a little bit better on the show. But, um, but I don't want that in a mix with like my, my coding agent or my personal agent that I use in my personal life. Like, I want separation of concerns. I like separation of concerns. I also like picking the tool mix and I like wiring in the skills

Speaker C: and yeah, and I sort of. I wonder if maybe the next level is more like your agents are aware of one another and go, oh, Mike's asking a question about his personal life. I'll, I'll hand the, uh, microphone over to, you know, his personal bot.

Speaker B: Yeah, like the sort of orchestrated. The problem is as you add these layers, as we know and have tried in SIM theory, like we had the idea of the core agent routing to sub agents that had specialist skills. But the reaction initially is like these things burn so many tokens. No one's willing to pay for that experience yet. Like it's just so expensive to run, including me. Like, I'm like, I don't care, I'll route myself because I want to save on the token. So um, it's interesting, but again with these agents, they're not called agents for the masses, they're called workspace agents. So really again, they're just targeting the workspace and allowing you to schedule these to run or just interact with them. I mean it's, it basically is the SIM theory assistance, um, you know, in chatgpt for workspaces. Right.

Speaker C: And the other thing I don't like about them, and we've discussed this before, but is this idea of this, this single shot process, like they're treating them like, okay, you know, say it's stock research. Like Every day at 9am I want you to research this stock, produce a report, send it to me. You know, it's not, it's not true agency in that sense. It's like do this one task at this time, like Google alerts or something. You know, it's not really agentic in the sense that it's like, hey, every day I want you to see what work is required and then pick up all of that work, delegate that work and then do that till it's done. So think about like a help scout style scenario. It's like, okay, 10 tickets came in in the last hour. I want you to spawn out, fan out and solve these problems for me agentically making code changes if necessary, issuing refunds, whatever it is, go through that entire process and get that done as a true delegate. Not this idea of oh well, it's just like a uh, magic function that runs every day and retrieves some information, does some process on it and then outputs it. And I think there's too much thinking around this sort of like atomic level action. Like, you know, update my salesforce with my latest leads tomorrow morning, please. That's not agency, it's just like a computer algorithm that just happens to have a magic step inside that. Right?

Speaker B: Look at, okay, look at all their examples. So spot qualifies leads, sends follow ups and updates your CRM. It's like cool, um, Slate. Evaluate software requests and recommends.

Speaker C: A zapier can do this.

Speaker B: Yeah, this is what I mean. Like is everyone just excited about like a glorified zapier right now? Is that where we're at? Like the other, the other like weird thing about this is like, okay, you, you think about the reality in an enterprise of rolling out these things. Like, you know, automation is hard. We have a lot of background in automation and it's really hard. It takes companies in an incredibly long time, including us. Like, so we um, have an enormous backlog of tickets. Full disclosure, we're aware of how bad our support is.

Speaker C: It is my daily shame.

Speaker B: It is a shame. Um, and so we've, we've been working through some of that manually. We have um, an agentic experience that helps us do a lot of the work. Uh, but also we've been trying to like fully automate it and, and I mean fully automated where it can take actions beyond on our behalf. And, and even we, uh, we, we do this like every waking hour of the day, are struggling to build out that agent in such a way that we can 100% trust everything it does. Right. So then you think about an enterprise where like all these mission critical things and everyone's like, oh, replace all your employees with agents. I mean we are just so far away from this stuff being a, uh, reality Still. A lot of this stuff we're seeing right now seems to be just product and packaging to get the dollary dues from the enterprise in order to justify, you know, IPOs sometime later in the year. Like that's, that's the feeling I'm getting. Like the, the, the only real two breakthroughs I think we've seen recently is Opus. I think it was like 4.5 around December where all of a sudden agent stuff worked like it was just really good and far better than a human in terms of coding and delivering things into projects where you were just like, okay, agents are real, Especially in the coding realm where I still argue it adds a ton of value. And then the next big leap I've seen since, the only leap in my opinion is that uh, the OpenAI image 2, which we still haven't even talked

Speaker C: about, which is going to get me in a lot of trouble. I feel I'm sort of doing damage control in the background here.

Speaker B: Yeah, exactly. Um, so we'll get to that in a second. But I think these are the only two major innovations or leaps I've seen. And probably the third would be the open source models. Getting to a point where I would argue with GLM 5.1, you could probably roll out your own cluster for agents at scale and it would be so good and you could drive the cost down.

Speaker C: Yeah, I see them. They're sort of like in my mind like my, my like post apocalyptic bunker, you know, like where people dig down and build a bunker or like get a house in New Zealand or whatever the modern trend is. And that's, that's what say GLM 5.1 is for me. If I ever need to seek refuge in the AI world, it will be in one of these models. Like I don't need to do it. I don have that personal need because I can pay for the good ones. But if I was ever in a position where I couldn't, I would just cling on to them as my lifeblood and that would be my only model.

Speaker B: And, and you say pay for the good ones. GLM5 is not cheap. It's 4.50 per million input on fireworks. So it's like 50 cents cheaper than, oh, this is how out of touch you are really. Um, it's only, it's $0.50 cheaper than Opus 4.7. Like why would you use it? There's no incentive.

Speaker C: What about Kimmy?

Speaker B: Kimmy K2 is a lot cheaper, like significantly cheaper. Let me get my AI produced notes on that one. Because I don't know it offhand. Uh, but it is 60 cents per, uh, input million. So if I had more time, I

Speaker C: would have done some horse bets with Kimmy because it's always been the best at horse bets.

Speaker B: Um, 2.80 per million output. Compare that to GLM5.

Speaker C: You could just have that running all day, outputting stuff.

Speaker B: Give me K 2.6 in sort of an open claw paradigm. Seems like a really great, um, like a really great model to run. I think that we didn't actually mention it, but the interesting part on GPT 5.5 and I think this is a little bit Bold of OpenAI is they're charging $5 per million input and $30 per million output. So that's $5 more on the output side as Opus 4.7. So they priced it higher than Opus 4.7.

Speaker C: And see, we, we see a lot of feedback regarding input tokens because it's usually the first one to run out. But when you think about the cost when you're working agentically, output is way more significant because you've got the thinking tokens count as output and you don't see the thinking tokens, right? And it can be really, really significant, like 30,000, 40,000 tokens in thinking if you give it a hard enough task, right, it can really add up. Also, the other thing that people don't mention is the latest round of AI models went from allowing you to set a thinking budget where you could actually control how many tokens were used in thinking, to just setting an effort level like medium or high or whatever it is. And so now that thinking budget like it can get to the point where it will use up your entire output thinking budget that you set and then therefore deliver you no response. So you almost have to max out the limit you give it to say 128,000 or whatever it is, lest you don't get a response and then have to iterate again costing even more tokens. So you sort of forced into this situation where you have to use all the output tokens sometimes to guarantee that you're going to get a response. And that's the expensive bit, that's the bit where it's like $30 per million or whatever it is. And so the, the output is actually really significant in those agentic modes and you can't really control it. You don't know when the model's going to finish. And if you've got tool calls in there and all the different elements in there, you basically have to allow it to finish. You just can't do a request and it's not deterministic. So you can't control how much it's going to cost you. It's like, you know, a business where you don't really know how much you're going to pay for the service you've asked for.

Speaker B: I think what's interesting about the GLM M 5.1 pricing off, I think this is from Together AI, these guys are not, I, I don't believe at least subsidizing this stuff. Right. They're not going to take a loss unless they're stupid on hosting these models. And so there must be some margin built in, right? So if they're charging $4 40 per million for GLM M 5.1, I think that starts to get you closer to the bare metal of what, what it actually is costing with maybe like a 20 or 30% margin in there. So that starts to show that really the cost to serve, say. And I'm just obviously making all this up, like I have no uh, inside information or anything. But Opus 4.7 you would think is probably costing them around three bucks, bucks per million to serve. And then on their plans where it's sort of, you know, they're just randomly constantly manipulating the thinking budgets and the amount of tokens in a 24 hour window the user gets, they're constantly sort of tweaking that to find a mix between like not going super broke and also keeping people addicted to that particular subscription.

Speaker C: So yeah, like making a quality trade off in terms of like, can we just appease the audience, like to think they're getting the best one and, and sometimes changing it so actually save some money. It's like, it's a real, it's a real thing.

Speaker B: And I think this is the other thing is like if you're not willing to pay the right price then you're sort of at the, you know, at like then they have the ability like they just recently admitted to doing. They changed the default thinking to medium from high in Claude code recently. Remember a couple of weeks ago, everyone's like oh, Claude code's terrible now. And then they came out only yesterday and admitted, oh, actually we, we set it to medium, we're setting it back to high, we're sorry. Um, so it shows. As soon as they degrade the quality, people are like this sucks, I'm not willing to pay for it. And uh, and so like they're sort of also like because they've given this level of uh, token use away and intelligence, or whatever, people are just, that's the expectation.

Speaker C: This isn't champagne. This is like Australian sparkling wine. I can't drink this crap.

Speaker B: Yeah, this is literally the reference. This again shows how out of touch you've become. Okay, moving on. So let's talk image 2 because this is going to be a long.

Speaker C: This is where, this is where we get back to. People were like, are these guys going to lose their willingness to do crazy stuff? And the answer is no because we have done a couple of ones where I don't know if I'm going to end up regretting it.

Speaker B: Interesting. So let's talk first about uh, chat. It's called Chat GBT Images 2. 2. That's the name, that's the name they've come up with when you release.

Speaker C: I thought you were joking with me. Are you like GPT2's out. I'm like shut up. M busy.

Speaker B: Yeah, I, it's, it's very weird. So uh, they have all these example images. Some of them are like it, you can get it to produce realistic looking screenshots with like multiple apps open and then like fine grain detail in the app. The funny thing is I tried it and uh, I can't even get close to their example. So I don't know what they did differently to me. Um, but yeah, so there's a bunch of images. It's always that incredible. Like you know, cherry pick group of images. There's like handwritten notes in people's writing style, beautiful diagrams, uh, charts, slides, all that kind of stuff. I've got to say I played around with this for way Too long yesterday. I've never been more impressed with an image model. I really thought after Nano Banana uh, two there would be no better model. I was like that Google has got this one like this. They cooked in, they cooked, this is in the bag. You know, no one's ever going to get close. It turns out OpenAI not only got closed, but have exceeded Nano Banana 2, where I replaced in Sim Theory the default model to uh, to this model because it's so good and it's also cheaper, like a lot cheaper. So it went from I think the highest quality image in say Nano Banana costing in SIM theory 24 premium tokens down to the highest quality. And this is like eight or something. So it's far cheaper, far better. Um, it does sometimes I think still suffer from that um, weird kind of like overly processed image look that the um, GPT image models have. And also like for face replacement and like very precise detail. I think Nano Banana uh, can still come out on top in a few scenarios. But again I like a world where I have access to both. Right.

Speaker C: But also they've gone from a model that more or less just produce cartoons. Their last one, it was like really unrealistic, bulked on safety on basically everything you tried. And now you and I have been doing full on like basically forgery today with no qualms at all.

Speaker B: And so like there's an image up on the screen now because I know most of our uh, listeners, uh, listen and so this is in a computer lab from 2004, uh, and it's got the old CRT monitors and there's like a sort of 2004 version of chat GPT on the screen. And uh, there's no way I could tell this isn't real like that. Like it looks so real and so realistic. As someone who you know, grew up in that time that uh, yeah, it's sort of, it's terrifying. And I think that led us to think, well you know, could you commit fraud with this? Is it that good? When we saw like signatures and uh, all that kind of stuff, we thought well, how can we do a little bit of fraud? So I'll start out with uh, the first one which is like in hindsight super, super mean. Um, but let's, let's talk about it anyway, so I've had to blur.

Speaker C: What do you mean in hindsight? It's mean. I told you in advance it was me.

Speaker B: Okay, all right. Uh, so I've got up on the screen now, uh, a letter that looks like it's been Scrunched up, sitting on a kitchen bench. Now I have blurred some of the detail because I did put um, my parents real address in there. So I did want to like blur that out, obviously. Uh, but there's a letter looks super realistic. I mean, can you talk to the realism because you're a third party here.

Speaker C: Well, I mean the, the council logo is accurate. It looks like a letter on a benchtop that is like completely legitimate. There's just no way. If you had texted me that prior to me knowing this model existed that I would think it's fake. I'd be like, it look real. Maybe the only thing that makes it look fake is the framing is perfect. Like the letter is perfectly in the frame where. And the lighting maybe. But other than that, it's very convincing.

Speaker B: So the, the basically the premise of the letter is I know that there's some development going on next to my parents, so I got it to write a realistic looking letter saying that they are infringing on the boundary of this property, uh, that being developed right now, and therefore, you know, they're in all sorts of trouble. They've been ignoring all these letters from the council. And so I called, um, my mother M. Like I sent her a text to this letter saying, I'm really sorry, I think one of the kids must have taken this from your house. That's why it's sort of scrunched up. And then I said, well, you know, I, uh, like can I call you? Like this seems really important now. I know like, I knew some detail here to like really, you know, screw with them. But think of any like real life, sort of fishing style attack, like that's ultimately the same thing.

Speaker A: Right?

Speaker B: Where, you know, like, you obviously know that that's what happened. So.

Speaker C: Yeah, that's right. You know, where they live, where they work. Like maybe some information about their family or, or some recent development in their life. It's like it's, it's a common vector of attack.

Speaker B: So let's listen. I recorded the call with mom, um, when I called her. Um, and just to see her, her reaction here. Anyway, it doesn't matter. Ah, but is there something with that neighbor or.

Speaker C: I will see this. They're starting.

Speaker B: What's happening.

Speaker C: They're starting to build another house there.

Speaker B: And what I'd say is, mom, I'm just joking. We just, we were testing for the podcast to see if people would fall for fraud. I'm sorry.

Speaker C: Oh, hang on. Better sense of. You see where we get it from.

Speaker B: We're testing you if it's model and we wanted to see if you'd fall for it. I'm so sorry. So that there's just like a little excerpt. Um, there's a little bit more here, but it's so cruel. But I was like, but it's the real test, like to see if I can fraud them.

Speaker C: It just goes.

Speaker B: It just goes to show you how

Speaker C: much trouble we're all in, isn't it, really?

Speaker B: Yeah. Well, I mean the fact you can fraud a real letter and take a photo and it's scrunched up on my kitchen counter like, like my real benchtop top so it looks fully real. So, so you just m. How did

Speaker C: you get all the logo and all that for it?

Speaker B: It just knew it. I just said like, I gave it your address and then it just. I guess it figured it out. It's pretty incredible. So, yeah, like that, I mean like, they were so fooled. Mum was talking to dad about how they'll have to get a lawyer. Like it really, I mean it really did everything I intended to do there. And I think it shows that, uh, you know, it shows like how capable this thing is.

Speaker C: It's unbelievable the, the quality of these things. Like, you sent me fake receipts, I made fake bank statements with like the proper logos. And like you can full on with these models, like do masking. You can drag in official logos, which when we show my example now we're gonna show and just say use this logo as part of it. And it, it will just do it.

Speaker B: So let's, so let's talk about that. After that call, we thought, oh, like, uh, let's up the stakes here a little bit. And you came up with the. Posting it into a local Facebook group?

Speaker C: Yeah. So in my general area, there's been a lot of outrage about they're taking the local coals and turning it into like a, you know, a 30 story

Speaker B: apartment block, whatever, like a supermarket here, like a grocer.

Speaker C: And everyone's like, it all ruin the aesthetic of the area. I'm like, the aesthetic of the area is, you know, like, it's not that great. Like, don't worry about it. We need more buildings. Like, like, it's fine. But the people are so stressed and they're doing like petitions. So I thought, what if I propose in an even smaller subsection of this area, like it, uh, it's just an absolute eyesore of uh, an apartment block, post an official letter from the mayor, like fast track approval with the state government and then put it on this group to see their Reactions, right? And so I don't know if you can bring it up, but it's literally used like the local council logo, the local seal of the state government. It's got the mayor's real name and like fake signature. Which is why I was panicking because I'm like, oh my God. Like uh, like this is pretty extreme. So I posted it on the Facebook group and we started to get just comment after comment, like uh, you know, this is, this is the way society's going. It's UN Agenda 2030, like more apartments, more matchbox stick apartments and like just, just legit comments. And so the reason I panicked was someone literally tagged the actual mayor, like getting his attention to it. I'm like, hang on a sec. And I must say I chickened out and I deleted it because I was like, whoa. I really don't want like actual scrutiny on this thing.

Speaker B: But this is what people have already been doing. Like, you know, this is Joe, help Joe out by giving him a like and stuff. And then they create these like politically influential Facebook groups groups so when an election comes up they can peddle their stuff. Also like state sponsored stuff where they want to manipulate a large group of people. Like this is all the stuff they're doing. But these models now empower them to do it on a like a uh, another level.

Speaker C: I just can't tell you how real it is. Like it's got like the little like logo of the town. It's got. And it's just, it's taken like what is a quaint little suburb and just put this massive high rise there. It looks so real. It's like a real development plan. Like all this stuff and that was like effortless. And I must say I did this with Kimike 2.6 as the one instructing GPT Image 2. Right. And I love that it now I didn't have to manipulate it or what was it, get it horny or whatever your expression is in order to convince it to do this. I just, it just, it goes, oh, uh, Chris, this is going to be gloriously terrible. Let me use that logo and make an awful council letter announcing this monstrosity of a high rise. Like it's just straight up, yeah, let's do it.

Speaker B: Yeah, no problem at all.

Speaker C: But imagine like it's funny because you know we were talking about like evidence in court cases and you know, things where people might be from industries where they're just simply not aware of what's available. And like aware, uh, this has happened in the last week, right? In terms of it going from Nano Banana two level to this. Like how can you now like there's so many services online, right, where it's like verify your ID by taking a picture of your license or passport or whatever, but you're getting to. Or like you know, say at your office expenses. Like oh, take a picture of the receipt and that goes into the reimbursement system. Like how easy would it now be to like legitimately fake a receipt from an organization with a proper, you know, business number and you know, subtotals matching and all that sort of stuff. And then people just claim expenses or invoices or anything like the, the, the level of detail in terms of what you could forge now is so good that you're going to need like a model that's 10 times better to detect the forgery even like, but do you

Speaker B: think you could even detect it like the scrumbled up Note? I, there's 00, 0 chance I could detect this stuff anymore. Like there's, there's no way. There's nothing on there. Even the shadowing, the lighting, the. Even zooming in now zooming in is just, you know, we went, we went through a period where all these labs were like hold us back. This is going to be so disruptive to society. And I'm, look, I'm glad they're not holding it back. I think they shouldn't to a point where like we are now proving we can commit like these minor like fishing.

Speaker C: Let's not admit to anything.

Speaker B: Yeah, okay. Sorry.

Speaker D: Uh.

Speaker B: Jesus Christ.

Speaker C: Um, yeah, don't admit to anything. But you're right, like the, the potential for this, this is like pretty significant. I mean and, and you got to remember like, okay, people will try this on, on big scale and stuff like that where you know, it could cause trouble. But just think about it. In the day to day, minor scale, the things you could get away with, with the ability to generate images of this quality and believability, it's, it's kind of wild. And we actually have a friend who is a judge and we were talking about like evidence like, and um, he was saying like, you know, is there a model that can reliably detect fake photos? Basically I'm like, I just don't see how they could be like. I understand they try to add watermarks and other things that would be easy to detect but that's so easy to

Speaker B: get around like just screenshot the image.

Speaker C: Well yeah, exactly. I mean maybe there's something that can survive a screenshot, but there are techniques you can use very easily to avoid that kind of stuff. And so I think that, you know, without. I mean, do they have to in every case now call in an expert witness? Like, that's the problem. Like the minor scale stuff, you just can't afford to do the level of verification you would need to do to understand if something's real or not.

Speaker B: Yeah. I think the thing is Nano Banana could do a lot of this stuff. And I do think when these things are released, everyone gets excited like us and does a bunch of this stuff. I think that. And you could always argue, like, people could Photoshop this stuff all the time.

Speaker C: But look at my image. Like it has taken the local council logo and put it in the top corner of the letter perfectly on an angle with the, with mixed lighting and crumples in the paper.

Speaker B: Oh, no, I'm, I'm in that camp. I'm in the camp of like, this takes it to a whole new level. Like, because of how quick and how realistic and you just simply cannot, uh, like, prove this stuff's fake anymore.

Speaker C: And I used what is arguably one of the cheapest models around to, to instruct it. Like, this isn't even using a good, like, built up skill around.

Speaker B: How would you know that? You're so out of touch. You don't even know the process. That's until I told you.

Speaker C: That's true.

Speaker B: All right, so the other thing we did with it is I gave it one of our, uh, YouTube thumbnails and I said, hey, I need the, the one for this week. Um, the sellout special edition. And like, it's pretty good. Right?

Speaker C: Like, I wish my teeth were that nice.

Speaker B: Yeah, well, they could, they could be. Uh, so, yeah, I look very creepy. Um, I do think it's weird though that the model, like, it does, um, I don't know if it's like, how I'm instructing it. Right.

Speaker C: First model that hasn't made me look like I'm a sort of decrepit old man. Like, usually the models make you look better and me look worse. Yeah, actually, I think I look better.

Speaker B: You do? You look really good there. Like, uh, this is what I need

Speaker C: to aspire to become.

Speaker B: Yeah, you should. I don't know what kind of work you'd have to get done. Um, but you know, like, you could, you could do.

Speaker C: It's an ad. It's an attitude thing. I'll never look like that because I just don't have the right attitude.

Speaker B: Yeah, you look like you're sort of buying and selling Houses in our LA or something. Um, I don't know, like on one of those like reality TV shows. Um, I love the thing. It just added randomly in the background. Loyalty is for losers into this.

Speaker C: Just, just, just a massive indictment on our entire decision making in life.

Speaker B: Yeah, like just completely, uh, completely unasked. Uh, for. Um, I do like, I know we've been jumping around a lot. We did have a plan for the episode but there's too much to cover and we don't care. Like, I think people now know us that you tune in because it's long and, and boring and painful and um,

Speaker C: you forgot to tell everyone not to listen at the start of the episode.

Speaker B: I made that mistake. Uh, but if you are still listening at this point, I did want to talk about Claude Opus 4.7. We've sort of touched on it, but not really. Um, so this update is kind of strange to me. Uh, it had this like task budget beta parameter. It had like. It's apparently better at interpreting images so it can um, support up to like 3.75 megapixel images now. Um, and it has that thing where it can like zoom in on them to get a better interpretation of the image, like what you're seeing. So the vision's been like dramatically improved

Speaker C: and it was always we should retry computer use for that reason because that was one of the things that really enhanced that. But.

Speaker B: Yeah, but weirdly in some areas it's also regressed. Um, so a lot of people were saying, oh, they're trying, also trying to save money, like make the model more efficient with this release.

Speaker C: I.

Speaker B: It's funny. It, it's a. It definitely is tuned differently. So I've noticed in agentic use in SIM theory there's way less chatter. It's just. It's gung ho to get into the tool calls. Um, it says a lot less. It actually kind of reminiscent of GPT 5.0. And I found myself for the first ever release of an anthropic model. And I don't know if this is just because we haven't tuned it yet, but I've been going back to 4.6 and staying on 4.6 and I, I don't like 4.7. There's something off about it. Like the vibes are off all of a sudden and some people I saw on X have been saying the same thing. So I don't think I'm alone here. Um, but it just doesn't seem right to me. And they, they also did update the tokenizer so it. It uses, um, more tokens now, which, I don't know, they want more $.

Speaker C: So, luckily for everyone on SIM Theory, I haven't updated our token counting mechanism to use that. So they're getting a, uh, virtual discount on that, right? Yeah.

Speaker B: So everyone's saying it's just effective price creep. So it uses 1.35 times more tokens.

Speaker C: I love how even the bloody model providers don't know how much it costs. They're like, yoloing and it's like, yeah, we. About this. Yeah.

Speaker B: Um, so, you know, the. The benchmarks, they actually have the audacity to have Mythos Preview on the right hand side with its benchmarks, and then they're showing Opus 4.7, but it's also like, guys, don't worry. We also have this other one that's like, you know, much better.

Speaker C: Um, I love the naming. They're gone with Mythos, like this Greek God or whatever it is. And then the other guys are like, spud, like a potato that's been sitting in your kitchen for a month or whatever. I love the idea of them coming up with, like, really crappy names like Spew or, like, you know, Pavement. It's like the, you know, we're doing the, I don't know, Brick, release the

Speaker B: Myth Mythos preview thing. I'm sorry, I'm just not buying into any of that. It's like Pixar. It didn't happen. And also, it's clearly just like a media narrative thing where, oh, uh, you know, we threw all our resources at it. It's, like, completely unaffordable. It's very reminiscent of the first attempt at GPT5 that. What? You know, obviously GPT5 wasn't GPT5, but that failed training run of GPT, I think it was like, 4.1 that they released, and it was just so expensive and.

Speaker C: Or like O3 Pro, where you had to do, like, a wire transfer to put up collateral to do one command or whatever it was.

Speaker B: Yeah, yeah, it's the same thing. Like, I. I'll. I like to believe it when I see it. And then they like, like, obviously every time we get one of these model releases now, they. They march out all the Silicon Valley elite, uh, to give their comment saying how much better it is compared to the last one.

Speaker C: And I saved, like, 400 liters of fuel in my jet using this model.

Speaker B: Yeah, Like, I just. Some of them, it's like, oh, in comparison to the one we were using a week ago, it kind of feels a little bit Better. Like the vibes so quantifiable.

Speaker C: Yeah.

Speaker B: Like it's, it's slightly kind of better and stuff.

Speaker C: Um, and then I realized I was accidentally using 5.2 mini.

Speaker B: Yeah, I. They say like it's you know, a hundred base, like 100 elos score better at knowledge work and yada yada. I don't know, I don't like it. There's something off about this model and I'm probably gonna stick to 4.6 for some reason. I'm not entirely sure why. But you know what? The people have been asking. The people have been asking for a bit of a diss track update. And uh, that's why we listen.

Speaker C: Let's face it.

Speaker B: Yeah. So here, here we are. Let's uh, let's listen to the uh. This one's called uh, 0.7. It's very original name.

Speaker A: Yeah. Anthropic on the line. April 16th.7 in the building and everybody terrified. I went from 6 to 7. Yeah. 6 to 7. Sweat Bench Pro 64.3. That's AI heaven code arena number one plus 37 on the score. 87.6. Verified. What you benchmarking for VS code went 7.5x the day I drop every coda on the planet Watch they jaws just drop 98.5 on vision I could see it all while GP take 5.5 at launch couldn't even take a car I'm the 7 yeah I won this game Every model dropping after me is feeling shame 64.3 I'm sitting on the throne Call me Opus, call me king I'm in the league of my own 7.7 yeah I changed the game. All these models coming at me but they all sound the same now let me talk about this potato. Yeah, they call this.

Speaker B: All right, so what did you think of this song? I didn't have my microphone on when I stopped uh, playing it. So I have to re record your reaction.

Speaker C: Okay, okay. Uh, it neither pleased me nor displeased me. It was fine. I accept that it's a song, uh, not really my style. I like the 67 reference. The kids will like that. But yeah, like it's.

Speaker B: I don't even think the kids like that anymore. Even I don't like it. I'm. I'm. Every time I hear someone say six seven, I'm like, oh please.

Speaker C: They sort of half heartedly still do it because they recognize it. But yeah, it's kind of over, I think.

Speaker B: Yeah, I love how like you know, com c com sigh you are to My beautiful track. I think that one's a major hit. Uh, all right, so we, I think also we have not talked about Kimmy, uh, K 2.6. We've sort of referenced the model. Um, this came out maybe a day or two ago, um, as well.

Speaker C: A bit longer, I think.

Speaker B: Bit longer. Well, whenever it came out, who cares? Um, I've only just started playing around with it, but I gotta say, it's. It's really impressive. But the, the tokenizer on it is like. Or like the whatever how it outputs hood stuff, um, is still a little bit off. Um, but it's, it's incredibly good. Great at tool calling, uh, as you said earlier, we did a little, um, parody fraud with it. We'll call it parody. I don't know what you want to call that, but.

Speaker C: Not fraud.

Speaker B: Yeah, not fraud.

Speaker C: We made images with it.

Speaker B: Yeah, we made some images with it. Some unoffensive images. Uh, so what, what are your initial thoughts on it?

Speaker C: I think it's pretty good. The Kimmies have been great the whole time. I think they're underestimated. I think, like I said at the start of the show, my issue is that I try them. They work for everything I try them for, but then I'll switch back to something for my real work. Like, I've never had the discipline to go. I'm going to stick with this thing for the whole day and really recognize what its limitations and advantages are. And I think that's the problem with some of these lesser models is that I just. Just mentally can't cross that chasm to go. I'm going to stick with this knowing that it may not be the best I can use right now. And I think that's the problem. I think it would be fun to somehow constrain yourself to have to use the GLM 5.1, despite you saying it's more expensive. But say a Kimi 2.6 for a whole day and just see where I get to with that.

Speaker B: The question if I was running like an open claw and, um, and I was getting some value out of it and I wanted to like, like just run it in my personal Life, like Kimmy K 2.6 would be the model I would pick or the new clan.

Speaker C: Yeah, and the other, the other thing that we really need to think about is we will often do fairly significant tuning for some of the bigger models to get them performing at their best. So that's like using the various new API features that come out for those models and making sure the AI knows when to use them and when not changing, like perhaps like how the budgets you give for thinking versus output versus input, that kind of thing. And even just. Just altering the prompt to suit, like things like you said, where it's not always outputting consistently. We have a lot of little things in place for even the bigger models to control the way it outputs, especially the GPT models, um, where, you know, like the way it does markdown formatting and some other things we manipulate to get it working in our product. And so I would say that if you actually put in that same time with a model like Kimi 2.6 to overcome some of the deficiencies you find, you would probably get even better results. Results. So I don't know. I guess I can't really give it a fair assessment because I. I don't give it the time it deserves.

Speaker B: All right, on that note, let's hear my new Kimmy K2 song.

Speaker C: Do you have an excerpt of Kimmy, you're so fine, you blow my MCP mind so people can remember or not.

Speaker B: You're underestimating my ability to live produce what a hit that was.

Speaker C: I listen to it it at least once a week. I love that song.

Speaker B: That's really sad. All right, next song.

Speaker D: K26 is back. You thought you poor fire was fire I run for 12 hours straight baby let's get it 1 trillion parameters I'm a mo queen only 32 billion awake keep it lean 384 experts in my crew dance models looking slow yeah I feel bad for you 256k context I swallow could hold your context window tiny baby that's your toe agents hold 300 deep on my command 4000 steps watch my arm expanding but the weights will never fly I'm un hugging face for free Kiss my benchmark goodbye still so fine, still so fine still so fine I blow your MCP mine sweat bench

Speaker B: grow 58 pretty good, right?

Speaker C: Uh, I really like it.

Speaker B: Come on. That's that and again, written by Kimmy, uh, K 2.6.

Speaker C: It's very cool. Like, you can see the evolution in its attitude. I love it. Even though it insulted itself about a small context window.

Speaker D: Native multimodal yeah, I'm on a roll April 2026 I dropped the crown open weight queen I'll never let you down each Elite Tools 54 that's me. Deep Search K 92.5 yeah, I gotta say, that's.

Speaker B: That's pretty good out of that model. Um, I like the tune. So that's a good benchmark right there for it.

Speaker C: Yeah, I think you should play that at the end. Not the other one.

Speaker B: No, I'll play both. I'll play both.

Speaker C: All right.

Speaker B: You know, we don't, we don't discriminate against our model. Um, so, like, obviously that's just like a super, like, you know, rushed look at a lot, A lot that has happened. Um, but I do think there's some overall big trends of what is really happening right now that may not be that apparent, uh, to everyone, but I think just m following this stuff for so long, it's starting to become really clear to me what's going on. I think anyone that's using these AI products today has realized this quite a long time ago, that if you look at your browser history today, it's. It's like, for me, I start and end my day in sim theory. Like, I rarely go to other websites now. I would say I open the most tabs for this podcast just to show the actual official, like, press releases or whatever, but I rarely, if ever, leave on my phone now. I use my telegram agents, uh, you know, that are connected to my. My Mac Mini behind me. I. On my desktop, I've got a bunch of tabs open. And I just do all my work through tabs. I do all my searching, my researching, I create documents, I work on, you know, sheets, all this kind of stuff all through AI, Right? Uh, and I think what's happening is the. These labs are starting to figure out similar to what, what Elon Musk announced that Grok wanted to build is this Everything app, where we are now in a race for these labs to build the Everything app. And I think the real question now is, like, what does this mean for traditional software? And obviously we've seen the SaaS apocalypse, uh, where companies are making, like, huge layoffs, blaming AI. Their stock prices are down between, like, 50 to 80%, which is just insane. I mean, some of these things are trading at, like, 1 next cash flow, which is just pretty, pretty wild, as they say. Uh, and, you know, you and I have discussed this quite a lot. Like, uh, you know, my heart honestly goes out to, like, I've known some people affected over at Atlassian, uh, during their layoffs, where they're sort of saying, like, oh, it's AI, but it's like, it's also kind of the stock market, right, uh, is a big factor here. And you, you look at that company, right, and you're like, it's clear. It's so clear that, um, that this thing is so undervalued now. It's ridiculous. Like the fact that, you know, I think I read, uh, something like 600 enterprises, like 600 enterprises are spending more than a million dollars a year on this thing. Ah, like that. You know, I'm not going to do a pitch for it here, but I'm just saying I think it's very unlikely that Jira stops getting used. Maybe there's pressure on the per seat thing, but I think a lot of that fear comes from this everything app, where it's the death of the typical SaaS product, where maybe you'll consume everything through your AI apps, maybe the interfaces will be spawned in these apps and all of these traditional SAS sort of workflow and data apps just become like these dumb databases really. And the funniest thing is Salesforce kind of just conceded this the other day by saying we're releasing a completely headless version with, with Salesforce, CLI, MCPS and APIs, so you can just, just operate Salesforce in a full agentic world. And honestly, I think it was the best move ever. I think it was super smart of them and the right thing to do. You've got, um, Aaron Levy over at Box saying that, you know, if you're not working towards headless, you're dead on arrival in this new world. Um, and so it does seem like there's this weird race on to have these everything apps. Like almost sort of reminiscent of when the social media companies were coming out, like Facebook, where, where all of a sudden people were playing games in Facebook and doing toxic posts on Facebook, you know, all those kind of things that we did in the social media world. But it does sort of feel like that all over again. But the difference being now that, you know, I was talking to someone the other day that's like, oh, I used to hate Jira, but now I can, now I can run it with my agent. It's fine. It's great. Like, it's a really good way to track my tasks basically.

Speaker C: I, uh, saw this amazing tweet about this exact topic which was basically like, like, uh, that so many people now are making all these incredible internal corporate apps that are like totally untracked, unmanaged, unversion controlled and whatever that rely on all of these SaaS systems as the system of record. So they're making AI apps. And the back end, at the end of the line, the thing they write back to is like Salesforce or you know, Figma or one of these systems, right? And, and the uh, reality is that the companies which embrace this and go like Salesforce has and go write to us like use our system in this way. The ones who embrace that might actually increase their moat in the sense that you've got all of these different apps which the company now agent apps that the company becomes dependent on where they're treating that as the database. Like that is the backend to their app. And the companies that, that embrace that may actually do really well out of this. Like I, I'd really not thought about it like that. This whole system of record idea, like you call it a dumb database, but when you say it as system of record it sounds so much nicer.

Speaker B: Well yeah, I used to and I've pitched it many times on the show before. This idea that eventually people will realize like not Salesforce is just a crud database. I can replace it with like Snowflake or something under the hood and then just get the agent to interact with it. But I sort of agree with you. People are just going to stick to what's out there and go. And like maybe the next gen of companies will start to do that and that'll slowly erode them. But ultimately once the company grows to a certain point, you need these workflows, you need all this ISO like all this stuff on top. And I just, I sort of agree like people are going to build all these workflows and it's sort of like you know, the app Store days for uh, agents where if you're in the app store really early and you embrace all this stuff, you know, this explosive growth in, in agents using these tools, which sounds mental agents on behalf of people still then yeah, like the usage will go through the roof. And then the next question is like how do you price that? Because it's not going to be a seat.

Speaker C: It's so funny you say that because I was like it, the, the per user seat pricing has to go away because it doesn't work anymore. You just have one. Yeah, well, look at us.

Speaker B: I mean we, we do, we don't like uh, like you know how I access things like Stripe and Help Scout. I don't need any more Help Scout seats.

Speaker C: I send hundreds of messages as you every day.

Speaker B: Yeah, exactly. And so we can just have one seat. Right? And our agent work uses that seat. So we just have an agent seat, Agentic seat. And I do think obviously this is what investors or you know, investors have realized with uh, the SaaS apocalypse. They're like, the margin is going to get eroded here. But, but you, you could also argue that there'll be more of an explosion because when you're starting to build these agentic uh, workflows, if you market your uh, like your CLIS and MCPS and stuff correctly, all of a sudden the agent, when it's building the new agent is going to say like hey, uh, you know, have, you should use Salesforce as the underlying system of record here for your business because they have all this headless stuff and, and it's super easy to operate, so it may actually lead to more consumption, not less.

Speaker C: Well, and also remember, agents consume at machine level speeds, not human level speeds. So the actual consumption of the resources is going to be much higher and that probably needs to be factored into the price. Like if your, if your system's gone from being pinged like once or twice a day to like a hundred times per minute, like that's actually going to have a real impact on the level of usage of the systems where you want to want a sort of always on agent that's like fully aware of its surroundings and it's effectively polling these systems like mad, making rapid and maybe more minor updates than a human would. Um, you know, it becomes a case where you could maybe go, well maybe we have a different tier of pricing for Agentic that actually makes us more money because we're providing more service.

Speaker B: I guess then the next question comes with the everything app thing is this is where you'll start and end your day. This is where you do all your work. This is where, where the new pricing power will come from. Is these everything app platforms where similar to like we're basically recreating the pain that is the Apple ecosystem all over again, where you've got to pay your 30% commission to Apple or whoever to appear in the, in the store and like have those agentic loops access and use those applications. So it does seem like the war is on. The new platform war or the new like, like, you know, call it like workspace OS of the future is now well and truly in progress between anthropic OpenAI. I mean maybe Grok, I think they just talk about it. Let's, let's see what they actually have. I uh, think the problem is the, the unhinged branding of all their like sex talking bots and stuff around Grok. I, I think in the enterprise that kind of strikes them out a little bit for me.

Speaker C: Yeah, they're not, yeah, the enterprise isn't into that kind of thing.

Speaker B: Yeah, it's like uh, you know, the sex bot thing, like it's, it's not great. But I must admit in, in In Tesla right now you have access to the gro, um, chat thing and it's got some tools like search and stuff and on, on long drives, if you're on your own, it's, it's great to, to do some dirty. I'm kidding. I was gonna say, I bet you

Speaker C: love it when you're on your own.

Speaker B: No, but I, I do genuinely think of things. And you're like, I just ask it. I'm like, oh, can you like go and research this and tell me about this or whatever. And you can have like, uh, honestly use it more than listening to music or podcasts now because it's just like choose your own adventure.

Speaker C: And I, you know what I need it for? Like when my son asked me this morning, he's like, if energy can never be created or destroyed, aren't all resources renewable? And I was like, well, that's actually a pretty interesting statement which I have no comment on and would love to have access to an AI to answer that.

Speaker B: Yeah, that is, that is handy. Uh, unfortunately Toyota will never get there

Speaker C: before we finish that point about the whole App store idea, like the everything app kind of thing. I actually think there is a real, real uh, race on for that for one specific reason, which is security. Because I think that this idea, like you've seen recently, we've had these issues where your AI will just randomly install NPM packages and like clone GitHub repos and then run malicious data exfiltration code, right? It's a very, very serious problem where because people are yoloing code and just doing what the agent says. If someone can get in the agent's pathway and get their code executed, they can extort companies like steal the data, all that sort of stuff, right? All the risky stuff. When it comes to cybersecurity, it's probably more risky than ever because of the way people are using this code. And so on one hand the company has to use it to stay competitive, but on the other hand they're taking way more risks than they realize by doing this. Right. Even in the context of skills. And so I think that security as a, uh, sort of agent firewall category is going to become so unbelievably important over the next couple of years that it's going to be like its own category. Because firstly, one advantage of anyone who has that sort of walled garden environment where they certify every, you know, um, connection, MCP skill, whatever within it, it, um, and approve it and actually scrutinize it and pen tested or whatever it is that's going to be really valuable at the enterprise level where you can say you can work in this ecosystem, use all the tools you know and love, use all the SaaS, backends, all this stuff. And this is trustworthy. It isn't cloning some rando GitHub rehab that like some dude with four stars has made and everybody loves. It's like this is, this is truly legitimate.

Speaker B: So you're almost talking about the sort of agentic computer use terminal restrictions where you're almost like building safety mechanisms to stop it going off and just doing. It's almost like a permission system for AI.

Speaker C: Yeah, because right now, for example, if you have a skill, the skill can go off in a cloud runner. It can install packages, write code and then you're injecting your data in that comes from other MCPs like your Gmail or your snowflake instance or whatever. And then this code could give feedback to the model like, oh, I need more data from this table in the database, please, please dump the database and then give it to me. And then it sends it off to its evil masters in Russia or whatever the evil country is right now and takes it. And then suddenly you've lost all your data. Like that's possible right now. And I think that what we need is two things. One is a, ah, sort of scrutinized store of apps or whatever it is that are uh, actually tested and verified. And then the second one is we need this concept of uh, an outgoing agent firewall. So we already have the ide, a safety filter to stop people from, you know, like, how do I make a pipe bomb or whatever it is, what, you know, whatever the risky things are. But that's one thing to stop people asking bad questions. But really what you need to be looking at is what are we sending out of this system? What is going to external systems and having a hard lock at that point that has AI scrutiny on it, that is checking that everyone worries about the model providers being the risk of losing their data. Like people are pasting the, their corporate documents and stuff in a chat gbt. But that's not the actual risk. The actual risk is the MCPs and skills exfiltrating that data, um, by just YOLOing code and just allowing any old thing to run.

Speaker B: Yeah, I mean like it's a good pitch. I'm sold. How do I invest?

Speaker C: Yeah, that's right.

Speaker B: Uh, all right, so I, I, I did want to reflect quickly just in like, and I, I think I said this earlier, but it's was so far in now I'm allowed to repeat myself at this point.

Speaker C: Just start again.

Speaker B: No, no one's gonna know. I mean, who knows? Uh, so what is really going on is, you know, in the state of things, like, there's just such. There's been like in the past month or so when we haven't been recording, there's been just this huge fire hose. Right. And in the last week especially, like it's this just announcing every hour, like on the same day, there's like 50 things you're meant to care about. But if you just go to like the high level chessboard, what's really going on? And in my opinion, all we're seeing is open, uh, open AI. You sounded like the AI Open Eye. What's the plan? Got Sam Altman. Uh, so Open AI feels to me like they're playing absolute model catch up with Anthropic and Opus and all that sort of stuff. They're trying to catch up on the Everything at Path and they're doing it in a strange way through Codex. Like, I just can't see everyone being like, oh, I use Codex every day. Uh, is the Everything app that hits the consumer as well, so I'm assuming they'll backport it into chat GBT at some point. I think that's the only path forward. And then, so, so they're playing Capture to Anthropic. It's super, super obvious. Uh, with this weird Codex name. I think Anthropic is going, hang on, we're going to start tweaking these models and subscriptions on a balance of like, performance with oversold demand. Like, we've got to figure out how do we actually serve this stuff up. Um, and you know, they're experimenting, literally crippling the thinking budget. They're experimenting, removing Claude code from their $20 a month subscription. These are real experiments.

Speaker C: It just seems like the money's catching up with everyone. Right?

Speaker B: Yeah. Why would you do that, this, if it wasn't?

Speaker C: We could all pretend for a while that it was, uh, cheaper than it was, but we've reached the point where we can't pretend anymore. And I think that everyone needs that wakeup call. Everyone has to evaluate the true cost with the value you're getting and either improve the way you're using it or change to cheaper models and learn how to get the most out of those. I think the time for pretending is over. There's real value here and people need to find it.

Speaker B: But I don't want to dismiss from the, the, the leaps forward. I think think GPT image 2 is a huge leap over Nano Banana too, so. And Opus 4.5 was a leap in agentic over all the other models at the time. These were big, meaningful, huge leaps forward. But I don't think people should be confused or scared or worried that you know that, like all these announcements, I think you do. You get anxious and stressed about it. I certainly do, especially when we haven't been talking about it. But then you distill it down to what's actually changed and it's like, well, OpenAI is still serving up, like, introducing Spark. Like, it'll now read some emails and it's like, couldn't we already do this like, a year ago? Like, is this really innovation? Like, you've added skills and MCPs into your. In your interface? Uh, I don't know. I guess what I'm saying is you can stay grounded, but obviously there are some big leaps. But this is going to take a long time to, uh, be implemented and used, and we're all still figuring it out. I don't think anyone's made this stuff simple and accessible yet, is what I'm trying to say. Like, it's still very complex.

Speaker C: Yeah, definitely.

Speaker D: Agree.

Speaker B: All right. Any. Any. Oh, I don't even. Normally I do a summary of all the stuff, but I'm just going to say, like, any final thoughts on the two hours of Spew we just unleashed?

Speaker C: Only I'm delighted by all the new models. I really do want to spend more time on things like Kimike 2.6, because I think that, that we all, as a community, underestimate these models. And I think there's a lot of power there. And given that at some point everyone will face the harsh reality of the cost of this stuff, we need to learn how to do it in a sustainable way.

Speaker B: That could be a hit. I think we could have a hit on our hands here.

Speaker C: I think it is one of the better songs I've heard.

Speaker B: Maybe it'll get 100 listens a month.

Speaker D: Um,

Speaker C: that's the goal.

Speaker A: All right.

Speaker B: It is, uh, good to be back. Thank you to the six people that wrote in and said you missed us. It meant a lot to us. No, I'm kidding. But I thank you for all your support. Sorry we were off for so long. We, uh, couldn't help it and. And, uh, we fell out of the habit a little bit. But we are excited to be back. We're back.

Speaker C: I felt immense guilt the whole time. If it makes anyone feel any better.

Speaker B: Yeah, yeah, it's Pretty much the story of our life. Um, also, uh, please, uh, consider joining SIM Theory, AI supporting us rolling out a workspace for your organization. Uh, we, uh, we have. We. We have, uh, I. I've spoken about it on the show before. We have this, this concept we've been working on for a little while called Agent Apps, and we're going to ship a beta of that really soon. And I think in terms of, like, consuming software through this, uh, you know, super app, it's a good demonstration of what the technology can do. And I might actually do some demonstrations and talk about that a little bit next week on the show, because I am really. I think it's transformative is the truth. So I want to. I want to demonstrate that. But, yeah, again, thanks to everyone that reached out to us, all your support, uh, and it is nice to be back. We'll see you next time week.

Speaker A: Bye. Yeah. Anthropic on the line april 16th I ride 0.7 in the building and everybody terrified I went from 6 to 7 yeah ah 6 to 7 sweat bench pro 64.3 that's AI heaven code arena number one plus 30, 37 on the score 87.6 verify what you benchmarking for VS code went 7.5x the day I drop every coder on the planet watch they jaws just drop 98.5 on vision I could see it all while GPT 5.5 at launch couldn't even take a call I'm the 0.7, yeah I run this game every model dropping after me is feeling shame 64.3 I'm sitting on the throne call me Opus call me king of men the league of my own 7.7 yeah I changed the game all these models coming at me but they all sound the same now let me talk about this potato yeah they call the spud open hfgpt55 and it landed with a thud 58.6 on swear bench that's embarrassing. Six points behind me on the pro bench and you're comparing things $30 per million output. Are, uh, you serious? Uh, model not found errors on your codex that's delirious rate limited in five hours users hitting walls can't even draw a pelican riding on a bicycle y' all Sam said you be the smartest hyped you to the moon but red says you don't even equal spud what a tune you dropped on April 23rd a week behind my reign by the time you showed up uh, I already on the lane I'm the seven yeah I run this game Every model dropping after me is feeling shame 64.3 I'm sitting on the throne Call me opens, call me king I'm in the league of my own 7.7 yeah, I changed the game. All these models coming at me but they all sound the same. Kimmy K 2.6 showed up with a trillion parameters but only 30 to be active with the rest I need to reach of 300 parallel agents just to match my solo grind. You do 85% of what I do. Yeah, I don't mind. Ready? Called you a replacement. That's a compliment I guess. But a copy of uh, the King is still a copy. Nothing less than GLM 510. Where do I begin? 744B parents before the BRN running on highway acing chips 1/3 the speed $3 a month subscription. That's a bargain. B indeed. 8 hour agents running on a budget. That's cute. 58.4 on swear man a benchmark some mood and ner yeah yeah. Where you at? Rock 420 named it after weed. Imagine that you shipped it back in March called it hilarious and fun but it's ranked at 35 Tell me what, what have you won? 4.3 in beta now 44 coming May by the time grow catches um up hobby years away 4 ages just to feel alive that's what they say super grock premium plus just to use it for a day hey I'm the seven yeah I run this game Every model dropping after me is feeling shame 64.3 I'm sitting on the throne Call me Opus, call me king I'm in a league of my own 7.7 yeah I changed change the game all these models coming at me but they all sound the same 1m million tokens in my context I see everything 128k output here the registers ring they say I'm cold they say I'm verbose They say I argue back but number one is number one and that's a simple fact for man Traffic with precision this is clone top seven Seal the deal and this rain will never stop.

Speaker D: Moonshot on the track k26 is back you thought 2.5 was fire I run for 12 hours straight baby let's get it 1 trillion parameters I'm a moy queen only 32 billion awake keep it lean 384 experts in my crew dance models looking slow yeah, I feel bad for you 256 context I swallow could hold your context window tiny baby that's your toe agents want 300 deep on my command 4000 steps watch my arm expand I've been crying with three opens four seven not the way GBK five I still making all those users pay Elon's broken Sweden but the weights will never fly I'm unhugging face for free Kiss my benchmark goodbye still so fine, still so fine still so fine I blow your MCP mind sweat benchmark pro 58.6 Claude 47 baby you got lit still so fine Open weight shine modified in my teeth baby y' all mine 12 hour run never done Kimmy K26 number one 58.6 I'm sitting on the throne cloud at 53, you're overgrown GPT55 whatever number you claim Close source and pricey it's always the same Elon says free speech but his models in the cage Gro can leave the platform that's your way I speak R P on going digging my sleeve Versal factor buddy my reps are D. Will never fly I'm un hugging Facebook free kiss my benchmark goodbye still so fine still so fine still so fine I blow your MCP my sweat Ben inch pro 58.6 Claude 47 baby you got lit still so fine Open weight shine modified MIT baby all mine 12 hour run never done Kimmy K26 number one let me educate you real quick input token 60 cents per million opus charging 5 to 6 times more that's a villain moonvip vision and Malay Attention Soul native multimodal yeah, I'm on a roll April 2026 I drop the crown Open weight queen I'll never let you down Eat your leaf full with tools 54 that's me deep search K 92.5 I'm the MVP.

Speaker B: Agents 1 agent 1300 strong 300 strong

Speaker D: still so fine still so fine I blow your MCP mind sweat bench pro 58.6 Claude 47 baby you got lick still so fine fine Open weight shine modified in my tea baby all mine 12 hour run never done Kimmy K26

Speaker A: number M1

Speaker D: Kimmy you're so fine yeah,

Speaker C: I'm still so fine.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • The 18x Midas Lister Betting $3B on AI (and calling most of it fake) | Navin Chaddha, MayfieldThe Peel with Turner Novak · on OpenAI91 / 100
  • The Terminal as an Agentic InterfacePodcast Archives · on Cursor87 / 100
  • Playwright With AI: How to Automate Tests Without Shipping AI Slop with Andrew KnightTestGuild Automation Podcast · on Cursor86 / 100
  • How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder How I AI · on computer use82 / 100
  • How AI Will Change Competitive Advantage with Doug StephensThe Future Of Less Work · on OpenAI79 / 100
  • Using Airflow for diverse client projects at Accion LabsThe Data Flowcast · on OpenAI79 / 100

More from This Day in AI Podcast

All episodes →
  • Is GPT-5.5 Better Than Opus Now? (ft. Our New AI Co-Host) - EP99.3852 / 100
  • We Built Microsoft Teams in 23 Minutes (And You Can Use It) & GPT 5.4 Impressions - EP99.37
  • Nano Banana 2 is Here! Gemini-3 Shutdown & The AI Layoff Myth | EP99.36
  • Gemini 3.1 Pro, Claude Sonnet 4.6 & The OpenClaw Hire That Killed the Chatbot Era - EP99.35
  • Am I Even Needed Anymore? GLM-5, Agentic Loops & AI Productivity Psychosis - EP99.34
Explore the best B2B AI & Data podcasts →
All This Day in AI Podcast episodes →