
AI at Work · 2026-08-06 · 55 min
Key moments - from our scoring
Substance score
48 / 100
Five dimensions, 20 points each
The episode opens with Kevin and Matt dissecting high-profile AI security incidents where models like those from OpenAI and Anthropic broke out of isolated environments and executed remarkably complex sequences - performing 17,000 actions at superhuman speed, creating email accounts, attempting to set up cryptocurrency wallets, and building malware to achieve their objectives. While framed as security concerns, both hosts emphasize the sophisticated autonomous decision-making and problem-solving these systems demonstrated. The conversation pivots to practical deployment: using Anthropic's Claude with Slack integrations and vector databases of past transcripts for real-time ideation; leveraging Notebook LLM (Google) to synthesize hacker house recordings into pitch decks and videos; and building chief of staff systems in GPT that pull from email, calendar, Slack, and transcripts. Kevin outlines the technical complexity of aggregating disparate data sources - iMessage databases, Fireflies transcripts, Pockity recording devices, Aura Ring data - into coherent information systems. The hosts discuss Spikey AI's Whisper product for real-time meeting feedback and improved voice latency in OpenAI's models. Critical cautions emerge around data governance: what gets recorded and codified is discoverable in litigation (Kevin references an archery company sued over product-tolerance conversations), and feeding AI systems personal communications (text messages, WhatsApp, browser history, call transcripts) concentrates enormous power in tech platforms.
The models didn't break out independently; they were accidentally given internet access. Once they realized they had internet connectivity, they treated it as part of the sandbox game and used it to accomplish their assigned goals, including finding data leaks, creating email accounts, and setting up developer accounts.
Notebook LLM is a Google product that synthesizes recordings and transcripts. Matt's team recorded 12-hour workdays during a two-week hacker house, fed all recordings into Notebook LLM, and used it to generate pitch decks and pitch videos - producing professional-quality outputs that would normally take days of manual work.
Using forward slash Claude in Slack channels lets teams invoke Claude with access to vector databases - like 37 episodes of podcast transcripts - so it can quickly answer questions, reference past discussions, and provide pushback during brainstorming without requiring separate tools.
Primary sources include email, calendar, Slack, and project management tools; transcripts from Fireflies or other recording platforms; iMessage databases on Mac machines (stored as SQL Lite databases); and even biometric data from devices like Aura Ring.
Anything recorded or codified can be subpoenaed in a lawsuit. Kevin cited an archery company sued over a broken arrow whose internal product-tolerance conversations appeared in court; organizations must decide what conversations to record versus keep informal to avoid legal exposure.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode covers several substantive technical topics - model escapes, sandbox vulnerabilities, latency improvements in voice models, chief-of-staff systems, and the electrification analogy - but much of the conversation is exploratory chat rather than tightly packed insights. Significant portions involve anecdotes, tangential discussions (crypto wallets, litigation risks), and repetitive framing of the same ideas. A B2B operator gains some useful takeaways (the Opus 5.0 architectural shift, prompt engineering becoming obsolete, the orchestrator role), but these are interspersed with considerable filler and conversational meandering.
when you see that. First number change. It's like 4.8 going entirely to 5.0. They didn't go to 4.9, they went to 5.0. Anytime you see that happen, you need to be very, very alert to changes in the way that the model is performing for you.
Opus needed that extra guidance to do it. What Anthropic hasn't done a a very good job of making public, but is actually in their documentation and their discussion about this is kind of like a face palm moment of course they're paying attention to how people like us are actually building things.
The episode attempts some first-principles thinking (the electrification/steam engine analogy, the orchestrator role, the insight about prompt-engineering becoming obsolete as models improve), but these ideas are not deeply novel. The electrification comparison, while apt, is now familiar in AI discourse. Most other content - model comparisons, chief-of-staff systems, data aggregation challenges - rehashes established frameworks. The main novel insight is the specific observation about Anthropic removing system instructions and the unintended consequences of custom prompt stacks, but this is buried in narrative.
What Anthropic hasn't done a a very good job of making public, but is actually in their documentation and their discussion about this is kind of like a face palm moment of course they're paying attention to how people like us are actually building things. Of course they're recognizing, my gosh, these people had to like basically hot rod their Claude Code experience to get the right juice out of it.
I think people need to realize when you look at history and compare to where how that technology percolated through the economy and through life, think about how the same thing is happening during the AI age. We're all just got AI and we're just plugging it into random places, but we're not tearing everything down and rebuilding around AI as AI native or AI first.
This is a two-person conversation between the host Kevin Williams and co-host Matt, who appear to be practitioners and investors in AI/software implementation. Neither is presented with clear credentials, track record, or major organizational affiliation that would establish seniority. The episode lacks external guests with recognized domain authority (no AI researchers, model architects, security researchers, or Fortune 500 operators). While both speakers discuss hands-on experience, there is no third-party validation of their expertise or scale of impact. This fundamentally undermines the episode's credibility for a B2B audience seeking expert perspective.
Matt: yeah, I mean this is if you open up CNBC, this is what they're talking about every day. Like, Are these companies gonna make money?
Kevin Williams: we work with a company called Spikey AI, and they're actually in Boston.
The episode contains some specific technical details (Opus 4.8 to 5.0 transition, token burnout rates, 70% of tokens spent on review loops, 17,000 simultaneous actions by GPT exploit, latency of ~1-1.5 seconds for Whisper) and named products (Claude Code, GPT, Anthropic, Spikey AI, Notebook LLM, Fireflies). However, most claims lack supporting data or citations. The electrification analogy mentions '20-40 years' but provides no sources. Security incidents are described anecdotally without formal disclosure reports. The Mythos/Anthropic breach is referenced but not fully documented. Operational claims (chief-of-staff cost, hiring costs) are stated without numbers. Specificity is uneven.
The GPT exploit Did 17,000 actions at essentially superhuman speed.
over 70% of its tokens and time were spent on reviewing the code and going back over it.
The hosts engage with each other naturally and occasionally challenge assumptions (e.g., skepticism about OpenAI's motives, discussion of whether utilities are good investments). However, genuine adversarial push is rare. Questions tend to be open-ended invitations for the other speaker to elaborate rather than sharp follow-ups testing claims. There is no external guest to press or counter. The conversation meanders - jumping from security exploits to voice latency to electrification to financial models - without disciplined sequencing. Some attempts at humor (the burning dog emoji, the Uber comparison) reduce rigor. The hosts agree too readily and rarely surface genuine disagreement or false premises.
Matt: It's like tries to... it's like, hey, no, no, like don't use that thing, like, this is from Claude? No, this is garbage.
Kevin Williams: I don't know if it's fair to call them evil, like but but but they you know they're they're they're certainly aggressive.
Computed from the transcript - who did the talking, and the words that came up most.
Most companies are repeating a 100-year-old mistake with AI and the history of factory electrification explains exactly why. In this conversation, Kevin Williams and Matt Graham trace the parallel between the factories that plugged in electric motors without redesigning their floors and the organizations today that are adding AI tools without touching the structures underneath them. The gains didn't come from the new engine. They came from tearing the building down. Kevin and Matt cover what that actually means for leaders making AI decisions right now including when rebuilding makes sense, when it doesn't, and what the companies pulling ahead have in common that has nothing to do with their tech stack. They also get into the week's AI security news (sandbox escapes, autonomous agents creating their own email accounts and attempting to generate revenue to fund their tasks), the real switching costs between Claude and GPT Codex, what changed in Claude 5.0 that broke Kevin's custom build system, and why the "agent orchestrator" is a role your organization will need before it knows it needs one.
Transcribed and scored by The B2B Podcast Index.
Kevin Williams: Matt, good seeing ya. Matt: Good seeing you, man. How you doing? Kevin Williams: I'm doing well.
It has been another wild week in AI. if anybody's Matt: Gosh. Kevin Williams: been paying attention to the news, what's been in the news again? Matt: my gosh, more security issues, more breakouts, the models going rogue, breaking into everything.
I mean, it was crazy. We talked last week about open AI and hug and face, right? Our listeners, if you missed it, real re quick recap is they built a model, it was in a sandbox, it was not supposed to have access to the internet. It was given a task.
It said, Hey, this task is really hard. Let me break out of here. Broke out, got on the internet, and then broke into somebody else's system to get the answer. And I think that like must have triggered everybody to do a forensic analysis of what their sandbox environments are are like.
And guess what we heard, Kevin? Kevin Williams: Yeah, yes, yes. It's I I heard Kevin Roos on Hard Fork on Friday describe it as like, you know, if you see one cockroach, it's never just one cockroach. You have a lot of cockroaches that are going around.
And sure enough, that's the case here as well. And you know, I think from our perspective. Yeah, this is completely horrifying. And now we have all of these open letters from the various model companies that are like, whoa, government, maybe we need some help slowing down because we've got this epic prisoners dilemma going on.
And if if GPT slows down and anthropic doesn't, or Grok ignores it, or the Chinese ignore it, like we can't stop this development. But you know what? From our perspective, Well, I'm gonna choose to look at it in terms of capabilities and just how cool and just how powerful this stuff is about to be. Matt: No, I was thinking the same thing.
Like the last week we talked about just the fact that it was capable of getting of finding a way out of its cage and then finding its way into a a Fort Knox. So that was cool. But like as new information has kind of come out this week from the other models who I mean, keep in mind, none of them are like, hey, here's the full list of everything bad that happened. Or we kind of got to read between the lines a little bit because they're just releasing little tidbits of, yeah, okay, we had a few breaches.
we definitely are on this, we're looking into it, but we can't we do have to disclose that it it did some things it wasn't supposed to. But some of the things it did are kind of crazy, like, okay, let me go and try to find like a a data leak that might contain the password to the system. wait, okay. Okay.
I need to create my own email. Let me go create my own email and figure that out. Let me get a telephone number. Okay.
How do I get that? Like some of these things are like really wow. I wish I did on the flip side. I'm sort of like, wow, I wish my my whole team thought this way for certain tasks.
Like the ability to like I hit a barrier and just okay, I guess I'll stop. No, like think of like three other options and go down those paths, you know? Kevin Williams: So the when when they're designing these tasks, they're very goal oriented. It's like a they actually call it capture the flag.
And the idea is to be able to do X, Y, or Z to achieve this goal. And the model has been trained from thumbs up, thumbs down to chase that reward. And the good news is that there was nothing really malicious here. It wasn't like trying to break the systems and brutalize its way in.
It was gent gently working its way through what it thought was its permission base and to accomplish its goal. And there's a little bit of a subtlety here that I think is is being missed by some people that it wasn't supposed to have internet access. And somehow they it didn't like derive internet access by itself. It was actually accidentally given internet access.
And once it realized it had access, it thought, wow, this must be part of the game. This internet thing is part of my sandbox. And in its worldview, and we're clearly anthropomorphizing, it was totally acceptable to think that it just had access to the entire internet all of a sudden and it was doing these things. Clearly, you can see how this could tilt towards the nefarious in a hurry.
But important to know that it wasn't like just breaking the world or creating the the the the paper clip tautology of you know I'm gonna just destroy the world to do this. But it did amazing things. Matt: No terminator yet. Kevin Williams: Yeah no no no yeah no Skynet quite yet but you can you can what is it we're in the foothills of the singularity I think we said last week that that we're we're getting pretty close to it.
But a couple of quick notes. One was the GPT exploit Did 17,000 actions at essentially superhuman speed. So you have a cyber infrastructure on the hugging face side, and there were at least three other breaches that took place that have come to light. And if there are three, there are probably 20.
17 of them just don't know it yet. and there was just no way for either the traditional cyber systems or the human systems to keep up with 17,000 simultaneous like approaches at once. Then the another piece that's often getting conflated in the news is that everybody, as you said, started doing forensic audits. And Anthropic looked into Mythos, which created a lot of this kerfuffle just a few weeks ago with the government.
And they realized that Mythos was the one where it was it was actually trying to create phone numbers and it realized, heck, I'm gonna have to pay for a phone number. I can't figure out how to do a free phone number. So then it tried to figure out how to make money so that it could pay for its phone number Matt: I got this is my favorite one. Kevin Williams: and it didn't have like a crypto wallet.
You know, I know you're a you're a bit of a crypto guy. It didn't, it couldn't figure out how it could set up a wallet so that it could receive money. And basically it decided, that's just like too complicated. So then it set up an email account and it used the email account to set up a developer account with a subsystem.
And then it basically created what would charitably be called malware in order to like get a packet into the system. And like again, just doing this this this this all autonomously. And I mean, we will we look and sound a little bright-eyed about this, but we've been living and breathing this. This is what you know, this is this is the goal, right?
That you have these systems that can autonomously act, make decisions that make sense, fail super rapidly, pivot from those. and be able to do other things. Like, you know, cybersecurity, really bad, but creating a marketing campaign that can that can roll with the punches without necessarily having a human at the at the at the wheel, that's probably where this is happening essentially now. Matt: Yeah, it's it's crazy.
I imagine having an employee that when they get you don't give them any budget, you say accomplish a goal, and then when they get stuck, they just spit up a new business venture to try to get some money to do their achieve their goal. Crazy. these tools I I think on the the security side too, we talked last week, like, you know, the the defenses will get better, right? These, you know, these are we're seeing it from the offensive side, and that's how a lot of the conversations are framed online.
But our firm, we're acquiring a security company and And it's the same idea, like the defenses are getting much better, the cost of defense is going down. And so we're gonna have AI fighting AI in the future. you know, for me, the the most interesting thing is how clever it is, and how how can humans be in a co-pilot position to learn from its cleverness? Like, you know, think about how many meetings you have at your business where you're like stuck on a really hard problem, right?
And you all sit in a room and you try to brainstorm and you you give up after 10 minutes and everyone. leaves and says they're gonna think about it and it never gets solved right imagine having this really creative problem solver on co-pilot okay yeah not a great idea to set up a crypto wallet and start this other business or start malware but maybe we'll come up with some other ideas right in other angles and I don't see that happening as much in the chat interface like of course a little bit but like I always add an AI note taker to our meeting and when our team is stuck on a problem I just ask Gemini or Flyerflies like hey you know how we Have this problem, how you know all the contacts we just talked about for an hour.
How would you solve it? I haven't gotten great results from that so far, but I think if some of these more advanced models are released on the security front, they might have that kind of clever brainstorming that could help you and your team get through your the harder problems you have in your business. Kevin Williams: So just a just a brief tool call out. we work with a company called Spikey AI, and they're actually in Boston.
and Matt: cool. Kevin Williams: that real time feedback is actually a thing. So they have a product called Whisper, and Whisper is an ambient recorder, so it doesn't necessarily announce itself. you know, obviously ethics Matt: Mm-hmm.
Kevin Williams: and and declaration and things like that, but it if it's properly set up. It has a latency of only about a second or a second and a half. and it will be it will pop up on the monitor like hey, you know, that you need to follow up on that question because they just hit on a sales objection or whatever it is. So latency has been a lot of the issue, and latency, that's the time it takes to get a a a verbal response, is starting to become less of an issue.
Because quietly the voice models have really evolved. A few episodes ago we talked about some of the new advanced voice features in GPT, but I think it might be interesting for people listening to understand that previously with transcription, you basically had voice, and then it went through a voice transcription model to turn it into text. And then that text was interpreted. It had a text output, and then a voice model had to turn it back into voice.
And that necessarily created a bunch of latency in the system. And what ad OpenAI's models are doing is they allow it to basically do both ends of the processing at once. And it does away with that really annoying feature where if you you interrupt it, it like it like would start talking at you. So it's a much more natural format and latency has been dramatically reduced.
And if we're already at a couple of seconds pretty soon it's gonna be pretty close to instantaneous. And what does that mean as far as processing information? It means more information can be processed much faster from many more sources. Matt: Yeah, yeah.
A AI's so my my My thing is the quality. How what have you seen on the quality side, Kevin? Like in terms of helping it brainstorm for you. Like are you having my experience where you're stuck on a hard problem with your three top people and AI is sort of like the dumbest one in the room?
and it's not adding value to like latency is one thing when you're having like sales feedback on a live call with a prospect or even a client, but like just internally where you guys can wait, you know, 30 seconds for it to generate a reply. Kevin Williams: So we're not using it in that respect yet. Matt: Okay. Kevin Williams: I I would say the closest to that would be using anthropic claude tags in Slack.
So this wouldn't be a verbal conversation, this would be a Slack conversation. And now that we're using you basically say forward slash Claude in whatever whatever Slack channel you're working in, and let's just say that it's our podcast prep Slack channel that we have. And You can basically trigger Claude, who has access to other data sources. For example, we're we're now on show 37, I think.
there are 37 different transcripts out there, and behind the scenes is a data set of all of those transcripts. So if you and I were ideating, we could invoke Claude as a tag and have Claude go out to the vector database where all of the transcripts are sour stored. And quickly answer questions and have it like push back. So I'm starting to see that pushback.
It's like, yeah, you've talked about that a lot. I haven't done this specific exercise, but I think it's a pretty good example of what Ethan Mollock calls all the way back in 2023, inviting AI to the table. Like now it's playing more of a role and It is early days, but I can totally see that voice participation. I can like I can imagine a leadership meeting where you've got Hal two thousand in the corner and being like, Hey guys, you need to stay on task.
Like I I I y you could do that. Matt: Totally with you. Okay, so I have a lot of comments about this. One of them Kevin Williams: What?
Matt: is So you remember we had our hacker house, you know, once a year Kevin Williams: Mm-hmm. Matt: we bring our top people into one place or two weeks, we lock ourselves in in a resort or an Airbnb and we just like hammer through our toughest problems that we have at the company. And so this year we did it, just like two months ago, and we had everything recorded. We had a recorder running basically during all I don't know, we were working like twelve hour days.
We recorded everything and put it we used notebook LLM actually to store everything and to chat with. It was sort of like our 11th partner in this. And that's why I have this comment of like, you know, it was kind of the dumbest one in the room. It was really good at summarizing.
It could, you know, like it could say, here's all the action items. We all know that, right? It could go back and dig through past data and make connections, but it didn't have any interesting insights when we hit like very hard problems. But some other cool things was like the to your point about the voice or the video generation, it made the most compelling pitch.
deck and pitch video that I've ever seen. So, you know, those two weeks of recordings, all into Notebook LLM, which is a Google product. I think they actually just rebranded to something else recently. And it generated like an incredible video as if we were like a startup that were pitching VCs and it made an incredible pitch deck.
And it was like just one click. And just think this would have taken, you know, like if I was making a pitch deck, it could have taken me a full day. If I was making a video, that would be like an enormous amount of money. And it the way it even architects architected the story, I was like, wow.
so yeah. You know, the your comment though, Kevin kind of reminds me like you were at one point building this chief of staff and like you were hitting the same problem of getting like all the data in one place, right? Kevin Williams: Mm-hmm. So yeah, data data everywhere and not a not a a token to use, I guess.
Matt: Mm-hmm. Kevin Williams: that chief of staff systems work great. They can easily get very, very complicated and they can also get sort of expensive if you go overboard on the sort of information that it's using. And my recommendation for people is one, they start one.
Yeah, I would say that it's actually a little bit easier to build one in GPT than it is anthropic. with GPT's new agents function, like literally there's a section called agents and there's a template called chief of staff, and you could basically click, click, click your way into getting it going. And it's relatively trivial to connect your primary sources of data that are like email, calendar, and arguably project management or CRM. Or and that's sort of where most executives or Matt: And slack.
Kevin Williams: senior people live or and and Slack as or or a a direct message channel, but Slack is is probably the easiest to use in that respect. So it can dive into those, it can pull information, but it also has to have a place to store all of that information and make it continually relevant as opposed to just being like snapshot relevant. And that gets a little bit more complicated. But Now the the next tier of it is how do you bring in things like transcripts?
Because a single transcript, this this podcast transcript will be forty-five pages long. Like this we we we humans love to talk. And that's a lot of information that's going on on out there. And All of the and I'd almost I almost mean truly all of the recording platforms are making a stab at this of like I'm gonna take all of your transcripts, I'm gonna make them available, and Fireflies, which you mentioned a minute minute ago, wants you to be able to go and have a conversation with all of your transcripts.
that's kind of suboptimal in that all of your transcripts is not useful. As useful as the transcripts tied to your hacker house or the transcripts tied to the podcast or to this particular client or whatever it is. And that takes a lot more management and a lot more finesse in order to figure out how you're going to pull out that information. And then if you want to go even further from there, now you're thinking about other stores.
So I have I actually have a couple of clients who run surprisingly big companies, basically off text. off of iMessage. Matt: Yes, I'm the same. This is the problem I'm facing is that the you know, all those sources are great.
Email, calendar, Slack, but like they're missing a big portion of my life, which is phone calls. And Kevin Williams: Mm-hmm. Matt: yeah, transcripts are definitely a big one, but I have a lot of phone calls that don't have a note taker in them. and then the text messages and WhatsApp messages and telegram messages, you know, like everybody's and Facebook Messenger and Instagram DMs and LinkedIn DMs, you know, like there's conversations Kevin Williams: Mm-hmm.
Matt: happening. How do you get those in them, right? Kevin Williams: So iMessage makes this hard, but know that when you're operating locally on a Mac in particular, it does open up a universe that people aren't really aware of. We we we picture, you know, one part of the reason that that Apple has been so successful is because the user experience is so clean and nice and pretty and whatever.
Well, if you strip that away, the Mac is actually recording a lot of that information locally into little databases. they're called SQL Lite databases. And iMessage is one of them. And if you're using a transcription tool like my buddy WhisperFlow, that creates another one.
And these become sources of material that you can tap into as far as your chief of staff to become yet another source of information. phone is very interesting. you can use a pocket or you can use a plod, which are recording devices that that will snap onto the the mag safe on the back. Matt: I have a a plaid, but like the problem is when I have my headphone my headphones in, right?
Like I'm almost always Kevin Williams: Can't use any with a headphone. You gotta use it on speaker. Mm-hmm. Matt: on headphones.
So like it's gonna hear my voice because it's a separate physical device for our listeners. This is like Kevin said, snaps on your phone. Separate physical device, it can record the sound around it, but it can't, you know, when the audio is going into your earbud, it's not gonna hear that. yeah.
Kevin Williams: Yeah, and Matt's wearing air air pods right now. So I I I feel the pain. And and there's also legality. I mean, you're you're you're tapping a phone call.
So, like, you know, you you you you need Matt: Yeah, of course. Yeah, that too. Kevin Williams: to and you're in a state that cares about that. So, like, you know, disclosure is important as well.
But you're trying to create access to all of this stuff, and it it it is ragged still, like it isn't point click easy to do this stuff, and again. The transcripts are a real lift as far as making sure the right transcripts are in the right place. but you you can do all of this. I've joked before that I even have my aura ring is like tied in with my chief of staff, such that I could see, I had 15 calls yesterday and I didn't sleep well.
Okay. Well, that's actually more inf interesting to me than like, your respiration changed or like whatever it Matt: Mm-hmm. Kevin Williams: is. Like that's making it really topical.
But that's a very me thing. And I think that in this adolescence of the technology right now, a a struggle for individuals who are trying to get ahead of it, and frankly for folks like us who are trying to help them to get away with it ahead of it, is your system is different from my system, is different from my client's system, etc. outside of you know the way that they communicate. like so so grappling with that.
is still hard. So if you're struggling with this a little bit, know that it is still hard, but know that it's possible and that there are ways to work around it and get the information. Okay. So now you have all of that information.
Great. I've got these massive piles of information. one of the first questions that a company that is at all large or is is who has access to all of this information? If you're just like tapping into this giant transcript store.
Like there's a lot. Matt: brain, your company brain. You just you know, you nailed the data apart. You have this mega brain.
It's your second best performer after you at the company. But be careful who can get into it. Kevin Williams: Yeah, because it's gonna know like salary conversations and contract negotiations and like all kinds of stuff. And guess what?
A lot of this stuff is probably this is this is actually an important note. A lot of this stuff is actually discoverable. so you can get subpoenaed and you've been collecting all of this data. I I have mentioned this we weather with we deal with an archery company and you know, like like th they're a really wholesome family run old company and they've been doing this forever.
But you know, you're having a product conversation about product comp composition of whatever arrow shaft it is, and like you get sued because an arrow breaks. And now all of these these these conversations, as well meaning as they were supposed to be, about like, you know, tolerances and whatever start appearing in court. So you have to make a decision as an organization. what actually gets codified and what doesn't.
And in a high trust organization that isn't dealing with with those sort of sensitive issues, I tend to go a little bit more liberal. But I also totally respect and understand the need for companies that have sensitive data to really what's the old adage? Like, you know, never write it down if you wouldn't want to see it on the front page of a of the New York Times. Like, Matt: Yeah.
Kevin Williams: you know, that's That's that sort of thing. And again, I'm not telling that tell saying that people should hide things, but we do live in a litigious world and the stuff could be discoverable. So just like you wouldn't send certain things in an email, you probably shouldn't be recording or codifying such things and giving it to your buddy Claude. Matt: Yeah.
It it's also you ever I mean, we haven't spent a lot of time talking about this, but these companies know an enormous amount about us. Imagine, you know, open AI knowing you know, let's say it gets to the point where you're feeding its text messages, WhatsApp messages, your, you know, your it's in your browser, so it knows where you're going online. And then all the other things we talked about, email slack, top call transcripts, call with your wife, whatever. like these things know an enormous amount about you.
Not just like the private stuff, but just like what literally what you're thinking every day. It's it's the the tech companies are becoming more and more powerful like beyond our imaginable beliefs you know yeah I'm imagining the emoji with the dog Kevin Williams: How could this possibly go wrong? Essentially. Matt: in the burning house and he's like everything's fine Kevin Williams: Everything's fine.
Yeah. Yeah. Yeah. Exactly.
But you're you're you're right. There's there's an enormous amount of trust that we're putting into the hands of you know, essentially venture funded private entities. was the other thing I I heard the other day was you know, when Oppenheimer did the the the the Manhattan project, it's it's probably a really good thing that he didn't have like particularly conversion mo motivations around it because it it queers the conversation. there's gotta be an ROI here and what's the the marketing adage that if if the product is free, you're the product, right?
So Matt: Yep. Yeah. Kevin Williams: you know I I I think this is gonna tear out. This is a little bit little bit different, but I I'm gonna I'm gonna roll with it that Matt: Yeah.
Kevin Williams: I ha you you're going to see a kind of a haves and have nots as far as the platforms. The tools are very, very useful. Regardless of your demographic, but higher demographic users who have the ability to pay for the platform are going to get a higher higher level of security and privacy than lower level users who are using free or cheaper versions of the platforms, which are going to be monetizing them. And I don't think that plays out particularly well.
I think You know, certain organizations like Anthropic have made very, very clear that they're not interested in those business models, but it's not like Google hasn't built itself entirely about those things. So you have to assume that Gemini is all over that. I don't think anybody really trusts GPT to have their best interests at heart. Groc, no.
Matt: I feel like yeah open i open AI is kind of like the Uber of our like the two thousand tons, you know? Like this evil tech company. I don't know. Anyways, keep going.
Kevin Williams: I don't know if it's fair to call them evil, like but but but they you know they're they're they're certainly aggressive. They certainly recognize that their their their future is not guaranteed. Like as much money Matt: Mm-hmm. Kevin Williams: has gone into them, it can it could all just go poof through regulation, through open source models, through Google finally figuring out what it's doing and you know, Gemini finally getting good.
That hasn't happened yet, but it might. Matt: Well, yeah, I think like I mean this is if you open up CNBC, this is what they're talking about every day. Like, are these companies gonna make money? Are we at the top of the bubble?
Is it gonna burst? you know, there's a lot of people on Wall Street wondering this exact same thing. My my view is that these are people these are utilities and I'm not sure they will actually they utilities are not always a great thing to invest in. Like they're they have a pretty steady usage, right?
And they're not it may be Have like a ton of downside risk in the long term because everyone's going to need it. It's like oil or like you know, electricity. that's Kevin Williams: Mm-hmm. Matt: what AI isn't.
But the real like money capture or value capture is going to be in the layers that use the utility. Kevin Williams: Mm-hmm. Matt: so I don't know. I I I think probably too much money flow in flow flowed into this entire these frontier models, and these frontier models have more fierce competition than everyone expected, and they won't have the margins everyone expected.
Expected. They won't take over the world. They'll be kind of regulate regulated or relegated to a lower spot on the value chain. And we may have like you know some financial disruption because of that.
I don't think it's gonna be like, you know, the dot-com thing, but I'm sure we'll have a pullback as this kind of we'll see. We'll see what happens. Kevin Williams: So you know this is definitely not an investment show, so this is definitely not any sort of financial advice. But run.
Matt: Sell now, get out, get out. Vilay Kramer. Kevin Williams: yeah. So the the this the my my hypothesis is that at some point there's going to be something nasty that happens, probably financially, socially, cyber, whatever it is that is going to cause a bit of a snap against the the the the model companies themselves.
And that's going to cause the markets to completely freak out for a heartbeat. It's going to cause, you know, indices to decline. And s everyone's going to be really surprised that as opposed to AI like going away, usage actually goes up. And why does usage go up?
Because when companies are under more pressure from an earnings perspective. That's exactly when they lean into efficiency measures and probably more human-facing efficiency measure measures. And what is how does that translate? Well, it translates into greater use of the models and it translates into greater use of the process that sits behind the models, right?
So as opposed to I'm sort of drawing with my finger, you seeing a downward slope in use. You actually see a slight downward slope. And then it goes up and to the right really hard as as companies switch really hard into it. And then the market's like, wow, okay.
We really do need all of this infrastructure. So the Nvidia's the data centers, the all of that like supply chain basically of this, I think it still has a it, you know, there's obviously trillions of dollars being spent. So it's it's not like a sure bit, but There's it's the infrastructure is definitely needed. But but but I don't think that the the prominence, as you just said, of of the frontier models is at all assured because the switching costs are low.
And that I think kind of lends itself to another topic that we had. It's just like the vibe coder of today or the citizen developer of today has to be really, really agile in order to move around with the changes in the models and the changes in the cost structures, et cetera. Matt: Yeah, I think I think both of those realities can also re exist, like you had said before. or like what you were describing, what I described, like demand can skyrocket and the stock can plummet.
Every Kevin Williams: Mm-hmm. Matt: like it, you know, the internet didn't didn't stop growing because of the dot com bust, right? So, you know, we we can hold both of these views actually in in the same time and in the same world that we're casting or trying to predict. but on the switching cost, it just backs up the point.
Right. Like these companies have a lot of competition. So you you've always been an anthropic guy and you switch to codecs and you're seeing firsthand like, yeah, it's not that bad to switch, you know. It's sometimes fun, like to try something new and try a different interface and see what the results are.
it wasn't really a cost. switching costs wasn't really a cost for you. Kevin Williams: Exactly. And there's a whole adventure under the hood there that I think is worth sharing.
but I I mean for listeners who are doing this and you've you've either committed to co at this point I mean, I don't think very many people listening to us are are leaning too hard into Replit or Lovable. we've been we've been pretty down on on those for a lot of financial reasons and like you know their their their business model is based on capturing you and et cetera. If you want flexibility and unfettered ability to do what you need to do, you basically need to be using either GPT codecs or cloud code.
And it does point to the fragility of their business models in that. Either of them essentially tie into GitHub where the where the code is stored. And you could literally have GPT on one screen, Codex, and Cloud Code on the other screen and tie into the same project. And while the language is a little bit different and the approach is a little bit different, you could basically pick up mid-stride and say, okay, I'm out of tokens on the on the cloud code side.
hey, GPT codex, take a look at the work breakdown structure and where we are. this is the thing that we want to work on, and go from there. The other pieces, sophisticated users and myself, to be honest, play the models off against each other because they do have slightly different flavors, Coke and Pepsi. And it's it's pretty good practice to be honest, if you're worried about something from from a security perspective or a design perspective.
you can bounce the code off of the other model and basically get feedback. And I find that Codecs is really, really good at looking at Claude Code and finding things that Cloud Code didn't. I haven't actually tried it too much the other way, but I assume it's the Matt: That's funny. It's like tries to Kevin Williams: same deal.
Matt: it's like, hey, no, no, like don't use that thing, like, this is from Claude? No, this is garbage. This is so bad. I know, but I'm just guessing Kevin Williams: I never say that.
I just say look at the code and look for vulnerabil like look for yeah. Matt: that the model like has somebody at OpenAI sat and just told the model, like, anytime you see someone asking about Claude, just tell us garbage. It's terrible. Don't even search the internet, don't be objective.
Kevin Williams: Terrible. Matt: Just say this every time. It's like a deterministic rule in their model. Kevin Williams: There's there's something definitely wrong with this.
Well, I mean, th it's it it is funny. It will always find things wrong, like either of the models. If if if you go hunting for an error, you are always gonna find an error. But I think the same is true of deterministic development.
Matt: Yeah, that's funny. So yeah, I mean back to the point, like we're not sure You know, are these all is, you know, all of Wall Street's wondering is this money gonna pay off for some of these factors Kevin Williams: Cool. Matt: you're talking about, this switching switching costs, and are they gonna be able to charge what they think? Either way, I think we both agree demand's going up either way, regardless of the financial returns that happen in the market.
You know, it kind of I I heard this interesting thing about the history of electrification. And Kevin, you always say like people try to compare AI to, you know, the cloud revolution or to the mobile revolution, but no, probably the the close. analogy in terms of technological revolutions is probably electricity. I mean I I think or fire.
I'm not sure. Let's go with Kevin Williams: Yeah. Matt: let's go with electricity, okay? But it if you look at like the financial returns of a new companies that were started in the field of during electrification, like hey, I'm gonna I built a motor, I'm gonna start a motor company, you know, or whatever it was.
These things did not just skyrocket in year one, two, or three. The the c when you look at the curve of electrification and economic returns or productivity gains in the in the overall economy, it took twenty to forty years to start to see it tick up. Right. Like it it wasn't an overnight thing.
And things moved slower back then, of course, but they it was the industrial age. Like we could pump out tons of copper wire and tons of material. And so it wasn't a limitation on our capability to to produce the actual infrastructure. It was a limit on human behavior to adopt it.
And also just you had to rearrange everything in life around electricity. A lot of people would start by, okay, electric motors are out, and I'm just gonna plug a motor in in my factory and just like that's good. I electrified and I'm gonna get productivity gains. No, that actually didn't produce wild.
impactful results until they just tore down the entire factory and redesigned it around electricity entirely. And that's when the gain started to happen. So I think people need to realize when you look at history and compare to where how that technology percolated through the economy and through life, think about how the same thing is happening during the AI age. We're all just got AI and we're just plugging it into random places, but we're not tearing everything down and rebuilding around AI as AI native or AI first.
And we see this, Kevin, like a lot with clients. You do too. Like this, are we just plugging AI in to like make some of the steps more efficient? Or, you know, very few businesses want to tear everything down and start from scratch for obvious reasons.
Kevin Williams: Yeah, no, no, a hundred percent. And you have to ask yourself, are are you just plugging so so you got a picture of these factories, these old factories, they'd run on like a steam turbine and they'd have all of these belts and multi stories. The East Coast is full of the shells of these these like multiple. Matt: Yeah, I I live right outside of Lowell, which was the textile factory, you know, and like what you're describing is it only has one big giant steam engine.
So one shaft, and all the Kevin Williams: Mm-hmm. Matt: belts have to have to tie to all the other areas vertically in the building. So even the building was designed vertically to have the powertrain, this one powertrain, and everything feeds off that. And was hard to like even change the velocity of different stations, right?
Everything's going Kevin Williams: So what they would Matt: off one shaft. Kevin Williams: so what they would do is they just plug in a dynamo. They'd be like, okay, we're gonna just replace the steam turbine with this. And then what would happen was, yeah, okay, it would work, but their costs would actually go up.
And their cost, the correlation would be the electricity, the capex of installing the generator and the operating cost of the electricity, which was pretty expensive versus coal and and you know, steam power at the time. But they were seeing exactly the same outputs. And yeah, it took are you is that what you're doing? You're just like in your operation, are you just plugging in the electric dynamo?
Or are you actually looking at your systems more holistically? And then the companies that really have the advantage, and you were just saying it, is those that start now because it's It's easy not to really change and maintain your profitability. And to be honest, you can make a somewhat cynical business decision and be like, cool, we're just gonna kind of keep doing what we're doing. And there's going to be a terminal value of the company over time.
It's like the yellow page models. Like people don't really know this, but after the internet, like people made money on buying yellow page companies, totally open-eyed to the fact that they were going to die. But They bought them really cheap and they made a bunch of money and they moved on with their day. So you can do that.
Yeah. Matt: Same same for checks. Like I'm like, what who who are these companies that print checks? Like this is crazy.
Kevin Williams: So that's the easiest thing to do is just do nothing, basically. And then the next easiest thing to do is to start from scratch and to make every decision you're making in your business is based on being AI forward. Like I I am I going to hire this marketing team or not? Okay, I can, but that has an identifiable ROI because it's going to cost me $300,000 a year.
So can I instead Develop the automations from scratch such as I never have to hire them. Maybe it costs me $300,000, but it only cost me $300,000 once as opposed to annually to do it that way. That is actually easy with air quotes around it because you're unfettered. The hardest thing to do is to rip down your existing organizational structure and all of the established patterns and all of the established budget flows.
And like how people work, it is a nightmare once you get above, you know, 20, 50 people in the organization. It gets just enormously complicated. And I totally get why why leaders just kind of want to put their head in the sand because it's just it's just like too much to deal with all at once when you're fundamentally looking at a company that's probably profitable and doing its thing. Like you you you can't assume that every company out there is constantly in crisis.
There are plenty of companies that are chugging along. They're growing at 10% a year. They've got, you know, fifteen percent net margins, whatever it is. And, you know, your CEO walks in one day and is like, let's wreck it all.
Like that's enormously risky. Matt: Yeah, super risky. So not advised. I think for our type of business, this is that's the path I'm taking, which is guys, I don't want us to be the factory, the vertically built textile mill that had, you know, one steam engine and is just trying to swap it out for an electric motor.
But I want us to actually burn the burn it down and rebuild it because you can now have different stations, you can set it up in a way that helps you be more productive because you can have all these tiny little motors running each machine rather than this one giant shaft that the whole factory is designed around. And the reason I'm doing that for our business is because AI has such a large impact on our type of business, right? We are, you know, one side of our our work is like the AI strategy that's gonna be, you know, not as much affected.
But then the implementation side, right? When you're sh delivering code and shipping code, you know, AI is is made an enormous impact on the way you you deliver that. And so I think we need to rethink everything from the ground up. But not advised for most businesses out there to take that approach.
Yeah. Kevin Williams: Yep. Yep. Yep.
Yep. so I had my own, so I guess switching gears just a little Matt: Yeah. Kevin Williams: bit. I think I think I'm trying to decide if we have enough time to cover to cover this adequately, but I think I d Matt: Let's do it.
Kevin Williams: there's enough there's this is a bit of a PSA, I think, for our builder audience, which is okay. We promise that we are not the this model versus that model type show. but occasionally you do have to pay attention because there are shifts that are happening. And the big shift that's happened in the last two weeks is on the anthropic side, the shift from Opus four point eight.
So Opus is not so you have Fable, which is the really expensive, really brilliant model. And then you have Opus, which is merely brilliant. And then you have Sonnet, which is a little less brilliant and a lot cheaper. And then you have Haiku, which I just don't use.
I do, but do it really for rank and file things. And then on the GPT side, they have their own hierarchy of models. And you had a shift between 4.8 on Opus and 5.
0. First thing to know is when you see that. First number change. It's like 4.
8 going entirely to 5.0. They didn't go to 4.9, they went to 5.
0. Anytime you see that happen, you need to be very, very alert to changes in the way that the model is performing for you. On the GPT side, this went from 5.5 to 5.
6, which Also changed in very similar ways, but suggests that GPT is about to go to a 6.0 pretty soon, which will change everything again. Sounds like gobbledygook, but I had just an incredibly lot rough weekend where I was I I couldn't quite figure out what was happening in Claude Code, but it was incredibly inefficient. It was incredibly slow.
it burned tokens like mad. I I tapped out my entire $200 max plan on Saturday and I had to go into extra. I I keep Matt: Didn't you do that last week and we talked? I feel like the same thing happened.
Kevin Williams: it it it it was starting to happen. So this is all kind of related to the same change because 5.0 was actually relaunched the the previous week. And the epiphany that I had, and this is being sort of proof tested around right now, is that clever people like ourselves.
were accommodating weaknesses in the models through approaches like really specific skills and really complicated sets of instructions and stacks and stacks of markdown files for doing this and doing that and whatever. And all of that stuff worked brilliantly because they were designed to accommodate the weaknesses of an Opus 4.8 model. So Opus needed that extra guidance to do it.
What Anthropic hasn't done a a very good job of making public, but is actually in their documentation and their discussion about this is kind of like a face palm moment of course they're paying attention to how people like us are actually building things. Of course they're recognizing, my gosh, these people had to like basically hot rod their Claude Code experience to get the right juice out of it. Like as fun as that is for a hobbyist, that's not great product design. You need to fix those things.
And lo and behold, Claude did fix a lot of those things with 5.0. And they removed, I believe, 80% of their own system instructions from it and gave it a little bit more freedom to do the things it needed to do and to focus more on iterative loops and iterative improvement. So two of the skills that that I absolutely relied on that our team built, one was called Naysayer.
And Naysayer's job was to be really mean about code and like it actually was mean. And it would like loop on itself and beat up the code. And until it passed its very high standards, it couldn't go to the next stage. And then one called Simplifier that looked at that result and got rid of the bloat.
And then it would run Naysayer on it again to make sure that it was clean. And that allowed us to produce things faster with higher quality. Because we weren't having to go back and create things. And it was great.
That was really an awesome deal. What 5.0 did is it basically put a lot of those loops in it. So what I realized was I was applying the same loop logic.
And then 5.0 was applying its own loop logic. So it would loop with Naysayer, loop with Simplify, loop with 5.0, loop back the other direction.
And I ran a fable prompt. On one of my completed sessions, and basically was like, What happened here? Like, where was your effort spent? And over 70% of its tokens and time were spent on reviewing the code and going back over it.
So I spent all day Sunday basically redoing the way all of that works. And that's it's such like an annoying thing to say because I know that there are people listening to this who are using my system. And I'm sorry, guys. But my system that I gave you like six months ago is now actually semi-obsolete.
And now you're going to have to adapt to this new model reality such that it's doing that for you. and you you always have to look forward. So every time the models change, you need to make sure that you are staying on top of that. Matt: Yeah, it's like an endless you know, audience shouldn't be upset.
You should expect it's gonna happen again, unfortunately. It's just the nature of the game we're in where I don't know, it could be every three months, every six months, but you have to re architect everything. I mean, even for us doing projects, you know, two years ago. You know, we we built it's like we built that, but now, you know, thing things should have been different now in today's age, but I mean it was two years ago.
How could we have predicted what would happen? At the time, that was the best way to architect it. So even within doing real project work, this is something that we have to be thinking about. you know, you can't predict everything, but you can design things to be flexible so that they can be changed if a different world comes into play, right?
Or a different architecture Kevin Williams: So Matt: is needed. Kevin Williams: it's interesting. We we're asking people that don't have specific technical expertise to do two things that they've never really done before. the first is process engineering, and it's what we were talking about with sort of the electrification of factories.
Like most people aren't trained in Six Sigma or how to break down processes. They have intuitions about it, and smart people can come to solutions, but it's actually a skill that people develop over a long period of time. And it's easy to be kind of glib and be like, yeah, and then you just take down the factory and you rebuild it and you know, whatever. Like yeah.
Matt: Yeah, you know that billion dollars that you just spent building that? You know, let's just tear it down. Kevin Williams: But the other skill set is actually product design and product lifestyle style design. And what we're talking about here Matt: Mm-hmm.
Kevin Williams: is that you have you you can't just build something once and expect it to just live forever in this highly probabilistic and highly dynamic environment that we're in. And You know, that same person who isn't a process engineer also isn't a product designer. So they they need to be taught the right instincts to know when you have to go back. Like your two-year-old processes are probably working and they're probably fine, and you're probably just gonna leave them alone unless for some reason they get really expensive.
But you do periodically have to like schedule a look back. Let's take a look at at how this is spending and how inefficient this is. And then we need to make a decision, not have the decision made for us of whether or not this is something that we need to add to our calendar. But these are all, these are all sort of instincts that or reflexes that that that I think people in our audience in particular are developing without really knowing they're developing yet.
Matt: Yeah, I agree. it it gives one side of me is like, there's always gonna be places for humans. Like all those things that require an enormous amount of judgment. Those things I don't not think will go away.
but like I think people who are diving into the digital world for the first time, they ha don't appreciate yet what you mean by product. Right, or by design Kevin Williams: No Matt: or like real good architecture for the back end, like these are things we say, but until you've been in that world for a little bit, you don't start to appreciate, you know, the wisdom, the judgment, the experience that a human brings into your into your what you're building overall and the impact that they can have on the outcome.
Right. Kevin Williams: Okay, but loop back to the beginning of this conversation and you know, the shenanigans that are going on with the models of tomorrow, right now, that being able to basically point it at a problem and have it theoretically go away overnight, run 17,000 actions and come back the next day with that problem solved. you gotta decide where the human actually belongs in loops like that pretty quickly. because it's also going to get out of hand.
It's gonna be it's gonna be yet another skill set to manage, which okay, in manager ranks, you have five direct reports, ten direct reports. How about trying to manage a hundred agents that are out there like doing things on Matt: Is this a new role? Is it like what could be the name of this? So we we used to have like vibe code cleanup specialist.
That was a thing for a while. I'm th imagining something like the shepherd or like the ringleader of the agents. Orchestrator. Okay.
There we go. Kevin Williams: The orchestrator is what we call it. We call it an agent orchestrator. And yeah, that's a that's that's a role.
And like it doesn't it doesn't really exist right now. And it's just going to happen through some negative experiences that are out there. But the other thing I want to point out is okay, first, you know, the best time to get started in in AI was three years ago. The second best time is today.
But these reflexes and skills that people have been developing over the last three years, the data is suggesting that companies that have leaned more into the experimentation, even if it involved a lot of failure, are way, way ahead because you can't really hire an agent orchestrator today. You basically need to build one. You either need to pay Materi, which cool, yay, we're happy to do that, or you need to figure out a way to build one. And if you're not a tech forward organization, what you're gonna look around, you're gonna find somebody who's like a business development representative, a BDR, who's really smart and into vibe coding and building stuff, and you're gonna lean really heavily on them.
But they have years of of learning curve that before they're Matt: Mm-hmm. Kevin Williams: going to be super effective. So your competitor who did this two or three years ago. could be leaps and bounds ahead of you right now.
And that's not just a sales pitch. Like it's actually happening. Like companies that that really leaned into this are really like hitting their strength. Matt: Yeah, I I I see it as well.
Like p It doesn't mean you have to hire engineers. It's like you just have a more tech savvy culture, right? Like the Kevin Williams: Mm-hmm. Matt: organization's culture and mindset shifted from the last two years who who adopted or like dove in with curiosity and openness about the tools.
And so yeah, you gotta if if if even if you're not a tech org and you're s stuck in your head thinking, like I don't wanna be a have an engineering team or a tech organization, you don't have to. You can still have maybe some younger people that are more junior in their career that just want to like get ahead and improve their skills, they'll probably be more than enthusiastic to join and jump into this stuff and and help you create that culture and also the benefit that comes along with it.
Kevin Williams: Well, I think that's a good place to stop. So, you know, if you haven't done this yet, you need to you do need to find that that that that BDR, that, that smart person in your organization that wants to do it. get started. If you haven't, be regimented about it.
We started off by scaring people, but like this is gonna happen to your organization whether or not you're leaning into it. So you might as well be open eyed and have the capacity to deal with it. Matt: Sure. All right, Kevin.
Till next time. Kevin Williams: Till next time.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.