
AI at Work · 2026-06-24 · 47 min
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
Kevin Williams and Matt Graham discuss the hacker house model - an intensive, in-person sprint where cross-functional teams (developers, product managers, sales reps, and creative staff) lock in together to solve hard business problems using AI. Unlike traditional hackathons focused on learning new tools, hacker houses tackle production-ready challenges: Matt's team at Rapid Dev completed more work in two weeks in Istanbul than in two months of remote collaboration. The episode covers how to structure these sprints (composition, duration, expected output), when to use them, and how they differ from broader hackathon learning initiatives. Kevin then pivots to executive onboarding, sharing a case study of coaching a luxury fashion brand's CEO and creative founder from zero AI familiarity to deploying Claude-powered agents managing their schedule - covering practical setup (Whisperflow snippets, Slack integration, coworking automation), the importance of starting with risk education, and recognizing where Claude appears across five-plus platforms (web, browser, Claude.com, phone, desktop, Microsoft integration). The conversation emphasizes the messy reality of implementing AI: technical friction isn't code complexity but integration UX, organizational bottlenecks where side-project ROI gets trapped in CTO approval queues, and the value of identifying internal 'citizen developers' - high-potential non-engineers who can build disposable automation.
A hacker house is an intense, in-person sprint where a company's top 5-10 people rent a house and lock in to solve hard business problems - not just learn new tools. Matt Graham's team saw more progress in two weeks in Istanbul than two months of remote work, staying intensely focused (sometimes until 4 a.m.) with hands-on experimentation rather than delegated homework.
Eight people: two top developers, one technical product manager (who speaks both engineering and business), two sales reps (who understand customer needs), one creative team member (for outside-the-box thinking), and the founder/leader. The mix ensures diverse perspectives on hard problems.
He starts with risk and security (passwords, exploit awareness), then maps where Claude appears across five platforms (web, browser, Claude.com, phone, desktop), teaches advanced prompting (context, role, interview, task), and finishes by designing a custom agent for their specific workflows - such as Slack integration and scheduled coworking tasks - before leaving them hands-on ready.
A citizen developer is a high-potential non-engineer (like a salesperson with solution-finding instincts) who can build repeatable processes and simple automations but isn't a professional developer. They're often the bright spots that unlock rapid ROI opportunities outside the CTO's traditional project queue.
Executives are busy and skeptical that onboarding effort will pay off quickly enough. Kevin addressed this with the fashion brand CEO by showing her immediate wins (like using Claude's image recognition on her dahlias in the backyard) to build confidence and reps in the technology before scaling to business workflows.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode surfaces a few genuinely interesting technical observations - using Claude Code's /goal for recursive qualitative work, the minority-report multi-model consensus pattern, and open-router as a fallback harness - but these are buried under long stretches of general AI adoption narrative, mutual affirmation, and re-explaining basics. The insight-per-minute ratio is low for a 47-minute episode.
Anthropic did itself a disservice by calling clod code clod code. it is much more than code.
the same technologies that were used in Cloud Code can be used for qualitative tasks and development
Reframing Claude Code as a qualitative reasoning harness rather than a pure coding tool is a legitimately underexplored angle, and the model-consensus 'minority report' framing is a tidy metaphor. Most of the rest - hacker houses, AI adoption journeys, organisational friction - recycles familiar narratives without adding a contrarian or first-principles twist.
Anthropic did itself a disservice by calling clod code clod code. it is much more than code.
The art is finding what is the lowest common denominator model that you can use for the most cost effective price at whatever volume you're operating to get the highest level of output out out the other door.
Both participants are genuine practitioners - one running a dev agency using these tools daily, the other an operator-coach with real clients and a live alpha product - which gives the conversation authentic ground-level texture. Neither has done this at exceptional scale nor brings an unusually rare vantage point; this is solidly SMB-level practitioner experience.
I have a commercial product called ⁓ lead works. and lead works is a dynamic lead magnet tool that like it gathers very ⁓ precise information about leads
we got more done in we got more done in two weeks than we did two months
The transcript offers some concrete detail - named commands (/goal, /remote control), a real product (lead works), an 8-person team breakdown, a 1,000-data-point benchmark example, and a $150k vs $3M ROI framing - but most specifics are anecdotal and lightly evidenced rather than backed by hard metrics or verifiable outcomes.
we had like two of our top developers, one of our tech most technical project project managers or product managers ⁓ who can do the technical speak and the actual business speak. Then we had two sales reps
we have a set of a thousand data points. Go ahead and run those, that is a simulation... Nine hundred and ninety-seven of them correct
Matt asks a handful of decent follow-up questions (pitfalls of goal loops, parallel agents, balancing hype vs reality) and occasionally redirects the conversation productively, but the overall tone is friendly and mutually affirming with little genuine probing or challenge; several interesting threads are left under-explored.
what are some of like the pitfalls you saw, like trying to set the goal? Did it just get stuck for hours?
how do you also balance like the hype? versus reality
Computed from the transcript - who did the talking, and the words that came up most.
Most people heard "Claude Code" and assumed it was a developer tool. Kevin Williams and Matt Graham from Rapid Dev are here to correct that assumption and explain why the goal-oriented loop logic at the core of Claude Code might be the most underused productivity feature available to business leaders right now. This conversation started with a weekend of building and a realization: the same recursive, self-correcting loop that hardens software can harden a marketing plan, a strategic brief, or a financial analysis. The name was always the problem, not the tool. Kevin and Matt also get into the hacker house model how Rapid Dev gets more done in two focused weeks than two unfocused months and what happened when Kevin spent two days at a Los Angeles dining room table taking a high-end fashion brand from zero AI literacy to a functioning chief-of-staff agent. The episode covers the awkward middle most organizations are stuck in: past individual AI use, not yet at production-grade systems, and not quite sure how to close the gap.
Transcribed and scored by The B2B Podcast Index.
Kevin Williams: Hey, we're back at ⁓ AI at work, and I have my friend Matt Graham from ⁓ Rapid Dev with us again. How's it going, Matt? Matt: I'm doing great, man. I'm so excited for this today, Kevin.
I got tons of questions. I I I'm really pumped for our conversation today. Kevin Williams: Awesome. So I I happen to know, Matt, that you've ⁓ worked your way back to ⁓ to Europe for this time an internal hackathon, right, with your team.
Matt: Yeah, it's a little it's a little bit different. It's a hacker house. So hackathon is usually you're trying to get the team to like understand some new technology and we can talk all about that too. But this one is a hacker house where we got some hard technical problems that we're trying to solve.
We got our best people in the room locked in an Airbnb. It's a little uncomfortable, but you know, like that's the best. We're all tied to our desk type in a way. And we just got done doing one of these in Istanbul ⁓ two weeks ago.
We got so much done in we got more done in two weeks than we did two months. And so we decided to do one more. Kevin Williams: Okay. Matt: ⁓ this month together in another country in Europe.
So very excited. Kevin Williams: So some of our listeners may not be familiar with this idea. ⁓ I've done this sort of thing in the past, both with collaborators and internal staff. But the idea is that you know, you you're going to rent a house, you're going to put you yourself in close proximity with core people of your team.
So basically, there's no escape. And everybody is very, very focused on solving some problems. And of course, you're solving problems with artificial intelligence. So What does that look like in terms of an agenda?
And ⁓ you know, how can our listeners think about a hacker house or, you know, a hackathon, however you want to phrase it? And and sort of what's the run of show? How long do these things go? What's sort of your expected output?
Who's there, et cetera? Matt: Yeah. So a couple a lot of things here. Tell me when you want me to shut up, Kevin.
⁓ so for when I think hacker house, I'm thinking we have a hard business problem that we need to solve and we need to get a variety of different views in the room of and people that are really, you know, within the top ten percent of performance at our company. And so, you know, we get together, ⁓ you know, we we honestly barely even left ⁓ the Airbnb. We only got like a few groceries, we didn't even go out to eat, we didn't see any of the sites. ⁓ Kevin Williams: Yeah.
Matt: But you want a variety of perspectives because these are tough problems, right? When you're on a call on Zoom trying to solve a tough problem, everyone's like, Okay, yeah, I'll think about it. Then we all hang up and don't think about it. When you're in a house in an Airbnb, everyone's like, How about this?
How about that? We actually put our hands on the keyboard and we start trying it. Let's go look this up. Let's try this.
Okay. A few hours later, none of that worked. Let's try a different angle. Right.
And we're just all super focused and intense. Focused means you're blocking out other tasks, other noise noise. Intense means you're like staying up until literally 4 a.m.
You know, we I was just talking to one of my my buddies who was with us. He's like, Yeah, I remember I was up until 4 a.m. staying up just because he was so into the problem we were trying to solve, like in that flow state.
⁓ so yeah, I think it's really good to get your top people in person locked in a room when you have a hard, hard business problem to solve, especially when it comes to AI, because none of this stuff is like out of the box. Like, I mean, there's a lot of hype around it. Everyone's like, ⁓ it's so easy. You just type this and prompt it and bam, you got it all ready to go.
It's not like that. Once you get into the weeds and once you actually try to build something production ready, that's going produce business value. ⁓ and so you got to get the team together. You got to get them in the same room and figure it out all at all together.
Kevin Williams: So what's the composition of the team? Is it is it a mix of like creative, technical? Like what like how many people do you have crammed into this place? Matt: Exactly.
Yeah, so we only brought eight. So we had like two of our top developers, one of our tech most technical project project managers or product managers, ⁓ who can do the technical speak and the actual business speak. Then we had two sales reps because they understand the customer need, they understand what the customer actually wants. And then we had two ⁓ had someone from the creative team as well who's just kind of thinking outside of the box and like totally different angles and thinks totally different from the rest of us, and then myself as well.
Kevin Williams: All right. And and and how often do you do these? I mean, I know you did one like just two weeks ago, but Matt: ⁓ usually we do Yeah, yeah. Usually we do like one once a year.
Like what's our biggest lever you can pull in your business? That's really gonna 10X you. ⁓ this year, ⁓ we like AI's just moving so fast. There's so many big levers you can be pulling.
And again, we made so much project pro ⁓ progress quickly. I said, let's keep the momentum going. Let's do one more. I don't think we need another one after this for a while.
⁓ but let's just all get back. We changed the crew a little bit slightly so we get some different perspectives, not when some new problems we gotta solve. Get them in the room and we're doing it again. Kevin Williams: So ⁓ from my own perspective, having done this ⁓ both internally and externally, for those of you who are running ⁓ like SMBs, you have just a a few staff and you're not quite sure what this looks like, ⁓ embracing your broader community, ⁓ be you in direct consumer or you're in fashion or whatever it is, having kind of a peer group that you can do this with can be really challenging as well because you can bring different perspectives from different industries and different approaches.
but the format is still pretty much the same, except it's less executional. It's more like aspirational and idea oriented ⁓ when you're working with collaborators. The other is ⁓ you know, big companies tend to get a little bit precious about this. I know that sounds awful, ⁓ they want to do a retreat somewhere.
And then they like do all of these other activities and things. And ⁓ do like the idea of having like defined breaks ⁓ where like, We don't work during dinner. Like we're all gonna have like a really nice dinner and we're gonna leave the house because leaving the house for a few minutes can be really healthy and it inspires like cross-pollination of ideas. And then coming back on the on the other side ⁓ can be really helpful.
But you know, it it it seems pretty obvious and most companies don't do this. Matt: Yeah, I think I think another spin on this is you i even if you're not trying to get a bunch of people in the room to solve your biggest business problem, there's also a lot of value of just getting some group of people in the room so that they can learn these tools. Like a lot of businesses aren't at the cutting edge of everything that's happening and they just need to kind of level up their employees in terms of like, hey, I'm I'm the CEO.
I need everyone embracing AI. I myself as a CEO don't know exactly what that means, but we're all gonna figure that out together. So here we go. Let's design a quick hackathon together.
It's not really going to produce like, you know, a particular business outcome, but what it's going to do is get us all hands-on keyboards, experiencing these tools, understanding, getting the creative ⁓ genius flowing. And we'll start to see like how all of us can be using these tools individually in our own individual workflows. Cause a lot of these tools actually are really good for individuals. They're they maybe break down when you try to orchestrate a bunch of things together to produce a bigger business outcome.
And that's where it starts to get more complicated. But individuals in your companies need to be embracing these tools. And you as the leader has can have some influence getting them in the room to get them over that initial hump of fear or anxiety or, you know, just not knowing even where to start, you know? Kevin Williams: So that's actually a perfect segue to ⁓ my last week.
I was in Los Angeles and I was hired by a pretty high end fashion brand, ⁓ by the ⁓ who's the creative genius, and her husband, who's the CEO. ⁓ ⁓ they actually bought two days of my time. And ⁓ were so serious about going from ⁓ to probably two or three on the scale ⁓ they cleared their schedules entirely. And these are people who were in Milan.
the night before and they're bouncing all over the place. But they just realized that they absolutely had to make the space to do this. And ⁓ we sat down at their dining room table for two days and worked them through from I mean, honestly, literally zero. Like we need to get the chatbot installed.
I mean no. No. So ⁓ I mean that's extreme, right? Matt: Like had they ever opened they had never opened ChatGP or like opened it but not used it.
No wow. Well all you CEOs listening to the show, you're not you're not out you're not out of the game yet, right? Like there's still still plenty of people who have no idea. Kevin Williams: Yeah, exactly.
That's it. There are plenty of people like and and I'm not I I I'm not ⁓ recriminating here. Like they're a very successful business. They don't they don't need AI to drive their business forward, but they're recognizing the opportunity and they're recognizing the importance of of leaning in and ⁓ understanding that that they don't really have the time.
And you know, this is ⁓ this is something that we hear all the time is you'll approach a a busy executive and you'll say, Hey, we can, we can save you all of this time. And the response ends up being something along the lines of, I just don't have time to figure out how to save myself time. And it's just like, you know, you need to, you need to you need to slow down a little bit. You need to make space such that you as a leader understand the technology.
And, you know, it starts, it does start with the basics. It does start with. Matt: Mm-hmm. Kevin Williams: being able to use a chatbot, being able to prompt ⁓ really efficiently, and then understanding some of the like agentic harness features that are built right into the platforms.
So by the time we left, they went from knowing very, very little to having chiefs of staff that were managing their schedule and allowed ⁓ the the founder in particular, who really doesn't want to check emails, certainly doesn't want to get into a project management tool. Like she wants to be able to just basically talk into her phone and have the phone communicate the missives to the team. And now that that's all in place. So now the flywheel is starting to turn and we'll next we'll start working with the broader staff.
And in an industry fashion that is is kind of lagging behind technologically, this is a brand that in six months or 12 months will definitely be pulling ahead. Matt: So, Kevin, real quick. So, do you two days, these big executives run a big company, they fly you into LA, you sit at their dining room table, you talk to the, you know, that's a lot of pressure. They're paying you a lot of money.
And you're probably like, you know, I gotta make sure I put produce a lot of value here for these folks. ⁓ just like, you know, you want to do a good job. I know you do. ⁓ so like, how do you structure that?
Is it like presentations? Is it like you have a you plug into their giant TV and like show your off your screen? Like, and how do you also balance like the hype? Kevin Williams: Yeah.
Ha ha. Matt: versus reality, you know, like it's I think it's especially in the creatives world where like folks I've seen this all the time. Their their eyes just light up once they start seeing this stuff and they're like, everything's possible. You know, okay, everyone go on holiday for a full week.
We're gonna let agents run the entire company. You know, how do you kind of like dial it back to reality? And also just the format of your presentation, how you approach it. Kevin Williams: Mm-hmm.
So, you know, it's an interesting question. And I think it's a good, it's a good takeaway for people who are doing this internally because I know we have listeners who are basically in the seat of being like the AI Zar in their organization. And maybe they don't have the executive authority to make things happen. So they're trying to influence things.
⁓ I'm I'm I think I'm an interesting critter in that ⁓ although I run a training business and we do a lot of training, ⁓ I'm more of a coach. So if you're going to bring me in, you're bringing me in as a as an experienced CEO and an operator. And I'm going to be able to collaborate with you one-on-one as sort of a as sort of a peer to show you how I run my business. And I tend to overshare sometimes.
I have to be a little bit careful about what I'm bringing up on my screen because it's not canned. It's I'm literally showing like how I run my own day, how I start my own day with ⁓ with. A morning prompt that comes out of a snippet from Whisperflow. So I literally use Whisperflow the dictation tool in Claude and say paste morning briefing.
And that triggers a whole bunch of other stuff that happens in my day. And I think that makes it pretty accessible. But in general, I I hate to say this because I don't love to scare people, but these days I start with risk, ⁓ particularly with high net worth people. ⁓ and you know, b executives, like the risk is real out there as far as these new models and exploits and things like that, such that people understand how that how that looks.
So making sure that their passwords are all situated and such. And then I tend to go into ⁓ a phase of where does AI appear? So people don't really think about this too much, but ⁓ even just taking clause. Claude appears in at least five places.
So Claude's just the chat bot that you use ⁓ online, right? Claude also appears in the browser such that you can interact with it. You have Claude Design to do Claude to do more advanced either coding or qualitative topics. ⁓ you have the phone, right?
And you have Claude. Desktop, which allows you to do automations and co-work and all of these other pieces. And I don't count very well because ⁓ the sixth is increasingly across the Microsoft universe. ⁓ I've never been a huge Microsoft fan.
I don't know a lot of startup people who are huge fans of like getting into bed with Microsoft, but Microsoft plus Claude are leaps and bounds ahead of Copilot. And the ability to integrate Cloud in particular across Microsoft apps is a huge eye-opener. Like if you're a CEO, you're this client, that's that's a huge deal for you because a lot of your job is information synthesis, right? And being able to integrate the LLMs across these systems.
From there, we usually go into advanced prompting. We teach crit, context, role, interview, and task. Go into depth on that at some other time. And after then it's about solving some problems.
So I sat down with them, I I did the education, and then I understood what their actual workflows look like. And then I scuttled back to my hotel and I spent the rest of the evening basically designing what a perfect agent would look like for them and then taught them how to use it. So it's just agent version one for them, but walking them through how do you connect to Slack? Matt: Mm-hmm.
Kevin Williams: How do you put an ⁓ a a scheduled task together in cowork that can do this? Where does the prompt go? How do you use Whisperflow and create a snippet such that you can talk about these things? ⁓ by the time I I ⁓ left, everybody was exhausted.
They looked like like limp noodles by the time I I left. ⁓ you know, they have all the tools they need to start actually rolling. Matt: Yeah, so it sounds like what I'm hearing is like, ⁓ would you say it's accurate to say there's a lot of individual use cases that you're trying to inspire them with and get them up to speed on versus, hey, we're coming in, we're doing a business, you know, let's look at the business, understand the what the workflows are, what are the major levers in your industry and this business that you run, and then let's see where AI can be useful.
It's it's a little bit, maybe some of that, but a lot of like, hey, this is how we're gonna set you up individually, get your hands on the keyboard, start using this. Then invite me back or eventually we'll reach phase two of your journey, which is we'll start thinking about how to either get your employees to use these tools or start thinking about how the business more broadly can be leveraging AI. Kevin Williams: I I I think that's right on. because there is no just set approach.
Like, okay, if you have one Gmail account and you're connected to, you know, straightforward like CRMs or something like that, it's it's it's pretty easy to advise people. But it turns out the real world is really complicated. ⁓ you know, if you're dealing with Microsoft in particular, there are hoops you have to jump through to be able to make the agentic harnesses work with them. ⁓ I think to myself, you know, I ⁓ I live and breathe this stuff.
⁓ ⁓ still a bit of a fiddly challenge for me to connect all of these different pieces. ⁓ unless you're really an enthusiast, it's understandable that people are really frustrated and they're not really embracing the sorts of things that. Matt: Mm-hmm. Kevin Williams: That guys like us are are are leaning into because it's it's technically still a little bit challenging.
Not technical as in, ⁓ I have to get into Claud Code and I to build this or whatever, but it's just a different approach and people aren't quite used to it yet. And then when you translate that into like corporate trainings, when we do bigger corporate trainings, ⁓ you definitely have to move pretty slowly and within the kind of guardrails of the tech that the organization offers. and give people time to like experiment and give them little adventures. ⁓ I love giving people adventures.
⁓ again this founder. Yeah, side quests, like side quests, yes. in her case, she have she's a gardener on the side. And ⁓ you know, showing her how to use Claude and image recognition in Claude such that she could figure out like what was going on with her dahlias in the yard.
Matt: Side quest. Kevin Williams: Like she was so excited about that. And that excitement, it kind of goes back to your ⁓ to the ⁓ you know, the hacker house idea that that excitement, she's going to use image recognition for doing gardening, which is going to give her the reps of understanding like sort of the limits of the technology. And then probably she's going to start using it for fabrics and like other things that that that are more tied to her her world.
But the entry point was running around in her backyard like ⁓ photographing flowers. Matt: That's so funny. So it's interesting. Like everyone's got to get used to the UX, right?
The UI, like there's and like you said, it's not technically complicated, but there's a million different things you hook together and you log in here, you log in there, and then you see inner different interfaces for all these and you need to figure out where to click. So that can be a little nerve-wracking for some people. And then ⁓ what the journey I've seen is, you know, the CEO or some sort of leadership part of leadership team is interested, you sit down. You help them get up to speed, get them using these tools.
They start to understand the use cases. They are start to understand the ⁓ interfaces. They become more ⁓ proficient with it themselves. Then they want to get their employees to use it.
So they maybe run a hackathon or they do some sort of workshops or training or seminars with folks like ourselves or yourselves. And then ⁓ after that, then maybe they're really becoming more AI enthusiasts and starting to build, like maybe they have some citizen developers, but they want to have you know, something more robust or more professionally built. So it's feels like there's a bit of a a journey there that a lot of companies are going on, starting from the top, then the the employees learning and then actually leveraging these tools in more complex systems and more coordinated fashions.
Kevin Williams: I love that term citizen developer. ⁓ I'm gonna steal that for sure. That's ⁓ that's a good idea because we're we we're we're that's a lot of our end goal is to identify those bright point of light. And they often aren't within the CTO stack.
⁓ they're they're a salesperson who happens to have like a knack for solution finding and they probably wouldn't make a great developer, but they understand what the product needs to look like. I'll say that it's hard. Matt: Mm-hmm. Kevin Williams: It's really, really hard to graduate from, okay, we're all using this stuff.
And sort of at a basic level, we're building GPTs, we have some repeatable processes. ⁓ maybe we're bridging into agentic harnesses. This is all very, very new, obviously. Over the last three or four months, a lot of this has become more plausible for organizations, but they're still a little bit stuck with how they build the thing.
⁓ particularly since, as we talked about in the last episode, a lot of this stuff is still pretty fragile. So having people internally building sort of like disposable products that you can use for like internal whatever the heck it is, that's great. But when you start trying to build like service features and things like that, ⁓ people get out of their depth relatively quickly. And ⁓ you know, my goal is to teach people how to fish.
Matt: Mm-hmm. Kevin Williams: But at the end of the day, we run our, you know, our embedded AI offices in organizations, specifically because the organization is doing its thing and it can't necessarily slow down to solve the problems that we can solve out of the box. ⁓ obviously that's a great opportunity for people like us, but I'm being authentic in that I I really do want to see organizations grow that capacity over time. Matt: Yeah.
Kevin Williams: But at the moment, it's very early days. And ⁓ I would say too often, I'm seeing hesitant, but too often I see see too much being layered into a CTO vertical. And then like all of the does the the inputs and the needs go into the top of the CTO vertical, into traditional planning, budgeting processes, which you know that's that's what organizations do. ⁓ But there's so many opportunities that sit at the side that would have rapid ROI, but the individual ROI of you know X project, like, ⁓ it's gonna save us $150,000 a year.
And the CTO's office says, that's great. But we're working on these things that are gonna save us $3 million a year. Well, you know, that person, that salesperson who came up with the $150,000 idea, they get it. It's gonna make their life better.
Matt: Mm-hmm. Kevin Williams: it's gonna hurt improve their P and L, but they can't actually do it and they can't actually do it themselves. So they're they're sort of stuck. And that's where these forward deployed, you know, teams like your teams or my teams come in.
Matt: Yeah, it's interesting because it's not just a tech problem and it's not just like the tools are too complicated for the individual. There's a lot of actually organizational momentum or friction, or you know, like that salesperson, he needs to access data from different parts of the organization. He's not a project manager. He's never quarterbacked a large project that goes outside of his domain of expertise.
The CTO doesn't really understand the problem that the the salesperson has. So if you leave it to the CTO, he's not actually talking to the customer or the client or figuring out the actual problem that's trying to be solved internally. So there is not just a tech problem. And that's what I would caution like a lot of CEOs to think about when they think about building something bigger or trying to empower their team, but they should recognize that, you know, larger companies have digital transformation offices that have like very long SOPs and processes for change management, culture.
Cultural changes that are going to happen, personnel, habits that need to change, and then just all the complex coordination that needs to happen across departments, across different teams to actually make it work. So as the technology get gets easier, that is not going to go away, right? And for smaller organizations that haven't dealt with those more complex implementations, they probably don't have an appreciation for how comp complicated the human part is for all this, right?
Kevin Williams: Yeah, absolutely. Like, ⁓ you know, it's I I like to say that ⁓ you know, as an organization, if if if we have free reign, we will go in and we will break an organization chart. And ⁓ that's we're invited to do so, but guess what? The org chart has opinions about that.
So the org chart is made up of people and budgets and egos and all of these other things, and it's super challenging. So A lot of this is actually leadership positioning around, you know, how do we think about this? How is the how is the human element applied in the mix? How do we empower people?
When, where, and how? ⁓ what is our tolerance for risk? ⁓ you know, the closer to a regulatory ⁓ industry you sit, the tighter you're going to be on a lot of this stuff. But it's super challenging, for sure.
Whereas if you're a twenty person company and ⁓ you know, you can throw down with ⁓ five of your top people and ⁓ like just figure it out and bully your way through, it gives you a lot more flexibility. Matt: Totally. So Kevin, I'm gonna switch gears here for a second. You said something before the show started here about loops and ⁓ what you were working on this weekend.
I'm really excited to kind of shift gears into that if you'll let me. Because I was just l listening to a podcast. ⁓ it wasn't a podcast, it's like a YouTube episode. It was the I forgot the guy's name, but ⁓ the person who actually invented Claude Code at Enthropic, and he was talking about how code is solved with AI.
And I thought that was whole super interesting. And he started talking about the loops feature and how you set the outcome or the goal of what you want to achieve and less about like the individual task layer. So I think a lot of our listeners probably are focused on like how can I get AI to do this one task or like a small sequence of tasks, which we'll call workflow. But I think the big unlock that's coming is set the goal or the output that you want to achieve and then have it like recursively loop and Iterate and try to achieve that until it does.
So, what you learned, tell me everything. I'm really excited about this. Kevin Williams: ⁓ So we talked about goals as a primitive ⁓ about a month ago, aroundabout. And you know, if you have just been painting by numbers, if you've been listening to this show and you ⁓ you know, you've downloaded Claude code and you've been experimenting, and okay, you've got your super base and you've got Versel and you're launching things, you're you're probably getting a little bit frustrated in that Claude code is pretty needy.
It has a lot of questions. It's conservative for very good reasons. If you've built it in the right way, it's going to constantly catch errors, et cetera. But it takes like care and feeding ⁓ as far as giving it, either giving it access or ⁓ giving it permissions to do things that in a lot of cases you're not going to really understand because you're not super technical, right?
⁓ but at the end of the day, you know, any workflow is based on a trigger. Something is happening, right? And you have an output at the end of the day. And you can spend a lot of time fussing over the intermediate steps between them, or you can orient the system more on what the output is supposed to look like.
And this is actually a different skill set to understand what that output is supposed to be. And people who are good at scope are really good at understanding outputs, right? So by using a command that is forward slash goal in Cloud Code. ⁓ you can dictate what it's supposed to look like at the end.
And it's not like you're writing the forward slash goal. ⁓ I typically would use Claude to help me write a prompt that's based on the goal that I need for getting it done. And it will loop through continuously, ⁓ like trying to iron out its own issues, and it can run for hours at a time. This is where you you hear the ⁓ like.
the the ⁓ the elite of vibe coding are always talking about how they have seven agents going at once. ⁓ this is kind of a lightweight version of that, but for most sort of amateur vibe c vibe coders, it's it's really eye-opening to see how those loops work. ⁓ ⁓ if you've connected it to your phone, so I don't think we've ever mentioned this here. ⁓ If were to I don't I don't want be like too totally cloud centric because a lot of this stuff can be done in GPT as well, but I happen to be really I'm a Claude user.
So in Claude, in your app, there is ⁓ there's a code section in there. And in Claude Desktop, if you forward slash remote control, it's going to then send the process to your phone. So you can leave, you could go to the gym. I happen to know that you're a you're a big walking desk guy.
We've had many calls on ⁓ on ⁓ on walking desks, but you can pick up your phone and you can see the process that's going on to see if it's gone awry. And with a goal, those might run for a really long time. Matt: Yeah, it's super interesting. So what what are some of like the pitfalls you saw, like trying to set the goal?
Did it just get stuck for hours? 'Cause that's, you know, one fear that I have is it's just burning through tokens. It's a nice trick for Claude to just eat up your wallet real quick. You walk away and you didn't get it didn't actually accomplish the goal ten hours later.
Kevin Williams: Well, to be clear, ⁓ internally, ⁓ most of the development that I'm doing fits well under a $200 max plan. ⁓ I'm patient enough that I don't necessarily need to bust into that extra usage. So yes, it's true that you can burn through an entire session, but it can make a lot of forward progress in that respect. And then, you know, it will have stalled out because it either ran out of tokens or it just, you know, it it it it you have a session period of use.
And you can just say continue and it'll keep keep going. ⁓ I would say that it does take a little bit of discipline. I you know, it we have to run the gambit in this show between opening people's eyes to the art of the possible and and like kind of being overwhelming and technical. And I'm trying not to be technical, but you also have skills, ⁓ we've talked about at various times on the show, ⁓ that that do functions.
So I I have a skill that one of my team members built called Naysayer. ⁓ what Naysayer is doing is it's looking for code issues. And unless it scores ⁓ out of 10 ⁓ Naysayer. So it's going through a process where it's looking for code iterations, ⁓ can't ⁓ to the next step.
⁓ when you run a forward slash goal, it's trying to go do steps A, B, and C, but has its own loops. So basically A won't finish ⁓ the naysayer ⁓ it passes its quality controls and then it passes to B and then it passes to C. Previously, I would have been able to have done A. I would have been able to do it with naysayer and feel pretty good about it.
⁓ now I feel pretty confident of doing of passing A to B to C, but only because I have those controls under the hood. If you run a goal that's like, hey, build me this cool app that does all of these things, I'm going to lunch, like you might get something at the other end, but it might also have a ton of issues. ⁓ so I there's sort of a caveat there. But beyond that, I think my issues come down to ⁓ sometimes it can be over cautious and sometimes it can be over cautious about things that probably don't matter in the scope.
But yeah, no, it's it's it's been ⁓ super eye-opening as ⁓ as just ⁓ kind of a liberator. Again, I I do feel like that guy sometimes who has his thumb in the laptop ⁓ before like I'm walking around the house so that Claude doesn't turn off. ⁓ but now I have it on my phone so I can legitimately like go off and do something else. Often my prompt will be I need to go to the gym, like by the time I'm back, forward slash goal, have all this stuff done.
Matt: So are you running multiple ones in in parallel? Are you you know, so are some of our developers will have this running, this running, this running, and they're bouncing between, you know, different goals or different loops that are running at once as they build out different modules, perhaps? ⁓ are you have you gotten doing that or you just kinda just focus over the weekend and yeah, what was it this weekend you were working on? I can't quite remember.
Kevin Williams: so okay, so this weekend, so I have a I have a commercial product called ⁓ lead works. and lead works is a dynamic lead magnet tool that like it gathers very ⁓ precise information about leads that come in ⁓ then sends out custom nurture emails at the other end. And it it's really cool. It's like as a former marketer, it's like the lead magnet that I always wanted myself, and now I have the opportunity to build it.
But it is a touch of a side quest. So I have to like keep it on my weekends. and you know, it's an alpha and it's it there are there are a lot of leads that are cranking through it. And, you know, there it's an alpha.
So so customers are coming and they're saying, hey, you know, it'd be really great if we had this feature or that feature, or there's this bug. But now that we're at the 95% layer, it's really, really scary because it's out there in the live. Matt: Mm-hmm. Kevin Williams: Customers are depending on it.
Like now I'm being tested as far as the infrastructure that I've built, as far as whether or not ⁓ it's gonna carry me through or not. So when I'm doing things like goals, my goals get pretty elaborate such that they have to reference back to things that it can't touch. So, you know, I I never want to mess with live customer products that are out there. So I have to I have to work on my goals.
Such that we get to the desired outcome. But the process, the timing definitely slows down. So things that you know I would have kind of flown through in the first 80% of development. Now I'm going at a snail's pace because I just can't afford to break anything.
and I think that's a that's a very common feeling among people who are who are building these tools. And it's also a differentiator that if you're going to be disciplined about it and you're going to build something for scale. That's entirely different from building a little app that you're gonna use like once and toss away. Matt: Yeah, very different from building an internal tool versus external tool as well.
Like if you're building anything that's gonna touch customers, vendors, suppliers, you know, you gotta be much more careful with that than if you're just building something for a marketing department, internally, a sales department, internally. ⁓ you know, they'll be more forgiving than your customers for sure. Kevin Williams: Yeah. So the other part of goals that we were talking about earlier today that I think is really interesting and I don't think has been particularly well covered or particularly understood is the fact that this is not restricted to code.
⁓ Anthropic did itself a disservice by calling clod code clod code. it is much more than code. ⁓ it is first applied to code and product development and you know, building things under the hood. But the same technologies that were used in Cloud Code can be used for qualitative tasks and development.
And what I am starting to experiment with, and I would encourage audience to experiment with this as well, are using these same recursive loops, the goal loops within Cloud Code to work on qualitative outputs. So maybe it's a marketing plan or a strategic plan or whatever it is. The The skilled LLM operator today is likely to go to Claude or to go to GPT. They're likely to establish context, their lab, the background, everything they need to do the task.
They're going to pull in documents. They're going to do this thing. And they're going to say, I want the best business plan ever. And they're going to hit submit, and their chatbot of choice is going to go off and it's going to chew on it.
And it's going to probably give a pretty good document, right? At the end of the day. Well, Now think about this goal function in that respect, that there's a bunch of iteration that could go into that. So imagine ⁓ having like ⁓ steel man perspectives that look at it and harden the arguments and loop back over itself such that it's not a 20-minute process, it's more like an hour or a two-hour process.
As it's looking at all the permutations of the product and making sure that it actually matches what you're doing on the outside. ⁓ But that's where we're headed as far as goal-oriented prompting and outputs. And step one or station one was coding of like, ⁓ we can we can build parts of an app. And then later in that station was, ⁓ wow, we can build most of an app.
And the app has these features. And now you're starting to look at knowledge work or financial analysis or other things that aren't code, but it's using the same sort of logic. Matt: Yeah, that that recursive nature, the loop nature. That's exactly what that ⁓ the founder of Claude Coder, the inventor, I guess you could say, was talking about.
You know, something that was interesting working with my team just a few hours ago was not only looping it on itself, but looping it to other models to check it, like that naysayer concept you were saying. ⁓ we so in a particular thing we're working on today, we had five different models. that we were balancing it between to double check like who of our five models. First we'll start with two.
If the two models agree, then proceed. If there's a disagreement between the two models, then go to the next three models and see if those three agree or disagree. And depending on, you know, your confidence that you want to have, you can set it to go back or to move forward depending on maybe one of the two disagrees, but the other two do. So, you know, that's it's again, we've talked about token costs and prices for all this stuff.
So you can't, you got to be careful. ⁓ but there are a lot of cheaper open source models out there that make this possible if you want to have like extra, extra confidence with particular types of loops that you may be running ⁓ in your process. Again, that's probably too advanced for this crowd, but just something that folks can kind of keep in mind. But either way, all this stuff is getting better and better.
Cost is going down. It may not feel like that with tokens, but it will come down. And the cost of will come down not because of the token cost, but because of the it'll get it right with less loops, right? Kevin Williams: Yes.
Yes. We call it a minority report, like the Tom Cruise movie, like like being able to to to get a consensus among models, but it can be really expensive. So the art of this is not just brute forcing the output. The art is finding what is the lowest common denominator model that you can use for the most cost effective price at whatever volume you're operating to get the highest level of output out out the other door.
⁓ and that's that was a gut check, what we just said for like, you know, the 500 people who are listening who are in this are like, ⁓ I'm just using Claude. And it is different. You can't just toss in ⁓ GPT ⁓ instead of Claude. They have slightly different prompt logic as far as making it work.
So that that idea of looking at results is great, but You know, if the last couple of weeks in the in the news cycle with anthropic in trouble with the government because of fable and and things like that, this is a real problem in that, you know, between anthropic's instability, like it goes down all the time. It's completely nuts how often it goes down. and the fact that the government can just turn it off. Like, if you've just built your business on top of anthropic, you have a risk, whether you appreciate it or not.
I mean, you have a risk if you've built it on GPT, ⁓ which is what lends itself to this conversation about open source. And I really want to put a giant asterisk next to that because although you and I could easily like BS our way through ⁓ know, going to Hugging Face and downloading a model and like setting it up and provisioning it and whatever. and thumbs up, great. Now we have an open source model and it's only charging us process.
Well, you've just added all this extra infrastructure. You either have like real processors that are running, which can cost tens of thousands of dollars, like locally, or you're running cloud infrastructure of some kind, which believe me, Jeff Bezos needs a new settee for his yacht. So he's going to charge you what he's going to charge you to do that, right? And these things change all the time.
They have to they need care and they need feeding. And look around your office. Do you have a bunch of network engineers around? I bet you don't.
So Matt: Mm-hmm. Kevin Williams: ⁓ we're kind of stuck between a rock and a hard place that the right answer for an organization at a certain level of maturity is to start learning leaning into open source models just as a fallback in case that something goes awry elsewhere. But it's it's it's not simple or easy to do that. On the other side, the easier approaches are through ⁓ like routers, like open router.
⁓ so ⁓ you can use you can build open router right into whatever you're building, and it can do model selection for you. So you can have a primary model and you can have a fallback model, ⁓ such that if there is instability in the app, you can do that. But that doesn't necessarily help you if I mean I've been saying this for years. At at some point, something bad is gonna happen here.
Like it hasn't really yet, but like we are at the edge of kind of scary with some of these models. Matt: Mm-hmm. Kevin Williams: And a bad actor will do something that involves infrastructure, security, a dam, like whatever. And all of a sudden there's going to be like a snapback.
And you could see all of your models just basically either temporarily or permanently just getting turned off. And ask yourself, have I built a business that can survive that? Like, I'm I'm loving it. I'm so efficient.
I have one person who's running all of these operations and there are a thousand things that are happening. What if somebody just like Took your toy away. Matt: Yeah, it's a little scary. I still think it's less than five percent chance.
Like I think it's still worth diving in and going forward, but it's something to be aware of, definitely. One thing, another thing that was interesting today is as we run the run our system, we we create a database of all prior runs that ran successfully with the correct outcome. And we keep that database so that if we do need to switch to model, or we want to check if like, you know, the flash or the The flash model or the deep research model is better or worse, or whatever the efficacy of different versions of the same model are.
⁓ we have a database of set outcomes that we know it should produce. So, okay, we have a set of a thousand data points. Go ahead and run those, that is a simulation. Did the efficacy get better because we ran a more ⁓ higher-end model or ⁓ a more expensive model, or we just used a more basic model, or we used a totally different model.
And we know, okay, great, it got exactly. Nine hundred and ninety-seven of them correct, right? So this is pretty high efficacy, something like that. Kevin Williams: So reaction is awesome.
That's great. That's what we do with ⁓ industries that are in ⁓ adjacent to regulations, because you have to have traceability, you have to have explainability. So what you just built from a quality perspective also can be applied in terms of a liability ⁓ or a regulatory compliance ⁓ perspective. So that's great.
But now roll back earlier in the conversation where we're talking about that gap between the guy in sales who's trying to solve some little thing. And the infrastructure that's required to do model switching, prompt tracking, improvements, like no chance that this guy is going to have that skill set. And frankly, if you're in a semi-technical organization, like you know, take the fashion house I was working with. Like, yeah, they have some web developers and such.
They don't have devs, they don't have like the ability to actually build those extra pieces. ⁓ it gets kind of aspirational in a hurry. to do these like really full featured functions without without some professional help. ⁓ and you really can't hire it right now either because it's all so fresh.
Like you see universities coming out with, you know, ⁓ technical AI, like etc. And that that will matrix those people will matriculate and they'll be on the market in a few years and you'll be able to start hiring for those. But we're in the awkward middle here where you you just quickly get out of your depth. Yeah, well awkward yeah, awkward beginning for sure.
Yeah. Yeah. And I and I hope that that that that that that's clear. I hope that that that's clear that this is not ⁓ this isn't even easy for us.
Like it it just changes so fast. And you know, we live and breathe this stuff. So so being able to roll with whatever that technological change is sort of the gig that we signed up for. But it's daunting for, you know, tech verticals that have had you know, static processes for a long time.
And you you could totally sympathize and understand why a CIO or a CTO would just be like tearing their hair out as far as figuring it figuring this stuff out. All right. Hey Matt. Yeah, fun as always.
⁓ I I hope that your your hacker house yields the things that you ⁓ you need it to yield. I think that that's a takeaway today. Like if you want to jumpstart your your your action, ⁓ you do need to develop that literacy. You need to get people people equipped with the right tools, you need to get a familiarity around it, and then you need to have good problems to solve that are that are something that are tangible and there's something that are doable.
And get people out of their normal context, get them focused on it, lean in and put a really strict time limit on it that this is what we're expecting on the other end. And honestly, what better way to learn than to dive in and do it? And it might end up flaming garbage at the other end, but I guarantee that it'll be part of the journey that allows you to start building things that are better and better. So yep, Matt, thanks for being on again.
⁓ you know, Matt's team is absolutely fantastic and they're on the cutting edge of this. He's at ⁓ rapiddev.ai, rapiddev.com.
of those domains work. Excellent. We have all of the things. ⁓ but great talking to you and ⁓ enjoy your time in Europe.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.