The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Modern CTO
Modern CTO artwork

Where Do We Draw the Line on Letting AI Build? With Zach Goldberg, CEO at Gruntwork

Modern CTO · 2026-07-02 · 47 min

0:00--:--

Key moments - from our scoring

Substance score

45 / 100

Five dimensions, 20 points each

Insight Density8 / 20
Originality9 / 20
Guest Caliber12 / 20
Specificity & Evidence9 / 20
Conversational Craft7 / 20

This episode explores the increasingly complex question of where and how AI should be deployed, drawing parallels to how society has learned to regulate technologies like knives through context-specific rules. Zach Goldberg articulates the central tension: AI companies like Anthropic and OpenAI face market incentives to release powerful models while the world expects them to prevent misuse - a challenge more complex than simple access restrictions. The hosts discuss how safeguards currently operate on actions rather than users (refusing "how to build a bomb" but answering legitimate questions), yet remain vulnerable to sophisticated circumvention by knowledgeable actors. The conversation shifts to practical enterprise value, with Goldberg detailing how Gruntwork uses Claude for code generation and building custom integrations (HubSpot workflows, personal finance dashboards) that would be economically infeasible to hand-code. Both speakers highlight a emerging paradigm: single-purpose applications built with AI assistance in hours rather than days, alongside the recognition that network-dependent platforms like Facebook still require traditional development. The episode also touches on AI's role in gaming (Steam's content guidelines), Anthropic's Fable model suspension following US government pressure, and tools like Cursor, LM Studio, and Superhuman that are reshaping how developers and operators build and manage workflows.

Key takeaways

  • →The most effective safeguards operate on the action level (refusing harmful requests) rather than user identity, making circumvention trivial for sophisticated users while ostensibly protecting against casual misuse.
  • →AI has collapsed the time-to-value for building single-purpose applications from weeks to hours, creating an entirely new category of software that wouldn't exist under previous economics.
  • →Enterprise value from AI integration comes less from code generation at scale and more from stitching together existing tools across fragmented software ecosystems (like connecting HubSpot, CRMs, and internal systems via Claude).
  • →The regulatory framework for AI remains nascent without established language, shared mental models, or clear dividing lines between acceptable and unacceptable use cases - leaving private companies to navigate competing incentives.
  • →Modern AI tools enable a shift toward distributed, personal software ecosystems where individuals maintain hundreds of lightweight applications rather than relying on centralized platforms designed for millions of users.

Guests

Zach Goldberg

Topics in this episode

ClaudeOpenAIHubSpotAnthropicCursorGruntworkFable modelLM StudioVertex AISuperhuman email

Questions this episode answers

How does Anthropic's Fable model relate to the Trump administration?

According to Anthropic's public statements, the US government asked the company to suspend access to the Fable 5 and Mythos 5 models after they were briefly available publicly, with Anthropic pointing to political reasons for the shutdown.

What's Gruntwork's primary use case for Claude and AI-powered code generation?

Rather than pure code generation at scale, Gruntwork focuses on building integrations and custom capabilities across fragmented enterprise software ecosystems, such as extending HubSpot's features or connecting multiple systems - delivering features that would cost thousands to build traditionally for just $3 in tokens.

How did Zach Goldberg bypass Claude's safety restrictions to unlock his PDF bank statement?

When Claude refused to provide tools for cracking PDFs, he worked around it by asking Claude to parallelize an existing open-source CLI tool, demonstrating that safety filters are vulnerable to indirect prompting by knowledgeable users.

What's the difference between vibe coding and the single-purpose app model?

Single-purpose apps are lightweight, individually-maintained applications built quickly with AI assistance for specific personal or small-team workflows, as opposed to monolithic SaaS platforms designed for millions of users.

What metric changed for Joel after switching to an Apple Watch with cellular?

His screen time dropped 65% after moving essential functions (Apple Pay, calls, texts, Tesla key) to the Apple Watch and limiting phone access to three scheduled hours per day, freeing him from constant social media interruptions while maintaining work connectivity.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

8 / 20

There are a handful of genuinely useful observations - the read-vs-write enterprise AI adoption distinction, the 'single purpose app' framing, and the software-reusability rethinking - but the episode is heavily padded with off-topic tangents (Apple Watch screen time, Fable/Anthropic politics, Wall-E, books-vs-audiobooks) that dilute the B2B signal considerably.

The conclusion we came from, we came to is there's actually a clear line in the sand where things have changed and where they haven't. And the distinction is read versus write.
I call it the spa trying to be funny. Rather than the single page app. It's the single purpose app.

Originality

9 / 20

The read/write framing for enterprise AI adoption hesitance is a moderately fresh lens, and the software-reusability rethink for single-purpose apps is worthwhile, but the bulk of the conversation - AI won't replace senior engineers, trust matters, vibe coding is fast now - recycles widely circulating takes without adding a contrarian edge.

do we care about something that already exists that can just be recreated with six minutes of thinking of opus thinking?
It would be, it's akin to I'm starting a new company and yes, I have Claude, and rather than have Claude write code in Java, I'm going to have Claude invent a brand new programming language

Guest Caliber

12 / 20

Zach Goldberg is a genuine practitioner - CEO of Gruntwork, maintainer of Terragrunt, author of The Startup CTO's Handbook - with real enterprise infrastructure experience at scale, but the conversation rarely drills into his specific domain expertise, leaving his caliber largely unleveraged.

We maintain terragrunt. Right. An open source infrastructure as code tool. Right. And there if I'm some, you know, platform architect that doesn't work at grunt work
We actually did a survey of CTOs, um, and CIOs at, uh, grownwork over the past several weeks about their adoption of AI

Specificity & Evidence

9 / 20

There are scattered concrete data points - $3 in tokens for a HubSpot feature, one hour to build a personal finance app, 65% screen time reduction, $30k startup budget framing - but the headline survey is only about a dozen leaders and most evidence is personal anecdote rather than rigorous business data.

we have a CRM, for example, at Gruntwork, we use HubSpot... I can now just build. It cost me about $3 in tokens.
relatively small in size? We did about a dozen liters. Um, the conclusion we came from

Conversational Craft

7 / 20

The host is personable and occasionally redirects well, but questions are frequently open-ended softballs ('What's your stack?', 'Are you going to update it?'), and the conversation drifts extensively into personal lifestyle topics with no productive pushback on sweeping claims about AI, books, or enterprise software.

Well, I'd love to talk about video games, but I think we're here to talk about CTO stuff.
You gonna run a prompt after this to have it update the book and deploy it to Amazon.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B56%
  • Speaker A44%

Most-used words

code20trust18value15claude14books14world13point12build12past11answer11knife11humans11software11access10story10tools10

Episode notes

What happens when the worst among us get their hands on the most powerful tools? Today, we're talking to Zach Goldberg, CEO at Gruntwork, about the messy edges of AI adoption. We discuss why "vibe coding" has created a new category of disposable, single-purpose software that never would have existed before, why the line between AI you can trust to read data and AI you can trust to write to production may be the most important distinction in enterprise tech right now, and why the real bottleneck on AI misuse has shifted from access to intelligence itself. All of this right here, right now, on the Modern CTO Podcast! To learn more about Gruntwork, check out their website here .

Full transcript

47 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Today we're catching up with past guest Zach Goldberg, CEO at Gruntwork, about all the intricacies of where AI use is best suited and lots more. You're listening to Joel Beasley, Modern cto. Well, I'd love to talk about video games, but I think we're here to talk about CTO stuff.

Speaker B: Game companies are technology companies. They've got really interesting challenges. The timelines, the crunch, the working with the artists. Like, I can imagine that is a really hard problem.

Speaker A: Not that I know about. Uh, it. We did a series like. This is the cool thing about the show going for 10 years, I could. Like five years ago, we did a series with some different game developers, I think, like Activision or Blizzard or something,

Speaker B: and like, how their world must have changed as well in the past year and a half. And can you imagine, I don't know, this maybe segues, uh, a little closer to my neck of the woods. Like, there's such a public backlash to the use of AI in, you know, games, right? Like, you know, Steam has a disclaimer, AI was or was not used in the production of this game. Even if you used it to make one asset. Uh, now your whole game is tainted and, you know, some percentage of your audience is no longer interested in your product, um, because of, you know, a shortcut you use to make some inconsequential thing. Uh, it's the total. The whole game has changed. Right. The world we live in now is totally different.

Speaker A: Where do they draw the line? Because, like, you can apply AI so broadly if you would like. There. There have been these auto. Before LLMs, there. There was major progress and the auto generation of worlds happening now. I don't think they would consider that

Speaker B: AI in this particular case. I, um, think it was Steam, Valve that published guidelines on it and they did draw a very clear line. It was something like. If you use generated content to plan or to conceptualize, that's fine, but anything that's. I. I'm probably misquoting this, but anything that's actually shown to a user had to be created by a human procedurally. Yeah,

Speaker A: that's right. You can't. Yeah, it's just. Honestly, I think it's. I don't know.

Speaker B: We're.

Speaker A: I think we're about to enter a bit. When do you think that's going to become political, by the way? Let's. Let's make some, Some future bets. Because I. Everything becomes polar.

Speaker B: I mean, is it not already, like, Fable is not available to the world because somebody in the Trump administration said Something to Anthropic. And they pulled it out. Right.

Speaker A: Did they really?

Speaker B: Are you not familiar with this? Yeah. So Anthropic's latest model called Fable, which is the next version of the world.

Speaker A: I got access to it for, like, a day.

Speaker B: Yeah, exactly. And then Anthropic me, you know, must have been public for 24 hours, 48 hours, something like that. Uh, and then somebody in the Trump administration. And at least this is how Anthropic tells the story. And I have no reason to believe that's not.

Speaker A: So you read it from them? Not like an article.

Speaker B: Yes. The Anthropic themselves is directly pointing the finger at the US Government, saying they asked us to turn it off. And so it is still off a week or two weeks later. Um, presumably for some political reasons. Yep, there it is. I'm sure that one more will tell you all about it.

Speaker A: Yeah. Interesting. Okay, so.

Speaker B: And politics and AI are one in the same.

Speaker A: Yep. Okay.

Speaker B: It's in the hero headline statement on the US government's directive to suspend access to Fable 5 and Mythos 5.

Speaker A: Yeah. I'm curious, because it's always nuanced, right? Like, there's always details. Uh, and so I'm. I'm always curious to know, like, okay, like, you know, if I wonder what the. The. The case was. Did, uh, they share that? Did they say, like, this specific use case is why? Or did they just say ambiguously that.

Speaker B: I would just be speculating.

Speaker A: Yeah, we would just be gu.

Speaker B: Yeah, I'm not aware that they did.

Speaker A: Yeah, there might have been something serious.

Speaker B: You know, I think what I find really intellectually fascinating is. I like your point. Nuance is the, like, the name of the intellectual conversation, almost philosophical conversation we're having right now. You know, as we started a minute ago, AI could be banned because, like, oh, AI bad. Because, you know, some ecosystem governor, governance body was like, oh, we don't want AI or whatever. Um, or it could be, like, fully adopted. And, like, our company uses only AI to write code. And the reality is, like, the correct answer is, well, there's nuance and shades of gray and, like, different things to take into account. And I find it exciting because we don't have a great language that we have not yet evolved systems and frameworks and language to describe that nuance that we all share. Right. So what is an acceptable usage and what are the consequences of that? And how do we talk about. This is a very actively evolving culture sort of as we speak? I, uh, don't think anybody really has the right answer. For that yet.

Speaker A: No, you're exactly right. The words will mature. We'll learn the discussions that are important. We'll find the dividing issues, we'll find the common issues. Like, for example, I like to go back to the, uh, the knife. Right. Because you can use it to kill someone. You can use it to save someone's life. It can be very contra. It's banned in certain places, but allowed in others. There's versions of, uh, it that are banned at the airport. Like, you can't get through security with a bladed knife. But inside the airport, right? There's auto closing.

Speaker B: Knife auto closing.

Speaker A: There's all of these, like. So the concept of a knife exists and all these various different context with all of this stuff around them. The size of the knife matters. Like, a pocket knife can be so many inches, but if it's longer than that, it becomes this other. It.

Speaker B: It.

Speaker A: It's just. There's so much to it. But that's old technology. So as humans, our generation, we grew up with it. That wasn't even, like, really controversial. It' like, oh, this is how it works.

Speaker B: You just know you don't bring a kitchen knife onto an airplane. It's just a thing, you know?

Speaker A: Exactly. So the generations. But then again, like, if you're in the 60s, you're bringing a pocket knife on the airplane because you have to carry a pocket knife because you need it for various activities throughout your day, and no one even questions it. So it's. For me, I'm interested in looking at that analogy, seeing how far we could take it, like, how applicable it is. Uh, but then also I feel like we're the adults when the knife was coming about.

Speaker B: Yeah.

Speaker A: You know, and then our kids are us, you know, they're the generation below that's just gonna grow up and be like, oh, this is just how it is. So it's like, how can we be good stewards of this, uh, and do the right thing? And the answer is it's incredibly unclear and it's difficult and we're probably gonna screw it up.

Speaker B: Yeah. And I think, you know, I don't work at Anthropic, um, or, you know, Google's AI division or whatever. So you have to imagine, like, my concern as the person who doesn't work at these companies is the adult in the room right now is we're sort of depending on the employees of these companies to make good decisions about safety, about governance, so on about when is it a good time to release fable and mythos, so on and so forth. Um, and the concern as an external party is where's the incentives? Right. Like Anthropic is in a race with OpenAI and Google and you know, these other companies making these big models. So there's, you know, there's a capitalistic market incentive going on to be the best, to be the best price to get the adoption, get the users, whatever. Right. Uh, but at the same time the rest of the world is looking at them saying, please don't screw this up. Right. Like let's, you know, make sure these things are not getting into the hands of the wrong people and you know, causing major security exploits or whatever other bad things happen. Skynet in the future, uh, with, you know, the direction AI is going and that, that feels very tenuous.

Speaker A: Are you of the mindset that like we should have a hundred percent unmoderated models and they like, like we shouldn't have pulled or Anthropic shouldn't have pulled Fable, like we should just let everybody have access to as much information as they possibly can. Or are you, like there are situations where we should limit it?

Speaker B: You know, it's, it's an interesting. I don't have a clear black and white answer to that question personally. Right. Nuance.

Speaker A: That's what it is.

Speaker B: Yeah.

Speaker A: Because my, my default is freedom. But then there's like, there are definitely situations.

Speaker B: A crowded movie theater.

Speaker A: Yeah.

Speaker B: Right. Like what is the equivalent here? What is the. Our social contract says we agree to bind ourselves to certain limitations that we all agree are reasonable. Right. I don't know what the equivalent of fire. And like, obviously I don't want a terrorist nation having access to the most powerful models that can exploit all of America's financial systems or power grid systems. Right. Like we both probably agree that would not be a good thing.

Speaker A: Right.

Speaker B: Uh, but how to do that? It's not quite as simple as saying don't bring a knife on an airplane. It's like once it's out there and it's available, like how does Anthropic or OpenAI whomever, like it's not like deep state or not deep state terrorist nation is going to like raise their hand and say, I'm a terrorist, I want your AI.

Speaker A: I think one unifying area that we could, we can just discuss a little bit is you know, our world is designed where you and I operating outside in reality going about our day. There is all the people exist from just got out of prison for murder, just got out of the mental institution all the way up to uh, prize winning physicist. We've got this huge spectrum of people, and I'm constantly surprised when I see, like, different signs that say things because, like, I saw. I saw a sign the other day in the bathroom that said, don't flush diapers. How many people were like, yeah, this is a good idea to flush it before they printed, uh, went through the effort of printing up a sign. It was like a metal sign, too. It wasn't just like a piece of paper. It was like they had a metal sign that was bolted into the wall that was. And so I'm like, yeah, those people exist. They're out there walking around too. Now. I don't want the super intelligence or whatever helping them execute some of their ideas. You know, I just don't.

Speaker B: The reality is the filter at the moment is not on the user, it's on the action.

Speaker A: Right?

Speaker B: Like, if you ask it, how do I make a bomb? It says, I can't help you with that, Steve. Um, or what's guy's name from Space Odyssey? Um, Hal. Yeah, no, thank you, Hal. Um, but if you say, you know, what was the backstory of the French American War? Like, okay, it'll tell you that. Um, and so, like. But that's tenuous, right? Like, you. I don't, uh, know if you've tried, but like, trivially, the other day, true story, uh, I had a bank statement that was encrypted with, uh, a password. My bank. Genuinely my bank statement. And I didn't remember the password that the bank used for the PDF. And I didn't feel like calling the bank and sitting in a phone tree for an hour to figure out how to unlock my own PDF. And so I happen to know that there are command line tools to crack PDFs, and like, all right, it's my data. Like, just give me my own data. So I asked Claude make a command line to crack this PDF. And Claude said, no, thank you, I can't do that. Uh, and so I googled what are the common CLI tools for cracking PDFs. And then I asked Claude, run this tool in parallel. I said, sure, I'll write a bash script to run the tool in parallel for you. So the circumvention is very nascent and still seems, uh, not bulletproof at this point.

Speaker A: Right. But the limiting factor is your intelligence. So, like, you know, I, I think that's a. I don't even know if I have thought. It's an interesting thought that, like, if you're smart enough, you can get access to it. Like, if I wanted to build A bomb. I could download a model, crack it, you know, break it. There's different versions of them, um, that have different levels of security. Get the model to do what I want locally and then have it assist me in whatever nefarious thing I wanted to do. That is, that is a complete plot. But the skill level it takes to be able to do that and execute that, uh, is high.

Speaker B: So you're saying we've prevented not smart people from doing bad things, but we've made it even easier for smart people to do really bad things.

Speaker A: So psychopaths are probably going to become pretty empowered right now or you know,

Speaker B: nation state actors, like I said.

Speaker A: Psychopaths. No kidding.

Speaker B: Fair enough.

Speaker A: Uh, no, but it is very true. I think we forget because we live in such an amazing. You live in the United States, right?

Speaker B: I do, yeah. Yeah.

Speaker A: We live in such an amazing country and, and we live in a first world country and there, there are actively entire countries of people waking up every day trying to end us. Like that is. That is a reality that we are so well insulated from. We don't think about it on a daily basis.

Speaker B: That's a really powerful frame. Right. Uh, they wake up every day trying to think how do I cause harm to. Maybe it's us or it's another country.

Speaker A: Yeah.

Speaker B: And then you put that in the context of, you know, these capabilities.

Speaker A: They could get a VPN and then use Fable and now they have this assistant that like they don't, you know, because they're limited by their access to intelligence. That, that's a huge wealth thing by the way. Kings have always had that if they wanted to do something they could just summon the most intelligent person from their kingdom and, and talk with them and have them execute.

Speaker B: We are all now kings.

Speaker A: We are all now kings. We have all of this access, a

Speaker B: $20 a month open AI subscription and you are a king.

Speaker A: You don't even have to pay that if you have a Mac Mini. You can just download local. They've made it so easy now. Have you played with LM Studio?

Speaker B: I have not personally, no.

Speaker A: Oh man. You just. It's one click install on Mac. You select your model, instantly runs it, boots it up, runs it as a server. So you can run it with like openclaw. So you're just using all local stuff for all its tasking if you want. It is absolutely fantastic.

Speaker B: Yeah. Now, um, I haven't done all that much in the way of personal automation, but I am very double down at like at scale code generation, like at scale, like Compared to what's in.

Speaker A: What's your stack?

Speaker B: Uh, it's for the most part Claude, just straight Claude, Um, little bit of the Claude projects. Uh, and then it's homegrown tools to stitch together things that look like Claude or Vertex. AI and other important softwares in our ecosystem. Um, is where I have found lots of enterprise value being generated. Like getting, you know, we have a CRM, for example, at Gruntwork, we use HubSpot. HubSpot has lots of great capabilities and there's 100 features I wish it had that it doesn't have that I can now just build. It cost me about $3 in tokens. Um, and features that integrate data across multiple places or take actions across multiple different systems. Um, and from that perspective, it's been phenomenal over these past six or eight months. The things we can do, just integrating better across our software ecosystems.

Speaker A: Yeah, I have created more applications in like the past six months than I have in the past five years. I need a utility for something. It is off. It is almost become faster, Zach. It is for me to build the application and with cursor and deploy it to Heroku or right onto my iPhone, whatever the app needs to be, than it is for me to spend the afternoon hunting for potential solutions to see if I need to build my own thousand percent.

Speaker B: Thousand percent. True story. Literally this past week, for years, I've managed my finances in empowers personal capital. Right. It's an online dashboard. You link your bank accounts, it has views. And I've always just been frustrated with the data visualization. Like, over the past five years, I've probably made half a dozen support tickets with this company, asking for small features and tweaks here and there to make it easier for me to understand my own data. And it's like I had the same frustration a week ago. And I'm like, you know what? This is just a single pane of glass SaaS tool that pulls in some data. An hour is what it took me to vive code a thing that replicated every single feature I needed. And now I can just, with two sentences, add any capability I want. And I think this is unbelievable that we can all just do this, um, and, you know, millions of these apps now.

Speaker A: You know, it's interesting, as you just said, that I've had a lot of conversations where people, you know, over the past year or two, things have matured quite a bit, but there's a lot of, uh, punching down at the, at the vibe code, like the, uh. And I do it too. Cause I mean I've built enterprise level applications and I'm. But then I, I, I started to think as you were talking, like, what if the future like Vibe coding does work for like personally, like if you just need to do this very specific task and you need it.

Speaker B: I call it the spa trying to be funny. Rather than the single page app. It's the single purpose app.

Speaker A: Okay. The single purpose app. Yeah. So what if we don't need things we don't need m. Not all things. We don't need most of our applications to run at scale, we just need them to run for us.

Speaker B: Mhm.

Speaker A: What if that's the future? What if my future is a hundred little applications that my agent is maintaining for me versus me logging into like a Facebook UI that's being maintained for 10 million people?

Speaker B: I think this is a new category that partially overlaps with the world of software that used to exist. I never would have spent probably the hundreds of hours to hand code a replication of a personal finance app. Right. But now that I can do it in an hour. So that's an application that does exist now that wouldn't have existed before. And so in the universe of software that could be written, we were only writing 50% of apps or whatever percent that humans wanted. And now we have expanded that bubble. Uh, which is to say there is still a very, my belief is there's still a very large set of apps that are not replaced or replaceable by Vive code app. Right. Facebook still has value because there's still uh, or if you believe Facebook has value or Instagram has value because there's still a network of a billion people that want to share photos. Right. And I can't Vibe code that. Right. Because that does have to actually exist that a billion people use.

Speaker A: But their interface though. Yeah, that's like. I think uh, they're gonna become more like a data store where the API is powerful. Yeah. Where I'm just. Everyone's plugging into the API. I don't even know, man. I've been running an experiment for 15 days now. So I usually don't like to talk about stuff until I've done it for a while. But um, you've inspired me. I got an Apple watch.

Speaker B: Uh huh. It's very pretty.

Speaker A: First. Thank you. That was the point. I was like, I want to look beautiful. I don't care about the technology.

Speaker B: No.

Speaker A: I got this cool green band though. Uh, but the purpose was I want to. I noticed that I was having self control issues, overriding my time limits. On social apps and I have to use them for work. Like I have to check the stats, I have to make sure the post happened. But then I would go in there to make sure the work happened and then I'm doing 15 minutes of like just junk, just junk. And so I said I need to do the screen less. But the thing is I need access to my wife, like for, we got the kids and like we need to be able to call each other. So I, I, I went with the apple watch about 15 days ago with the cellular and everything. My screen time is down like 65%. I just, I can leave my phone at home. Like I just put it away, I just put my phone away unless if there's something I need my phone to do. And so like my, my what it can do my Tesla, it's my Tesla key. It can do my Apple pay, it can call people, it can text people, like anything that I'm required to do.

Speaker B: Does it. Were you able to integrate like your important work related social media stuff into the Apple Watchers that still requires.

Speaker A: No. So I've just, I have compressed that down to scheduled screen time. So now I've got like a three hour block every day where I can use my phone and my computer and I have to get everything I can get done in that. You know this as an entrepreneur you have to figure that. So I have to get everything I can get done in that time and then I use a lot of that time to make stuff more efficient. Like I recently found Superhuman, Ah, email. Have you used that before?

Speaker B: Mm, mhm.

Speaker A: Oh my gosh, dude, Superhuman. Do you currently use it or do you just play with it?

Speaker B: I'm not currently a user, but there's folks on my team who use it every day. Yeah. And they rave about it. I've dabbled with it.

Speaker A: Yeah, it's got this new stuff in it, uh, where you can like tell the AI, like look out for these types of emails and it then puts them in a priority section for you. Brilliant.

Speaker B: You know, uh, so that, that's the next question though is when do you hook up your Mac, Mini, openclaw, whatever to your email? And rather than relying on Superhuman's prompts, if you want to customize it and do your own workflow, Email's another API to integrate into.

Speaker A: Well, absolutely. Well that's how I started actually.

Speaker B: Okay, I started. So Superhuman's better than what you got?

Speaker A: Yeah, well for, for the intended use case, M, uh, what I was doing had more breaking points than just like I built this Infrastructure to do all of that to API and then Gmail to read them all and then to prioritize and move them into. But then the Superhuman had a better interface. It had snippets for quick replies and then it had keyboard shortcuts and it was so polished that I was like, yeah, I'll give them 40 bucks a month. This is saving me time. And now I don't have to maintain. Which is another behavior I think is going to happen too a lot with people. Zach. I think we're going to build stuff, but like if your personal finance app, if you went to go prompt it and then it was like, hey, there's also this out there that this other person did that Joel built, that achieves that. Would you prefer to use that and not maintain your own project? You could click, yeah, let's give it a shot.

Speaker B: Yeah. Uh, it's interesting how like I had a very fun story in developmenting, developing this personal finance toy. At one point it needed API access to some banking thing and the AI Claude told me, oh, there's a wrapper that already exists around the API. Uh, and I said, okay, great, use the wrapper. It's probably been tested. It said, no, I'm going to just build it from scratch because I'm going to use TypeScript and have types and the other one isn't maintained all that great. It didn't quite say no, but it pushed back and I was like, that's interesting. So given the opportunity to use something off the shelf that has some sense of validation, but maybe it had some minor trade offs. It chose to just, no, I'll just write the 500 lines of code from scratch. That was its instinct as cheaper. And one could say, well, perhaps it just wanted to burn those tokens. Um, you know, my instinct was like, just use the thing that I know works. Um, perhaps I shouldn't be surprised that its own homegrown 500 line version actually did work on the second try. Uh, so it's this idea of software reusability for these single purpose applications. I think is we really need to rethink that. Um, do we care about something that already exists that can just be recreated with six minutes of thinking of opus thinking?

Speaker A: I think that's the question every SaaS company is asking themselves right now.

Speaker B: Where is real value generated with software today? I don't know. Do, uh, you have a thought on that?

Speaker A: No compliance. I think that's gonna be a huge one when there's regulatory and compliance reasons. Why, like when the technology could one Shot prompt something. But I can't because I have to go to the government to get some form done or something like that. I think that's gonna be, uh, like if I was investing, I'd probably invest into that because I think long term that's got a really good shot.

Speaker B: Yeah. So at gruntwork we sort of have a perspective. Um, there are things we do that are vibe codable replacements. Right. Because they're table states as part of a broader ecosystem. Um, but where there's still lots of value is foundational tooling that's expected to last. Right. Where there is value in large amounts of people all understanding the same software. So, uh, yeah, there is. We maintain terragrunt. Right. An open source infrastructure as code tool. Right. And there if I'm some, you know, platform architect that doesn't work at grunt work, I work at whatever company. Um, even if I'm using an AI, I'm still relying on foundational tools. Like I'm writing code in Python or Typescript or Terraform or Open Tofu, whatever. Right. In terragrunt. Uh, and I'm going to need to hire other people and I'm going to use AIs and all of these people and AIs need to have familiarity with these tools. Right. It would be, it's, it's akin to I'm starting a new company and yes, I have Claude, and rather than have Claude write code in Java, I'm going to have Claude invent a brand new programming language and all of my programming will be done in this new programming language that's only used at my company. And so now I'm dependent on exclusively the AI model's ability to use that tool. I can't hire people easily who've mastered that tool. And there's no guarantee that the next version of the AI will be as competent with that tool. And so there's still a lot of value in these common layers that we reuse. That needs to be well understood. Uh, and those things, because the blast radius is so large, need really deep thought as to how they should work. Uh, you don't want to just say, hey Claude, what's the next feature of OpenToFu? You want to actually think, how is it going to work for the million existing users and how is the next million users going to use it? And like that's still a very contemplative process.

Speaker A: Uh, I've noticed that I have made the AI write everything that could be written in Rails. In Rails, because I've got a decade of experience in Managing Rails projects and I can catch it in little mistakes and understanding how it would scale in production or under load. Uh, and so yeah, I think you're right. There's definitely value in being able to talk about it with other engineers. But at the same time I'm not uh, an iOS native developer and I did an iOS app and I wrote zero lines of code. It is a fairly complicated app that had to kind of uh, it couldn't, it can't be published to the App Store. For what I needed it to do. I had to kind of like work around some things and it could be a local app for my phone. Uh, but I didn't write one line of code. It was a hundred percent just me talking to the thing in cursor and then it running and then me using it and then providing feedback. And I think we're there, I think we are there for, I mean I don't think, I know we're there for individual personal use.

Speaker B: Mhm. Yeah. And that's, that's this, this is where we're missing the taxonomy. Right? 100% agree. For the single purpose app, the individual app, we're absolutely there. Right? Half an hour. An hour. With an AI, you can build your iOS app, Android app, web app, whatever it is. Uh, but quite clearly I still pay full time gruntwork. The company I work for still pays full time software developers to develop enterprise software. And so I'm very confident I could not replace my team with only AIs. Uh, but I have a difficult time explaining the line because obviously there's a spectrum here and what is the label on the x axis in that spectrum? Right. Is it complexity? Is it size, is it scale, is it number of users, is it how long I expect it to last? Um, like there's a number of ways to look at that and answer the question, could an AI do this on its own or does this still need a team of humans overseeing and collaborating and supervising? We could point to like a number

Speaker A: of, I would label that like what's coming up to me right now. I would label that trust to achieve outcome. So there's a number of outcomes that have to happen at your business and you can, you know, because you're an entrepreneur, uh, like me, we own businesses and employees and stuff. We know that you can, we know how much load a single person can take, right? Like you can tell when they're overloaded or not. And so we say, okay, I have trust for that person to achieve these two outcomes. I know if I put a third that's gonna stretch on, quality is going to drop. But I know that in my mind I wake up in the day, I know these two things are important to the business. I know Mike's got it. Okay. And then you, uh, as an orchestrator you have many of those items that you will then spread across multiple people. And then you have these direct personal relationships with trust with them to achieve the outcome with the advanced AI. And I think that's what is happening.

Speaker B: Yeah.

Speaker A: Because if you did trust the AI to achieve all of the outcomes, I think we're not there yet. This is like full self driving with Tesla. Have you ever done Tesla full self driving?

Speaker B: Yeah, of course.

Speaker A: You've got to like build up this trust.

Speaker B: It's not trust. I think trust in and of itself is multidimensional. Right. I trust that if I ask AI to implement feature X, will it implement, you know, and I define half a dozen acceptance criteria, will it get those six acceptance criteria? I think for reasonable circumstances, like there could be some trust there. However, my definition of those acceptance criteria is almost certainly imperfect. There's probably six more AC that should be there that I didn't think of. And if I assign that to my senior engineer, he's going to find those 6ac, he's going to think of, okay, we'll build this feature now. And by the way, you didn't think of these other edge cases of you know, real world user workflows that might happen, but also there's other features coming down the road and you know, the how is it going to interface with X, Y, Z? Um, I have no trust that the AI sees around corners in that. Correct.

Speaker A: Uh, and that's why I focused on trust to achieve like an outcome because the outcomes are very complex. It's not a feature. It's not, it's not just like one specific thing. It's this culmination of all of these things in their environment, whether they're interacting with customers and going with their gut after talking to multiple customers. So like I think the human's role right now is achieving outcomes with the advanced technology. And I mean you could probably use those same words 10 years ago. But the thing is, what's compressing is the number of humans that you need to achieve the outcome.

Speaker B: Yeah. And there's still the same problem we've always had of defining the outcomes you want clearly knowing what outcome you actually want to get to, uh, uh, understanding what to build harder than actually building it. And that's more true now than it's

Speaker A: ever been before taste, like trusting their taste. When you know Mike's got the problem, you understand that there's going to be stuff that you can't even see that are coming from left field, things that are going to pop up midstream of implementation, all this stuff. And you have to trust that he has the right head on their shoulders to make these decisions.

Speaker B: Yeah. And I do think the other. So there's. Trust is a big part of it. But to add another element to the mix. A million tokens in context is enormous. But I think empirically it still pales in comparison to what a human can keep in their head and the quality of the context that a human keeps in their head. I still think we're several orders of magnitude off. Gruntworks software compared to Fortune 500s, I'm sure, is very small in terms of lines of code and total amount of complexity. But it's still huge compared to, you know, the individual application and completely terrible, like unusably poor when thinking about the system at large. Right. So it's still like, back to your point, like defining an outcome within a context that is understandable within the AI scope.

Speaker A: Have you looked at the systems that are designed to do that?

Speaker B: We've played with it a little bit. Um, I wouldn't go so far as to say that, you know, I'm the expert on the larger system, um, tooling that's out there, like Codex spaces, things like this.

Speaker A: Uh, Josh, who is that? That company that did a sponsorship with us. That's like, that's what they were specifically doing, but they were building tools specifically for that use. Case of massive code bases at enterprises to understand them and be able to work with them. I didn't play with it myself, hands on. Um, but I talked to some of their users and customers and stuff and it was pretty.

Speaker B: So you, you feel confident whether it were there today? We will be soon enough.

Speaker A: I think we have. I think we've got all the tools in place that we are and I think the momentum is there. I think look at the past. How wild since our conversation in 2024, like, we could not do that personal finance thing that you did in an hour.

Speaker B: 2024, glorified autocomplete. Like, do you remember glorified autocomplete two and a half years ago to now go spit out 3,000 lines of code.

Speaker A: I keep, I keep trying to challenge it more and it just keeps delivering. I'm just so impressed by it. Have you seen Blitzy? Yeah. Oh, so Blitzy is pretty cool. They, they have A number of example use cases, but they essentially write like if you have a, you're an enterprise, you're going to do this product, Sprint, it'll get you like 90% of the way. There's like writing all the code and planning it all and everything. And then you just have to go in and do like the last couple human things.

Speaker B: It's pretty, pretty fascinating with these kinds of tools. I wonder what is it they actually do. Right? Is it like, where, where is the value being created here? Is it chaining together existing agents and some like workflow, uh, glue on top of it all? Like, is there a defensible intellectual property there or is this really just gonna be part of the next version of Claude in 18 months?

Speaker A: Both. I do, I think that they, you know, you've got, that's the fun thing. So one someone gets a little bit farther ahead and then they see that and they're like, oh, that's the thing. And then all the competitors can catch up in like a month's time. It's wild. Like, have you seen the release of gems and then projects and then like just the maturity of these systems accelerating at such a degree. It's, it's, it's mind blowing. So it's almost like the, whoever can have the coolest imagination and then you watch everyone like run to it. Like a breath of fresh air. That's a better word. Yeah. Uh, then all you have to do is say, go tell the model that you want your thing to do that now.

Speaker B: Yeah, yeah. So to your point, so I'm not an investor in blitzi, uh, maybe buying, buying puts on these small company stocks because to your point, like it just shows up in, you know, your mega cap tech company in three months from now.

Speaker A: Yeah, but then that's where trust comes, right? So if you're an enterprise and you're like, okay, there is this company Blitzy that has done this work with a lot of other enterprises and sectors that are sensitive, like maybe HIPAA compliance or banking or something like that. I would rather bring them in than my team try to do it with Claude because it gives you this sort of, uh, as executive running hundreds of millions or billions of dollars of value, it gives you this peace of mind back to trust. Like, I would rather have the people who've been doing it since you could do it, uh, than start it myself.

Speaker B: Yeah, it's interesting. We actually did a survey of CTOs, um, and CIOs at, uh, grownwork over the past several weeks about their adoption of AI and Kind of the build versus buy question. Um, you know, building is now cheaper than ever. Right. And so, you know, has the equation changed? Like, do we see changes in patterns? Um, relatively small in size? We did about a dozen liters. Um, the conclusion we came from, we came to is there's actually a clear line in the sand where things have changed and where they haven't. And the distinction is read versus write. If there's a problem that a company has, that's basically information aggregation. Right. Like pulling data from my monitoring system, from Datadog, from aws and maybe like reading notion docs as well and synthesizing that into a thing like this sort of read heavy use case, Vibe coded everywhere. Everybody's got an internal platform with agents that speak all their tools. With writing, we, you know, actually go change my infrastructure, go deploy this application. Uh, we still see more hesitance to Vibe code, um, at Enterprise there, at least based on, you know, our limited data size. But it makes a level of sense. Right. To your point about trust, writing has, you know, you trust that it's going to work and it's not going to cause problems. And so you don't necessarily want to shoot from the hip on that with something that somebody generated at 9 o'clock last night. You want the vendor who's willing to, you know, put their reputation behind it, a warranty behind it, whatever, an sla.

Speaker A: Mhm. I think you're exactly right. I just like you and your personal finance app. You know, if you're just reading data, it's easy to spin it up real quick.

Speaker B: Mhm. But you better believe the very first thing I checked though, as I started this exploration was what are the right verbs in these APIs that uh, this application will have access to? You know, obviously the assumption is we were looking for zero was the correct answer to that question.

Speaker A: Startup ctos, you wrote a book about it. Let's give a shout out. What's the name of the book?

Speaker B: The Startup CTO's Handbook.

Speaker A: Startup CTO's Handbook. Now you wrote that pre2024, correct?

Speaker B: Yeah, it's, it's sort of a mark of pride that the book was published just about around the time ChatGPT became commonly used.

Speaker A: Yeah.

Speaker B: Uh, so 0ai in the production of that book.

Speaker A: And how do you think has any of the advice, do you think it's changed? Would you, would you change? Are you going to update it?

Speaker B: I do think there are missing chapters. There are new chapters that need to be written about leadership in the world of AI. Um, but I will take a stand that the existing chapters dealing with leadership and people management and how you do hiring and how you think about product development from a putting the user first, those things don't change. The nuances perhaps of running a hiring process are perhaps different in a world of AI generated resumes, but the core values of what it is to think about a hiring process and what your optimizing for. Uh, and when you performance management internally leading a high performance team. I don't think that has changed. Um, with the advent of AI how you interact with humans should not change because the humans are now using faster tools.

Speaker A: What's the chapter that you think is missing that you would like to add to it?

Speaker B: I think it's predominantly how we think about staffing early stages at startups. The there's been a, a common question I get asked is you know I'm an entrepreneur, I want to start this new company. What's the, you know, I've got $30,000, how do I spend my 30 grand to get a prototype? Uh and you know the common choice is well I can hire somebody expensive in North America or I can go higher or offshore for 110 the cost. Um, what's the right answer? Right. And uh, I think I would often advise people like spend your money as frugally as possible to get to the mvp because whatever it is you're building, you're almost certainly not going to still have in two years from now. You don't know what your company is on day one. Uh, and so don't overspend and over build early on. That's probably changed almost certainly to you should be by coding the first version of your applications, right. And you should be putting it in front of users and that should cost basically nothing. Um, and what is all the nuance to getting there? What is that process of vibe coding and getting in front of a user and iterating to find value, um, early on. And it also comes back to our earlier question of can you generate value with just a SaaS software application nowadays? And like what is your angle to not be replaced to your point by some other company who can also vive code, you know, 60 days from now. Um, so I think there's a lot to be said about thinking about value long term and thinking about how you use software to make people's lives better uh, in this world. Uh, so I think a few chapters on there is probably warranted, um, as far as your role as a cto,

Speaker A: you gonna run a prompt after this to have it update the book and deploy it to Amazon.

Speaker B: No, I confidently will say that, um, I care about the quality here and, uh, my level of trust for an AI that is very low. And so at some point, I'll take a month or two off of work and really think about it and talk to other and put some blood, sweat and tears into making something quality there.

Speaker A: Do you think humans are still going to be reading books in 10 years?

Speaker B: 100%, 1,000%. Why, um, why does somebody read a book? What are the use cases? What are the user stories for? Why somebody reads a book. Right. I'm 14 years old in high school, and I have a math test coming up, and I have a math textbook. I think, am I going to use an AI to teach me math, or am I going to read the textbook? There's no right or wrong answer there. Like, both is almost certainly the answer. Ah, why do I read a book? Maybe I want to be entertained. I'm sitting at a beach on vacation and I'm interested in a story. Right. Humans are wired for story. We enjoy story. It activates all parts of our brains and empathy and connection and identification, all these things. Um, I'm not going to use an AI to just vibe code a story. I want to go read Brandon Sanderson. I like the way he tells stories. Um, and we could probably have, you know, we could enumerate 20 of these, and I think some of them sure could be replaced. Uh, but, you know, bringing our conversation full circle and the nuance, I think, yeah, people will definitely still read books. You disagree? You think where books are gone, save all the trees?

Speaker A: I don't. I think habits are hard to break. So the people who really like books. I'm not a book person. Uh, I'm a story person, and I'll do audiobooks and movies and stuff like that. But, uh, I haven't ever really been. The only time I've gone to books is when I needed, like when I was learning programming, I needed to learn how to. So I had no problem going to the books to learn. Now I would just ask the AI. You know, it's like, oh, why? You don't have to wait for the book. You don't have to wait for the podcast interview. You can just go talk to the

Speaker B: AI now for adult learning. Absolutely.

Speaker A: Yeah.

Speaker B: Um, it's interesting. So, like you, I've been an audiobook listener for a long, long time. Um, and I think there are. There's a culture, a subset of the world who believes audiobooks are not as good as real books. Um, and actually there have been a number of peer reviewed papers published in the past couple years looking at this question, right? Do you learn or do you absorb or whatever verb, as much information or as well with audiobooks as with paper books? Uh, and the answer, surprise, surprise, nuance. Uh, when it comes to hard intellectual subject matter, that is, um, uh, the adjective they use where basically it connects and builds on each other. Then you want paper, you want philosophy, mathematics, physics, where to understand physics concept X, you have to also understand A, B, C, D all the way through X. Uh, whereas for story, audiobooks and paper books actually are absorbed and people answer questions and have the same retention. Very comparable. Uh, and uh, the mechanism that is proposed is with uh, buildable material you naturally go back and reference. You're looking at this page and it's like, what happened on the prior page? Like, how did we get here? I didn't understand that concept. Very easy to flip back, rescan two sentences and okay, now you're good to go forwards, mimicking that process in audiobook. Like, sure, There's a rewind 15 second button, but how often do you push it and does that actually take you back to that one sentence that you missed that was critical to understanding the next phase? And so the hypothesis is that for casual listening, for things you're just absorbing, for fun and entertaining, it's okay to be distracted while you're doing it. Audiobooks are excellent and just as good as paper for things where focus is required and learning is the paramount objective. Paper still has an empirical edge. So the research says at this point.

Speaker A: Yeah, that makes complete sense. Right?

Speaker B: Seems intuitive.

Speaker A: Yeah, 100%, I think. Yeah, let's. I still think books are going. I think they're gone. Uh, I think, I think you heard it here.

Speaker B: Books are dead.

Speaker A: No, I think it's, I think it'll take a couple generations and they might come back, they might make like some small resurgence or whatever. But the thing is like I can go learn something interactively with an expert on the topic and for me conversationally, that's just how I consume it. So I'd say for people like me, uh, that, you know, books have been dead for a while.

Speaker B: So I disagree with you. However, I will also argue with myself. When was the printing press invented? 15th century, something like this?

Speaker A: Don't know.

Speaker B: Right. So like, how long have humans been reading books? A couple hundred years. Right. How long have humans been telling stories and learning things? Tens of thousands. Yeah, right. So books in like the grand history of, you know, way humans communicate with each other is actually a very modern invention. And so is it possible that it's a blip?

Speaker A: It's possible that it's like a CD and there's just like a better way to. For it to look. We're all obviously going to neural implant and be able to stream thoughts. So it's not even.

Speaker B: Wall E is the future wall.

Speaker A: Well, I'm going to be the only person on board that spaceship with a 6pa. I'm not. I'm not letting my body go.

Speaker B: Your discipline is admirable, sir.

Speaker A: I'll be running that ship.

Speaker B: You know, I don't know if I could, you know, have the same life expectancy and just sit in that couch. Man. That's tempting. Ah.

Speaker A: Uh, no, no, I've been overweight. Oh, you feel, like, horrible. It's not. Have you been like.

Speaker B: No, but AI is going to solve that problem. We're going to have the pill that allows you to be overweight and still feel good.

Speaker A: Ecstasy.

Speaker B: I'm not naming the drug now.

Speaker A: This was fun. Yeah.

Speaker B: Just hang out.

Speaker A: Thank you so much for listening. And if you found this episode useful, please share it with a friend or a colleague who you think would get value from it. And if you have topics that you would like to hear discussed on the podcast, either add me on LinkedIn or send me an email. Joeloderncto IO Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • #291 Why Most AI Projects Fail to Deliver ROI, Sinohe Terrero, CFO and COO, EnvoyGrowCFO Show · on Claude91 / 100
  • 512. Is SpaceX Over or Undervalued, Why Consensus Kills, How Chewy Beat Amazon, and the GameStop Saga from a Board Member (Larry Cheng)The Full Ratchet (TFR) · on Anthropic86 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on Claude86 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100
  • How to Sell Against a Competitor Already in the BuildingSales Leadership with Fexingo · on HubSpot85 / 100

More from Modern CTO

All episodes →
  • Why do 81% of Mature AI Programs Get Rolled Back? with Daniel Morris, CPO at Sinch
  • How Snapchat is Evolving in the Age of AI with Saral Jain, SVP, Head of Engineering
  • Why Kubernetes Still Feels So Hard & What to Do About It with Steve Francis, CEO of Sidero Labs
  • Why Does AI Momentum Stall in Large Organizations? with Kyle Lagunas & Allyn Bailey
  • Are Apps Going to Get Axed by Conversational AI? with Jacques Klick, Director of Market Strategy at Sinch
Explore the best B2B Engineering & DevTools podcasts →
All Modern CTO episodes →