
Facilitation Lab Podcast · 2026-06-29 · 49 min
Key moments - from our scoring
Substance score
60 / 100
Five dimensions, 20 points each
This episode tackles the real friction behind AI adoption in engineering organizations - not the technology itself, but the human side of change. Douglas Ferguson and Peter Bell examine three distinct patterns of pushback: engineers dismissing AI because they tried older models, those resisting the shift from hands-on coding to managing agents, and those with ideological concerns about AI. Rather than treating these as obstacles, Bell frames them as opportunities: skeptics can become quality gatekeepers reviewing AI-generated code, while craft-oriented engineers can encode principles into agent systems. The conversation explores how this shift mirrors historical technology transitions (punch cards to higher-level languages, manual assembly to compilers) and introduces new role archetypes - harness engineers focused on building frameworks for agentic systems, and members of technical staff (MTS) working across traditional silos. Drawing parallels to DevOps and developer experience, they discuss how platform teams must build context management systems, orchestrators, and quality gates. Bell emphasizes the importance of experimentation with clear boundaries: identifying 3-5% of early adopters per core team who can pilot agentic workflows without company-wide disruption, while maintaining different quality standards for internal tools versus customer-facing systems.
They should ask those engineers to try again with current models like Claude or Codex 5.x, as the capability jump from mid-November onward is dramatic enough that past impressions are no longer accurate.
Shift their focus from writing code to encoding taste - have them review all AI-generated code to ensure quality, define design principles (domain-driven design, thoughtful naming), and train agents to follow those rubrics, effectively scaling their expertise.
Harness engineers who build context management systems and quality gates for agents, and members of technical staff (MTS) who solve customer problems using any combination of coding, design, and product work without traditional front-end/back-end boundaries.
Security vulnerabilities in third-party open source require patching within 5-15 minutes to prevent exploitation; humans cannot patch CVEs that fast 24/7, so autonomous agents scanning and fixing become operationally mandatory.
Identify 3-5% of early adopters from each core team (not just safe projects), give them unblocked access to experiment, and ensure different quality standards for internal versus customer-facing systems based on blast radius.
Our reviewer’s read on each dimension, with quotes from the episode.
There are genuinely useful practitioner frameworks here - the three-bucket pushback taxonomy, 'rich get richer' in DevOps, token usage as diagnostic vs. success metric, the dark factory CVE argument - but they're surrounded by a lot of mutual agreement, throat-clearing, and conversational padding that dilutes the rate of new ideas per minute.
there is a absolute divide between two types of engineer and leaders. There are the people who have taken a day to go build anything in Claude code or something. Similar since then. And there are people who haven't.
if you can't create psychological safety, nobody's going to tell you you how their job works because otherwise you might replace them with a machine. And nobody's going to tell you that they're scared about being replaced by a machine.
A few fresh framings emerge - 'get AI infected' as an adoption mechanism, the adversarial 'cranky old Sam' review pattern, and the mandate-trap framing - but the conversation leans heavily on well-worn analogies (technology adoption lifecycle, punch cards, self-driving cars, TDD) and standard change-management vocabulary rather than genuinely contrarian or first-principles arguments.
artisanal software development. I would be surprised if that's a real profession for most people in the field a small number of years from now.
just as you had to get test infected, I think there's huge value in getting AI infected where a coworker just sits down with you pairs on a couple of features
Peter Bell is a credible hands-on practitioner - O'Reilly author, active builder of agentic harnesses, advisor to CTOs - but he operates more as a consultant and observer referencing others' transformations (Angie Jones at Block, Sam Scalacci at Microsoft) than as someone who has personally led one at scale inside a large engineering org.
I'm writing a book for O'Reilly called Scaling AI Adoption in Engineering
I've got my harness to the point where I spend at least 20% of my time saying what would be three ways that you could economically and token efficiently improve the thing you've just shipped
The episode is above average on specificity: named practitioners with their roles and companies, concrete adoption figures, real token cost numbers, and named internal tools (Amplify, cranky old Sam) give the conversation genuine evidential weight, even if data isn't always sourced rigorously.
we looked for 3 to 5% apiece. That was our number, 3 to 5% of the engineers. And they were spending evenings and weekends. They'd installed Gastown on their personal computers
Sam Scalacci. I first came across him when he was SVP engineering at Box. Now he's a deputy CTO at Microsoft. He has helped them to build this tool called Amplify
The host organises topics well, contributes his own concrete anecdotes, and makes useful connective pivots (DevOps → harness builders, mandate trap → metrics), but rarely challenges an assertion or presses for precision; most exchanges end in mutual validation rather than productive friction.
I'd love to Dive into those three a little bit more deeply and maybe let's start with the I want to be a coder category
Yeah, and it's really interesting too when you think about cohort analysis. You could look at that in a number of ways
Computed from the transcript - who did the talking, and the words that came up most.
In this episode of the New Friction podcast, host Douglas Ferguson speaks with Peter Bell, founder of Gather.dev and author of the forthcoming O'Reilly book Scaling AI Adoption in Engineering. Bell draws on his work running invite-only peer communities for senior engineering leaders to diagnose why most organizations stall out in AI pilot mode rather than achieving meaningful transformation. The conversation maps three distinct patterns of engineer resistance - skeptics burned by early models, craft-focused developers who resist the shift toward managing agents, and those with principled objections to AI - and offers concrete tactics for reaching each group. Bell and Ferguson explore how AI amplifies existing organizational health: strong DevOps practices compound upward while process debt scales its dysfunction. They examine the mandate trap, measurement via token usage as a diagnostic rather than a performance metric, and the non-negotiable role of psychological safety in any serious adoption effort.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Welcome to New Friction. I'm Douglas Ferguson. AI just made execution almost free. So why are organizations still stuck? Because the friction didn't disappear. It moved and it multiplied. It's no longer in building, it's in deciding what to build, how to align, and how to move forward when the path isn't clear. That friction, the human side of change, is what this series is about. Each episode, I sit down with leaders who are living it, navigating the real challenges of AI, uh, transformation. Not the tools, the people. The task that took two weeks now takes two minutes. The work isn't the bottleneck anymore. The conversation before the work is. That's the work this show is about. I'd like to introduce you to my conversation partner today, Peter Bell.
Speaker B: Welcome to the show, Peter Douglas, thank you so much for having me. I'm excited to be here and to kind of have this conversation.
Speaker A: Absolutely. And, you know, we've been in conversation a lot over the years, and it's been a lot of fun recently diving back into this new era of hosting conversations around what matters for leaders.
Speaker B: Absolutely. Well, I've really been kind of getting into how agentic AI is transforming the sdlc. Right. The software development lifecycle. I'm writing a book for O'Reilly called Scaling AI Adoption in Engineering, which is like, the hard part's like, I got 300 humans. How do I get them to work and be different? And then I've also been doing a lot of builder work myself. I've built my own harness and dreaming systems and orchestrators and tools like that. And it's been really fascinating just to see the impact of this, both on a personal level and also at the larger organizations where I'm discussing with the CTOs how this is impacting their teams and their priorities.
Speaker A: Yeah, I've been, uh, along on that ride and building some of my own stuff, and we've been comparing notes, and personally, I've found it illuminating to dream about the future a bit more. Being able to get hands on and making some things that are useful for me, and then I can more easily see how this is impacting organizations.
Speaker B: One of the really interesting things has been fundamentally, the models changed in mid November last year. You got like 4.5. You got, I'm forgetting now, the version of Codex, uh, was it 5.2, 5.3? And what's really interesting is I've found that there is a absolute divide between two types of engineer and leaders. There are the people who have taken a day to go build anything in Claude code or something. Similar since then. And there are people who haven't. The people who haven't are still like, man, you know, I've got like a VP of AI. They'll take care of this stuff, or my platform team will figure it out. And the ones who have done are like, wait, we need to call on all hands. We are not going to be writing code in a year and we need to ready the organization for our new future. So I feel like if you're not doing this, this is one of the few transitions as a leader where you can't just manage and lead through it because your intuitions are going to be wrong.
Speaker A: Yeah, that's an important point. It's funny, not only have I seen it from leaders, but I've also seen leaders really struggling, struggling to bring some of their engineers or just any talent along because the talent, the individual has made an impression of what AI is based on earlier versions of the model. You kind of touched on that a little bit. But I think it's really important to really, uh, sit there for a moment because it's an important thing to wrestle when you've got folks that have made a determination about what a thing is, but it's evolved tenfold over.
Speaker B: Well, I feel like there are three broad patterns of pushback and we're seeing this in everything from juniors to some of the most senior engineers, product leaders across the board. One is, I tried it last summer and it sucks. Um, I can do faster, so they just need to try it again. That's manageable. Another is, I understand that this works, but this is now not the job I want. I love being the person who goes on stack overflow and figures out just where the semicolon should be. I love the crossword puzzle, nature of design, elegant API. I don't want to be babysitting a bunch of agents. That's hard because the job has changed. And artisanal software development. I would be surprised if that's a real profession for most people in the field a small number of years from now. Another third is, AI is bad. It's going to world apocalypse, Destroy the population, destroy the planet, climate impact. And there are lots of very valid reasons to have an issue with AI at the same time it's here. I don't expect it to be going away anytime soon. And so the question is, as a leader and with responsibility to shareholders, you need to help the people within your team to build the outcomes you need. And not everyone's going to self select into that population.
Speaker A: Yeah, I'd love to Dive into those three a little bit more deeply and maybe let's start with the I want to be a coder category. That's just uh, my quick little name there for that. But I've seen two versions of this and so I'd be curious if you've seen other layers or variants as well. But this notion of people being in love with the craft and being passionate about the craft, and the interesting thing about them is that they, you can actually use that, ah, as a value or as something to lean into to help them explore and experiment with using AI. Because you can say, well, how can we evolve the craft of what does the craft look like when we bring in these new tools? Right? And if the craft is really about quality, like what, what are the principles, um, baked into the craft rather than the practices? You know, this comes back to the agile days, like when say they were doing agile because they were doing story points or uh, user stories, but they weren't really practicing the values or living up to the principles. And so it's like helping people distill down their craft into principles that then they can apply in the AI world. I think that's really fascinating.
Speaker B: Seen a couple approaches that uh, you've absolutely nailed it. One approach is the, oh, you don't believe AI code is any good? Great. What I want you to do is review all of it and tell us how bad it sucks. And that effectively, firstly it means you don't ship crappy AI generated slop, which is wonderful. And secondly, it means that you get a training corpus that means that the AI stops generating slop in general if you implement it well and manage your context and deterministic quality gates and adversarial, um, multi step reviews. I mean you need to engineer this correctly. But what it allows you to do is it allows that person to say AI code sucks and for the business to get better by them being precise about how and where it sucks, which then improves the quality of generation. And then I think the other point, the one that you doubled down on, which I love, is this idea of saying we're effectively moving up another layer. We can now take infinite pains with these things. You really think that we need to have thoughtfully, orthogonally designed names? We need to think about like Eric Evans domain driven designs and ubiquitous languages and rich, uh, domain boundaries. We can do all that stuff, but rather than doing it, what we're going to do is teach an agent to care about it and then ensure that we have a review step to make sure that the quality of our Architecture is better than we'd probably have time to ship if we're doing it manually. So yeah, it's absolutely true that we get to encode taste, which is actually an amazing way to improve the quality not just of the software you write, but of all the software that is then written by the factory that you're helping to train.
Speaker A: Yeah, I mean there's kind of no end to the depth that you can go into if you really explore what these new capabilities unlock from a perspective of craft. Because ah, you look at test driven development, right, that's something that everyone kind of aspired to, but how many people actually did it? This is a thing that, where you can instantaneously have tests if you have well defined boundaries and well understood contracts and that can be, you can get them for free now, you don't have to spend the time and then uh, and I, I think it really could put us in a position where we're, we're being more thoughtful, more mind about our design processes, some things that we might have spent more time designing if we had the time in the past. And also kind of pulling on that thread more, uh, something I've seen be really effective for organizations is where if folks are really pushing back and saying like I can write better code or I don't trust this thing, have it start doing the code reviews ideally in an autonomous fashion. Like sure, still have human code reviews, but have the humans look at what the AI is discovering as far as the bugs and issues and concerns. And when you've got senior engineers where the AI is discovering bugs in their code, you know, it starts chipping on the ego a little bit and they start to realize, wow, this thing is actually pointing out things I should have noticed and they start to give it a bit more credit.
Speaker B: Absolutely. And I think it goes both ways. I think that you get value from the humans reviewing the AI code because it improves as long as you are capturing the context, capturing the transcripts and feeding them into a world designed context management system. Uh, and then absolutely, I think we're going. It's bizarre, but it feels a little bit like self driving cars in a couple of ways. In the one way it's like it's really hard to get 100% of the way there. Anyone who's like dude, like I just switched on Claude code and now we're generating production ready code is either doesn't know what production ready code is or hasn't looked at the code they're generating. You can absolutely one shot or few shot stuff but it's a real engineering effort, at least today, to ensure that it is of the, the quality and maintainability you'd want. So it's hard to get all the way. But the other thing is, eventually, I don't think eventually people are going to think 20 years from now the idea like wait, Granddad, they used to let humans drive. Like, I mean, how did that work? Didn't they get drunk and look at their cell phones and kill people? Isn't it much safer to use Waymos? And I feel like we're going to have exactly the same with writing code by hand, which is. So wait a minute. In a mythos level environment where you can get CVEs coming in and they maybe need to be patch in 15 minutes before autonomous agents are basically scanning, seeing what your tool stack is and exploiting a, uh, close to zero day. Humans can't patch CVEs in under 5 minutes, 24 hours a day, but agents can.
Speaker A: Yep, yep. It's pretty soon going to be zero hours.
Speaker B: Right? I mean you really need to get to that point where uh, to the extent that you're using third party open source tools, which I do still believe bring value, the downside is there's lots more value in exploiting them. Uh, the upside is there are lots more people trying to ensure that they're not exploitable. So there's the trade offs. But to the extent you're using that, you need to have a dark factory in place by, I'm going to say sometime next year for most of your key production code. And if you're not working on that today, you're going to be in trouble. In Q4 when somebody asks why your competitors are going five times as fast and you're like, uh, we're investing. Give us nine months, we'll catch up to their velocity.
Speaker A: Yeah. You know your story about the self driving cars and it makes me think about how we used to write programs on punch cards. And what was that transition like if you were really into how to provide instructions to a computer via punch cards? You know, your future wasn't very bright. Right. In retrospect it seems kind of absurd to think, oh, I'm going to hang on to this punch card thing because that's the way it's done and, and, and that's my identity as someone who uses these things. But you know, uh, I think at the time there might have been folks that surely felt like this is not
Speaker B: really programming well and it's because technologies don't start off perfectly. It's just when we moved from assembler to higher level languages, there were absolutely professional software developments who were like, huh, these automated compilers are fine, but then they're going to work well on a 16 kilobyte memory system, on an embedded system, or I want to double check that they're using the registers effectively so that they're adding and storing in the appropriate register. I think I can do better than that. Uh, and for a period of time they were right. Uh, today I'd be pretty hard pressed to think of anybody who's literally hand coding hexadecimal to add two registers together to get business outcomes.
Speaker A: Yeah. And with the compiler optimizations, uh, you'd be hard pressed to find someone that was better at it.
Speaker B: And that's the point, right? It went from, uh, it sucked for almost everyone to it was good enough for some people, to it was good enough for most people. But there were edge cases where it didn't work to there was no good reason for a human to be doing that. And I think that we're going to see that in terms of encoding in C or Python or Rust or whatever it is you use to program. Uh, the only difference is I think the timeframe is probably going to be compressed given how quick all of these sigmoid curves are stacking on top of each other. Because so many people are working so hard to improve everything from the underlying silicon all the way up to the harnesses and the context that we wrap the models with.
Speaker A: Yeah, and also it's a reinforcing loop. So the advances in AI, uh, create advances in the other areas which creates more advancement, you know, advancement, advances, advancement.
Speaker B: And the crazy part is like, we could stop now. I mean, if we were just to say like, okay, 4.8 is good enough, 5.5 is good enough. There's so much overhang just with the current state of the art models, we could be kept busy for the next decade. And I hear that they're not stopping, so my guess is it's going to get even better.
Speaker A: Uh, that's my stance as well. Now the other piece, uh, because I said there were at least two things that came to mind when I thought about subdividing this. I want to be a coder perspective. And this one is about pushback on the need to be a manager in this world of working with agents. And so some folks opt to go down the management track, others opt to go down the principal engineer track or something similar, depending on what the terminology is at your organization. And um, this agentic workforce is pushing everyone into this kind of management model. You can't just do it yourself. You have to be able to delegate, you have to be able to review work from others. And um, I think you and I were not that long ago talking about how it somewhat feels like having an army of interns that you're managing. Right.
Speaker B: It's really interesting because I think you're absolutely right and the perfect first level analogy for this is hiring. And I think it's why a bunch of honestly old ex technical people like me are having the best time in our lives. Because like, wait a minute, minute now we can actually ship exactly what we want and bring our ah, understanding of systems design and engineering rigor and like, you know, good practices but also shape. Merge that with the fact we spent 30 years asking humans to engage and build things for us and we kind of get. It's not quite the same as managing humans. Right. I don't actually. You're not going to need a sick day because your mom died. Like that doesn't happen to Opus. Uh, but on the other hand a lot of the delegation patterns and a lot of the patterns about how can I use generalized language patterns? The models are good enough. Now that I've got my harness to the point where I spend at least 20% of my time saying what would be three ways that you could economically and token efficiently improve the thing you've just shipped. I don't even. I will review the answers but more often than not I'll be like, yeah, go with number two. And so I'm not even proposing what they should do. I'm simply doing what I would do with a very smart. And by this time they start to feel like a junior to senior engineer, not an intern, which is tell me what would make this better for the definition of better that I give you and what would be a token efficient way of shipping that this week?
Speaker A: You know, I think to that point acknowledging the fact that, you know, these delegation skills, these managed tracking, prioritization, all of these skills are going to be super critical in the future and making sure that we spend time upskilling and investing and ensuring that our individual contributors are ready to start taking on those duties. And also that's something to even look for when we're hiring new folks. Like do they have those skills? Do they have an innate ability to do some of these things? Are they kind of wired that way to begin with?
Speaker B: I think you're right. This is fundamentally going to change not only the interview process, which I think was already broken in terms of managing the flood of inbound resumes and the validation of competence and fit process. I think that's going to change, but it's also going to change fundamentally. Fundamentally what we're looking for. And I think we're moving towards this world where in R and D you're broadly going to have platform and feature as you do now. But the way it's going to look is you're going to have what's effectively harness engineers who are primarily thinking about how can I capture more, you know what, how could I create a step where effectively uh, an instance of something that looks very much like Kent Beck meets Martin Fowler looks through our code and says whether it's good or bad. Right. How can I extract those insights, capture them in a context efficient way and help agents to generate a rubric for it and then manage the pipeline for managing that. So that's going to be people who are effectively either on the platform team or are embedded in stream aligned or feature teams and then I think the feature engineers are going to look much more like mts. Like we're already seeing this member of technical staff where we don't have front end, back end product engineering design so much as people who are uh, deeply understanding the customer problems and trying to frame experiments and features design to deliver value to those customers using whatever combination of product and coding, front end and back end is required.
Speaker A: M. Yeah, the piece you're talking about there around um, the harness builders or maintainers made me think a bit about DevOps. Right. Because there was a certain um, group of organizations that treated DevOps more as developer experience, uh, especially if you're thinking about developer experience of internal developers. And then there are some folks that thought of it more like site reliability or you know, the cloud version of sysadmins. But the organizations that we're thinking more around, how are we making it more um, streamlined and more enjoyable to do work as a developer, as an engineer? Uh, I think you think about that role, that definition of DevOps and it very much is this kind of universe of how are we building the harness, how does it make for a great developer experience and provide all the tools and functionality we need to excel and create basically agentic teammates.
Speaker B: Exactly. And I think we're going to see that there's two components to this and the companies that already have some kind of DevOps platform org will be in the best place. And as you called out, if they happen to have already renamed a subset of that DevX or developer experience, they're killing it because they're thinking about the right things, which is how can we help stream align teams, feature teams who are shipping the stuff our customers want? How can we make them stars? How can we make them succeed better? And I think that what you're going to see is, uh, that the platform team's going to own more most of the harness design and management. But I could also imagine, at least in an intermediate period for the next couple years, 18 to 36 months and maybe ongoing, you're probably also going to have embedded harness engineers or devex engineers within each team because what you're going to see as well turns out that a, uh, good harness, a good set of steps in a pipeline for throwaway react code for an internal admin dashboard is probably different from the level of quality and the type and definition of quality you have for the ST you use to build your customers every day. And so you're going to find that different teams working on different projects will require different subsets of. You're still going to have the same basics, orchestrator context management, but the details of the rubrics and the steps and the validations are going to be very different across your org depending upon what people are building and how much it matters.
Speaker A: Yeah, Security and uptime guarantees are totally different when you're looking at internal tools versus external as well.
Speaker B: Yeah. Or even similar to speaking with Rob Zuber, the CTO at Circuseon, he's like, look, you know, it would suck if our admin dashboard went down for an hour. Like, we don't want to do that. But if we can't run continuous integration runs for an hour, we're going to be getting phone calls. Right. They're two different things. And so you have to look at, uh, the blast radius of the changes you're making.
Speaker A: Yeah, that's an interesting point around even as we're experimenting with AI and what ways we might leverage it because, you know, there's some obvious use cases around. Oh, there's, we can have it review, pull requests, we can have it like sit here and help generate code. But there's tons of nascent opportunities we haven't pinned down or, you know, identified. And that's going to require a lot of experimentation. But there's a lot of folks in organizations that are afraid to experiment because they don't know what the consequences are. You know, they haven't been given the latitude, hasn't been spelled out. And I think being very clear where the no fly zones are and where there's ripe opportunity for Experimentation gives people a bit more confidence when they do fly and experiment. So they can avoid those no fly zones, but you know, lean in in certain areas.
Speaker B: A lot of this is the context. Right. You know, when many, uh, a couple of separate things I'd say. The first thing I'd say is don't expect 100% adoption. That's not a realistic goal. I was speaking with Angie Jones, who did this transformation off the last year at Block, uh, Jack's company before she moved on to the Agentic AI Foundation. And she was like, we looked for 3 to 5% apiece. That was our number, 3 to 5% of the engineers. And they were spending evenings and weekends. They'd installed Gastown on their personal computers. They were like doing all the things and we made them, we unblocked them and we elevated them. Although the one interesting thing she also said is, she said, you know what, we also made sure that we picked them from all of our core teams across the company. So that rather than just saying, huh, it worked for that admin dashboard or a little bit of app modernization, but it wouldn't work for hard engineering problems. They weren't given that chance. The good news with that is it meant that they created an org wide transformation very, very quickly. But to give you an idea of the level of executive support that required, that basically meant Angie had somebody from like legal and security in her team seconded to her. And it was kind of along the lines of if they couldn't either approve or reliably say, no, no, we can't do that because, you know, model hosted in China, probably not a good idea from a security perspective. If they didn't have a clear they could show or couldn't approve a tool in a small number of days, they would just pull the lever and it's like, okay, we're going to stop everything. Do we need to get Jack in the room? And because of that they had such strong executive support, they could get things done. Uh, I've been talking with other companies where they're still, you know, the CISO has just approved Copilot recently for 10% of the engineers. And I'm like, they're going to get exactly the outcomes you'd expect from that level of support.
Speaker A: Yeah, you know, the varying levels of support is an important issue. You just pointed out there's also even lack of understanding around what governance is. And when you've got different folks in the org thinking different things and expecting different things and there's a lack of alignment there and then there's not great governance being provided. You kind of get this perfect storm of like everyone being kind of frozen and afraid to do anything because they don't quite understand. Uh, so not only does it take good solid governance, but great, great uh, communication as well to make sure that people understand what that means and how to apply it.
Speaker B: Yeah, and it's really interesting because I'm, I'm all in, I'm all the way AI pilled, you don't have to be. And like, I mean even in the book I'm talking about this idea of pick a lane. It's like, you know, we have a technology adoption life cycle and there's going to be no different for this than anything else. And there are valid reasons to be an innovator, an early adopter, uh, early or late majority, maybe a laggard. For example, let's say you're in the business of, of ski resorts. You literally, you run a bunch of ski resorts. Your biggest business risk is an AI. It's climate change. Like that's what you need to deal with. It's, it's whether or not you're going to get enough snow. And honestly if you're six months or a year late to the party and you're like we're just going to wait till Microsoft folds it all into 365, uh, you probably could have made a little more money and shipped a few more features earlier but who cares if however you, you're Shopify or you know, like Wix, like a website builder, you probably need to be an innovator or early adopter otherwise you're probably not going to be here in five years. And so it's important firstly to pick a lane that is consistent with the business risk and opportunity you have and then secondly to make all of your communications consistent. Otherwise you get into this. The, the worst anti pattern is where the CEO is on CNBC to say and tell everyone how AI killed you are whilst the CISO is still saying nothing but copart.
Speaker A: Yeah, and that kind of gets into this issue we hear time and time again across clients and folks at these executive dinners we've been hosting. Is this issue around trust. And it cuts both ways because you've got folks that don't and we talked earlier about people that don't trust the AI. And then also you've got trust in the organization and that's partially because of the phenomenon you're Talking about where CEOs going on, you know, C span or whatever, saying some things and then ah, might not be in agreement with ciso, whoever else. But also things are changing so rapidly. The company's message is gonna morph a bit because, hey, there's, there's new things. We understand now and I think that, um, a lot of individuals are getting, are really frustrated by that and it's not necessarily a company's fault, but if we don't pay attention to that and shape that narrative and make sure that we're consistent in it and make sure that people understand why it might be evolving, maybe acknowledge, yes, we understand and we said this last month, but now we know this and so we have to take that into account. I think it goes a long way to be transparent around the thought process. Not just like, what does everyone need to know?
Speaker B: I think that's it. And messaging is always hard in large organization and change management. And this is a huge change management issue. Plus the fear is, wait a second, am I just encoding my taste so this thing can replace me? Are you going to need the same number of engineers and product managers? There's very valid reasons to be scared. And at the end of the day the transition is happening. And the thing that's most. What's really interesting to me is and we saw this in the DORA report like last year, the rich get richer in every dimension. And what I mean by this, if you've already got good DevOps practices, turns out if you're doing agentic coding, but Sally still has to FTP the files to the server, you're only going to get so much acceleration, right? There are continuous integration, continuous delivery, test coverage, feature flagging, so you can decouple, deploy from release and run experiments in production, whether using Datadog or Honeycomb, like the telemetry that you've got, the observability data so you know what's going on in production, all of those are more critical than ever. And the reason I thought of this is the other thing that's more critical than ever. Blameless postmortems. Uh, an environment where everyone, if you can't create psychological safety, nobody's going to tell you you how their job works because otherwise you might replace them with a machine. And nobody's going to tell you that they're scared about being replaced by a machine. And they're just going to tell you that co pilot doesn't work very well and that's not going to work for anyone.
Speaker A: Yeah, I love that point. And I want to come back to your comment. The rich get richer. And I just want to be really succinct there because you can be rich with process or you can be poor with process and those rich with process will get richer. Right. It will amplify that wealth of, uh, process that you have. But if you're poor and you haven't invested in process, you've got some dysfunction. It's going to amplify that dysfunction. So not only does Sally FTPing the file over prevent you from really leveraging the agentic workforce to help out in those areas, it's probably indicative of some other like, you know, process debt that you have that's maybe going to get scaled in its own. Right. And what about the edges of those moves? Right. None of that can be integrated. And so I think that's the thing we've been encouraging people to think about and that's what we really mean by new friction, right, Is Sally moving FTP file is going to present itself as serious friction in our ability to become the next level organization.
Speaker B: I feel that the other thing also is it's investing in your team not only in terms of don't expect 60, 80% of your team to jump straight on board. You find the 3 to 5%, the Coalition of the willing, the people who are going to spend their evenings and weekends doing this, not because you tell them to, to, not even because you want them to, but just because what would be more fun? And there's a certain point in your life as a builder where playing with this stuff is just fun. And that's great. Um, but then you need to then build the tools and the trainings and the systems to help at least the 60% in the middle to make that move across. And you've got to give people time to win. If you're like, hey, we need you to ship everything, which is still critical, oh, but also take 15 hours a week to go learn this new thing that's not going to work one way or another. It's going to break with anybody who has kids or family or commitment or parents to deal with. And so you really have to give people the time to adopt. And the good news is you actually don't usually lose velocity, but you need to give them the permission to lose velocity for a quarter so that you can speed up in the next quarter.
Speaker A: Yeah, it comes back to that psychological safety piece you mentioned earlier. You have to make it safe to experiment in a number of ways. Right. It can't be a cider desk project. They have to have reserved and protected time. And then also they have to be treated with, I would say, respect and encouragement when things go wrong. To your point, if you miss a deadline because you're experimenting with AI, well we need to step back and look and say, what did we actually learn stuff? Does that mean we're going to beat the next deadline by 50%? Okay, well it all comes out in the wash. But if instead we just have an immediate reaction that's bad, we need to be punitive here, then we're really going to miss the vote. People are going to stop experimenting.
Speaker B: And I feel like I remember Etsy back in the day and it's different, but I think it's comparable. They used to have this commit on day one policy and they still do and I think it's much broader now. But this was maybe 10, 15 years ago and most people were like, wait, you're letting somebody who knows nothing about your systems like commit some kind of. Even if it's just like a nominal fix to like a button on day one, what if they break things? And the feedback, the answer from Etsy was if your system is so fragile and brittle that somebody with good intentions can break it on their first day at work, you should be building your systems, not putting more gates in place. And I think that's the way to think about all of the agentic engineering as well, which is we need to build both the culture of psychological safety and support, but also these deterministic and adversarial gates, these tools to make sure that if somebody does make a mistake, you catch it early and quickly and it's unlikely to go to production and B to waste two weeks of their time trying to figure out why these prompts don't work.
Speaker A: You know, I joined a startup years ago and my first day on the job as cto, the junior engineer who had just gotten promoted before I came on online, they promoted him from, he had just wrapped up his degree at UT and so they converted him from intern to a full time engineer. And it was within my first week, week and he managed to delete the production database and um, luckily there were processes in place, we got it recovered, et cetera, et cetera. And I was posting, I can't remember, it might have been hacker news or it was somewhere that I posted just like, oh my gosh, first week on the job. Ah. And then of course someone commented like don't let them near production systems anymore. And my comment was, was I have more confidence in that individual on the production systems now than I do some of the other folks because that experience, like watching them go through it and them doing what they could to uh, a The fact that they reported it, they didn't try to hide it. The fact that they were just terrified and you know they're going to be walking around on pins, on needles anytime they're on a production machine. And I think that's the lesson we should learn. It's like not how do we punish someone but how do we actually learn
Speaker B: from our learning machine mistakes, I guess. Last anecdote for that. So I remember um, one of the, I think it was the first CTO of the, of the United States, uh, who there was a guy who helped to turn around healthcare.gov um, I believe it was back in the time thinking Obama days maybe. And what was interesting is he gave this talk at uh, at a group I was involved with and he said he couldn't. So imagine you brought in, this thing is months behind schedule, it's not working, it's a piece of junk and you need to fix it using the same people with different budget. You can't change out the team. How do you turn it around? And the first thing he did, he actually kind of seeded it where one of the people stood up and said, I lost some data from production. And everyone's like, you know, contractors like Washington D.C. i mean this is like, okay, you're never going to get another federal contract again. And he said, great, let's take a moment, let's have a round of applause for that person for being honest and sharing. Great. What did we learn and how can we build processes so that doesn't happen again? And that will was, he said, the turning point where they could start to build the psychological safety. So people were focused on sharing the problems they had so they could fix them rather than hiding them and hoping to run out the clock.
Speaker A: Yeah. So important, especially in this era of AI where we're moving so quickly and adopting new things, creating those environments is so critical.
Speaker B: Absolutely.
Speaker A: So I want to switch gears a little bit. You kind of touched on this a bit when you mentioned that in reality the adoption is about 3 to 5% and yet we see a lot of organizations with these top down mandates and we actually refer to it as the mandate trap. Uh, it's one of the things we're noticing right now. We're trying to coach any of our clients away from any of those behaviors. But I'm curious what you've noticed and specifically when you think about this adoption rate of 3.5%, maybe how to get it up. How do we measure success in this AI world I guess is kind of what getting a, uh, because that's how we get past the mandates, is being able to measure the process, I think.
Speaker B: Absolutely. So firstly, I should uh, clarify. I think you're going to see 3 to 5% of super adopters and then you're going to see, I mean the, the know Steve Yakey got into a lot of trouble on Google, uh, by talk, on Twitter, by, by X, by saying, you know, 20% of people are killing it, 60% of people will come along and 20% will never touch it. And the same is true at Google and anywhere else. And first level Rand numbers, he's about right. There's a few people who are killing it, a bunch of people who are willing to follow along and a small tail who just have no interest in going. So first thing to do is drop the. Not like fire, but like don't focus on the last 20%. What you do is you elevate and you unblock that first, uh, 3, 5, 15%, whatever it is, you make sure that they get the tokens they need, the support they need, the resources they need and you elevate. Then you ask them, great, now you're doing this. How can we do this as an org, join a council, do lunch and learns? Can we build a small devex or platform team that has shared skills and shared resources? How can we get uh, more observability and capabilities within our platform? Uh, so a lot of this is about unblocking and supporting. Another part then is creating a path for the middle, the kind of quiet middle who just want to go home in the evenings but aren't opposed to AI. You just need to tell them how to do it. And then the other piece of this is in addition to that, you need to support these groups in figuring out what problems they have. So you're talking about management and metrics. The success metrics are honestly business ones. And you can take proxy metrics as long as you don't, you know, performance manage them. You learn something by token usage. If somebody's not blown through a $20 a month plan, they're probably not using it enough. But if somebody spent 8,000 bucks last month and has spent 6,000 this month and their output increased, they've probably improved the efficiency of the levels of the models they're using. So it's not that token maxing is good. But it is okay to know how many tokens people are using as a diagnostic to put them into populations, which you can with adoption in different ways. The true success metrics are pretty straightforward. It's revenues, it's Customer, uh, retention, it's all the numbers you care about. Uh, the challenge becomes that it's hard to map those to a particular feature, deliverable or a particular agent. So I think the main thing to do is the AI, uh, token usage and stuff like that. That is a diagnostic to help you to cluster people around common failure patterns of adoption so that you can give them the training and support to, to learn how to get through. Oh, that's the kind of thing we see when somebody's still in an ide. We should teach them how to use skills with agents. That's the thing we see when somebody's waiting for one agent 15 minutes at a time. We should show them how to use multiple agents and so on towards moving them towards using a dark factory. Um, and then the other piece is classic Dora DX Core 4 space metrics. I think still have a place. They're not the answer, but things, things like cycle time, mean time between failures, uh, PR rates, all of those can be gamed, but if there's no reason to game them, they can be useful diagnostics to help you to see how you appear to be doing.
Speaker A: Yeah, and it's really interesting too when you think about cohort analysis. You could look at that in a number of ways, Right. You could look at any of our standard metrics like cycle time, a few others you mentioned as it relates to folks that are know, heavily using AI, not using, barely using it to not using it at all. Because then there's an interesting story to be told there, right. It's like how are we, what kind of outputs are we seeing from these individuals and different teams as well? The other thing around cohorts is fascinating to me is what are we noticing as far as the friction that we might be seeing from each of those cohorts? And because you mentioned, you know, looking at, at, well, what signals are indicative of someone, you know, still being in the IDE or what, whatever, some of these, uh, types of behavior shifts that we're looking for. And I, I think likewise if we diagnose where are the sticking points in the organization and what are those indicative of? You know, it's like, hey, you know, if, if we're going to be investing in the, in this 20% that's really leaning in and we're trying to remove obstacles. Well, let's actually make note of the obstacles that are running into. Right. And then how do we codify that into repeatable patterns or, or, or better ways of supporting them?
Speaker B: Absolutely. And ah, I think we're seeing that, that what's nice is I think that 5 to 15, 20% what they do is they're actually as lot the obstacles they usually run into are uh, self induced by their company. It's and I have lots of friends with CISOs but like security compliance, government audit, governance audit risk. It's those groups that are designed to, to keep things the same so we don't break it all, which is a noble mission. But we need to understand that there's risk to not changing and support those teams in being enablers and not blockers. And then once you've got that in place then what they can do is they can. The truth is what they're doing is hard and it probably is requiring evenings and weekends. But what they can do then is synthesize the good practices, create standard skills libraries, create a standard uh, factory harness and standard adversarial reviews, improve the quality of the observability and The DevOps and CI and CD pipelines, all the things that are going to make it easier for other people on the teams to then kind of join along. And the other thing it feels to me like I remember when you mentioned TDD earlier, Test driven development. You had to get uh, most of us got test infected. We'd read the books, we kind of saw the stuff, it didn't really make sense. And then you paired with somebody from like a Pivotal Labs or ThoughtWorks back in the day and you have like oh that. And after two or three hours of pairing it made perfect sense. And just as you had to get test infected, I think there's huge value in getting AI infected where a coworker just sits down with you pairs on a couple of features and shows you how they're leveraging skills, how they're jumping between agents and how they're building these kind of pipelines so that they can start to trust the quality of code that's being shipped.
Speaker A: Yeah. And that's the kind of, I mean sure we see the CISO friction all the time. Right. Like oh, we really want cloud code, security is only approved co pilot or whatever. Right. And so that certainly is an uh, obstacles we should as leaders be trying to remove. If, if we've got folks on the team eager to, to push things forward in ways that uh, are responsible and secure, then we should pave the way there. But I think the, the latter half, the stuff you got into I think is a little less not obvious. Right. Like it's totally clear if like they're trying to use a tool that's not Available. But what about these, like, models and patterns? Because you can't go just grab a book off the shelf, you can't go read about Spotify's model of how to do this. Right. And so are we creating opportunities to sit with peers and see how they're each using skills and, and, and really look at like, what, what is some of that minor friction that they're running into that's not maybe, um, apparent or, or where they're scratching their head a little bit. A great example. Example, a buddy of mine mentioned that someone, uh, on his team was, um, he's a VP of engineering and someone on his team was, had one shot, this piece of code that was like, I don't know, it was like 250,000 lines of code or something insane. And there. And admittedly the engineer came to him and said, you know, this seems to be working. But I'm like, I can't even fathom how to read this much code. I'm stuck. And so then that became a conversation around, well, this is some new friction. Like, look at this thing. It's like, it's brilliant. It seems to work when we poke it, it does the things we want it to. But like, we need to understand this better. And the thing they came to was, what if we then use the AI to decompose this into more meaningful, smaller chunks that are easier to read. We're still get this out the door way faster than we ever would have previously. But let's just be a little, let's induce some slowness here because we want process and we want care and quality. And I think, think that's a great example of the types of friction we should be listening out for and helping our teams work through because that's what's going to create the models of the future.
Speaker B: Exactly that. Uh, so there's a guy called Sam Scalacci. I first came across him when he was SVP engineering at Box. Now he's a deputy CTO at Microsoft. He has helped them to build this tool called Amplify, which, it's like the harness, the best harness nobody seems to know anything about. They don't promote it very much. Uh, but what's interesting is he's got this, uh, Sunday letters from Sam on Substack and he's got a couple things that he's built into his pipeline. One was cranky old engineer, which is basically the salty old engineer that'll never work under load. You've got to ensure that there's a fallback and you have, um, you Know like back offs for your retries against the API or whatever it is. Uh, but now he's got cos, ah, cranky old Sam, which is based on cranky old simplicity as well, which is basically saying hm, hmm. And it will literally go back to the agent that's generated a code that's passing the functional test, that's meeting the performance requirements, saying would there be a simpler way to do this? And proposing unifications, um, and simplifications and simpler ways of solving the same problem. And it turns out that a lot of this is just re prompting and loops. And if you're willing to burn the tokens, most of the problems the agents solve, you can get other agents to tell them how to fix effects.
Speaker A: Yeah, it's amazing. It's so fascinating. I mean and to your point, willing to burn the tokens, I think that's going to be a conversation that uh, evolve even more so over the next six to 12 months. You know, as we're seeing the, the cost of tokens rise, more competition against models, the IPOs are certainly going to influence us because now the market's going to be a driver and have a voice in the cost of these tokens. So it's going to be a fascinating look and see how we start to optimize around token consumption and where and why to use them.
Speaker B: And I just want to throw one thing in there because what happens is every so often people who want AI not to succeed and I get it, I understand why will be like, oh, token costs are going to become crazy so we're just going to hire humans to do it. If there was one piece of generalized advice I could give is don't bet against a mobile models, I don't think that's a good long term bet to take. And there's no question that sanity is going to prevail. We just, Tokenomics is a real thing now just as FinOps is for Cloud, we've started by, let's say we're just going to run all this Kubernetes stuff in the cloud and it's going to be perfect. And then the CFO comes calling like why did our operating cost go up by $12 million last year? Oh, uh, we forgot to switch off the, we had this test run and we forgot to switch it off for six months. Right. That was like a million and a half. And so then you start, started to bring sanity and improve operation and then you know, spot versus reserved instances and all the rest. We're going to do the same here. But Here it is. You should never use a model to do something that code can do perfectly well. Long running. Don't use supervisor model, use deterministic pipelines. If you're extracting text from a PDF, have a python script, extract it and just put the text into the model. Don't burn the tokens that are 4.8. And it's all about there's I just ran a bunch of evals. I had a bunch of stuff running on Opus that I've now downgraded to Sonnet and then to Haiku with evals and test sets and they just auto tuned the prompts until it would work. So all of this is just a simple engineering problem. We know how to engineer the costs out of stuff. So yes, token costs are real. And no, don't believe that's somehow magically going to stop this. It won't.
Speaker A: So what you're making me think of is that, ah, Chaos Monkey might be coming back, but in the age of AI,
Speaker B: I think it's going to be so many of the things we've seen before are coming back just at another level and it's going to be really interesting to see how they all play out for sure.
Speaker A: Well, as we come to our end here, want to give you an opportunity to leave our listeners with a final thought.
Speaker B: Absolutely. I'll give a two for one. Uh, the first thing is do it yourself. You shouldn't if you're a cto, you shouldn't be writing production code and blocking that for three months as you go into performance review season. But if you're not spending time building with these models, you won't get the right intuitions and that's the only way to keep up. And second, have a sense as to where we're going. This isn't about Copilot, this isn't about ide, this isn't honestly about the kind of interfaces we're seeing now. We are building systems that will write the software and our job is to identify the experiments and the verifications and build the toolings to make that work. And understand that until you fundamentally rebuild the SDLC for agents, you're not going to get the kinds of ROI that you'd expect expect out of these systems.
Speaker A: Important words. Great to be chatting with you today, Peter, and looking forward to talking again soon.
Speaker B: Douglas, thank you so much for the invite. So much fun.
Speaker A: Thanks for listening to New Friction. If you enjoyed this episode, share it with a leader who's in the middle of this right now. They'll thank you for it and if you want to go deeper. We bring leaders together through executive dinners and virtual meetings. Masterminds. To learn more about our work or to inquire about exclusive executive events, visit folterscontrol. Com. I'm, um, Douglas Ferguson. See you next time.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.