DevOps Paradox · 2026-09-16 · 54 min
Key moments - from our scoring
Substance score
62 / 100
Five dimensions, 20 points each
The AI productivity paradox reveals a critical disconnect in software delivery: individuals feel 20% faster while measurably shipping slower, deployments and tickets increase, but overall team throughput and value delivery don't improve proportionally. Jeff Kaise argues that AI acts as an amplifier - magnifying both good practices and bad ones. When legacy pipelines already plagued by technical debt process more AI-generated code, failure rates spike and incident management consumes resources meant for innovation. The real bottleneck has shifted from engineering velocity to product definition, organizational topology, and flow distribution. Kaise emphasizes Value Stream Management principles: measuring not just code output but value realization across the entire lifecycle. He highlights that product managers, architects, and CTOs must now understand flow distribution - the balance of work allocated to innovation, technical debt, risk, and defects. The conversation explores how AI is collapsing traditional role boundaries; developers become product designers, product managers write code mockups, and testers extend into all disciplines. However, this transformation demands new skills (evals, data orchestration, code comprehension) and threatens those unwilling to evolve. Legacy systems and decades of accumulated technical debt remain anchors slowing innovation, even as coding becomes cheaper. The real work now lies not in development speed but in discovering requirements, modernizing systems, and orchestrating teams around customer value rather than metrics like lines of code or deployment frequency.
AI amplifies existing structural problems in organizations - when legacy pipelines with technical debt process more code faster, failure rates increase and incident management consumes resources, offsetting individual productivity gains. The real bottleneck has moved from engineering speed to product definition and organizational design.
Flow distribution is the mix of work a team allocates across innovation, technical debt reduction, risk management, and defect resolution. AI changes this balance because teams often use AI only for new features, ignoring debt and bugs, which eventually destabilizes the system.
Roles that only synthesize backlog items or write code to spec are at risk; the new world requires product managers to run evals, understand code, orchestrate data, and maintain product taste - while developers must understand business and participate in definition. Those unwilling to expand their role face obsolescence.
The challenge isn't cost but discovery and prioritization - companies must reverse-engineer legacy systems to understand original requirements, decide what to modernize vs. replace, and balance debt reduction with innovation using flow distribution metrics.
No; AI should be trusted proportionally to pipeline stability and failure rates. Even conservative estimates suggest AI can safely handle one-third of work, but this requires monitoring software delivery performance and the right balance of feature work versus debt and risk reduction.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode surfaces a genuinely valuable observation: that individual developer productivity gains from AI may not translate to team-level throughput or value delivery, supported by DORA metrics showing increased instability. However, this core insight is repeated across multiple speakers rather than densely packed with novel claims. Much of the middle section (team topology, role blending, flow distribution metrics) covers well-established Lean/DevOps frameworks already familiar to the target audience. The token-spend analysis is moderately novel but takes 15+ minutes to develop.
Are we truly more efficient? Do we have more throughput? Or as the DORA report talked about, creating more instability also? Yes, there's problems that are happening. We're not really seeing the throughput, we're not really seeing the increase in quality.
The majority of the spend then turned into what should we build? What's it look like? Uh, how do we define it? What's the right thing?
The framing of AI as an 'amplifier of good and bad' is not particularly novel (widely discussed in AI ethics). The observation that specification/definition is now the bottleneck is sensible but not counterintuitive for operators who've read recent product management literature. The Value Stream Management framework and flow distribution concepts are directly borrowed from Lean Manufacturing - not new to DevOps practitioners. The episode lacks contrarian takes or first-principles challenges to mainstream AI adoption thinking.
AI is an amplifier and it will amplify the good and the bad.
The new world of what needs to happen next in this transformation is not an engineering problem. It's a product problem. It's a definition problem.
Jeff Kayes is a credible practitioner with real executive experience (worked at Planview, Plutora, helped found Value Stream Management Consortium, now at Allstacks as an executive). He speaks from direct observation of enterprise AI adoption and customer token spend patterns. However, he is also currently pitching a product (Allstacks), which introduces commercial bias. He is not a founder/operator at a major successful AI-driven company, and his authority rests on advisory/consulting vantage rather than hands-on scaling experience.
Part of what I do in my job is I talk to a lot of different companies and customers and walking them through, or not walking them through, but I'm, um, part of their journey as they rediscover their new processes.
We were looking at their token spin and saw something interesting.
The episode includes some concrete examples: the Tesla parking incident, the text editor accidentally built by agents, ProductCon attendance, mentions of Allstacks customers' token spend patterns, and references to specific tools (Langfuse, Linear, GitHub Copilot). However, there are very few named customer examples, no specific revenue/business metrics, and most quantitative claims ('90% of developers using AI') lack sources or attribution. The token-spend insight is the most specific data point but remains somewhat vague about the actual figures and companies involved.
One study found that developers who felt 20% faster were actually measurably slower.
I had a friend that bought a Tesla, thought it was super cool, he told it to park itself and it promptly parked itself and crashed in the garage.
The hosts (Darren and Victor) ask reasonable follow-up questions and push back on some claims - Victor challenges the assumption that humans inherently know better than AI, and Darren questions whether reading full specs is realistic. However, the conversation rarely becomes adversarial or forces the guest into a corner. Most pushback is gentle. The hosts could have pressed harder on the Allstacks commercial pitch (which takes up significant airtime), the actual evidence for the 90% figure, or specifics about which enterprise customers experienced the token-spend shift. The conversation is collegial rather than probing.
There is one thing I strongly disagree with what you said. Really strongly. You said humans know it. I worked with humans who built stuff that nobody wants because it was in a requirement.
Does even matter anymore other than token spend. Seriously does.
Computed from the transcript - who did the talking, and the words that came up most.
#368: You're faster with AI. Your teammates are faster with AI. More tickets closed, more deployments going out than a year ago. So why can't anybody in the building answer a simple question: are we shipping more value? Jeff's take is that the bottleneck didn't go away, it moved. Code got cheap. Deciding what to build, why, and whether it actually worked did not. That's not an engineering problem anymore. It's a product problem. In this episode, we speak with Jeff Keyes, Field CTO at Allstacks, about where the bottleneck went once AI made coding cheap, and why product people have to pick up the tools next. Jeff's contact information: LinkedIn: YouTube channel: Review the podcast on Apple Podcasts: Slack:
Transcribed and scored by The B2B Podcast Index.
Speaker A: This episode is sponsored by Allstacks.
Speaker B: Are there more deployments going through? Yes. Are there more tickets being closed? Yes. Are there more confidence in the ability to take an idea and vibe it through to come up with something? Absolutely. Are we truly more efficient? Do we have more throughput? Or, as the DORA report talked about, creating more instability? Also, yes, there's problems that are happening. We're not really seeing the throughput, we're not really seeing the increase in quality. In fact, while individuals report better individual productivity, as a team, we can't really answer the question of are we shipping more value?
Speaker C: This is DevOps Paradox, episode number 368,
Speaker A: the AI Productivity Paradox. Welcome to DevOps Paradox. This is a podcast about random stuff in which we, Darren and Victor, pretend we know what we're talking about. Most of the time. We mask our ignorance by putting the word DevOps everywhere we can and mix it with random buzzwords like kubernetes, serverless, cicd, team productivity, islands of happiness, and other fancy expressions that make us sound like we know what we're doing. Occasionally, we invite guests who do know something, but we do not do that often. So since they might make us look incompetent, the truth is out there and there is no way we are going to find it. Yes, it's Darren reading this text and feeling embarrassed that Victor made me do it. Here are your hosts, Darren Pope and Victor Farsek.
Speaker C: Victor, you use AI for daily development. I use AI for daily development. A lot of developers are using it. 90% are using using it. Most of them are saying it's helping them be productive. Is it making you more productive, Victor?
Speaker D: Let's put it this way. If you tell me no more agents for you, I'm growing, uh, tomatoes. I'm not going back to software industry. That's executive decision from my part.
Speaker C: One study found that developers who felt 20% faster were actually measurably slower. So here's the question for today. What if a whole team is confidently shipping nothing faster? On today's show, we have Jeff Kise from Allstacks. Jeff, how are you doing?
Speaker B: I'm doing fantastic.
Speaker C: Did I say your last name right?
Speaker B: No, Kai's. We have to make it messy. If. If you speak more Latin languages, people say que because it looks like Reyes. And in English, most people say keys, but it's Kai's.
Speaker C: And I even have it phonetically in front of me and I did it with a dash key, so it's Kai's. Did our story sound even familiar there of what's Happening. In fact, let me ask it a little bit differently. We're not debating whether or not AI is helping type code. Right. We can see it typing code really fast. What we're really trying to figure out, is it helping us ship any faster? What do you think, Jeff?
Speaker B: For sure. Are there more deployments going through? Yes. Are there more tickets being closed? Yes. Are there more confidence in the ability to take an idea and vibe it through to come up with something? Absolutely. Are we truly more efficient? Do we have more throughput? Or as the DORA report talked about, creating more instability also? Yes, there's problems that are happening. We're not really seeing the throughput, we're not really seeing the increase in quality. In fact, while individuals report better individual productivity as a team, we can't really answer the question of are we shipping more value? There are some reasons why that I believe are part of that story. But is AI helping developer productivity 100%.
Speaker D: Everybody says that on individual level AI is helping, whatever that means. The logical conclusion must be that then there is something structurally problematic on a company level, uh, process level. Right. Because if I'm more productive and you're more productive and Darren is more productive, but when we combine our productivities, that does not turn up to be more productive as a whole, then my only explanation can be, okay, your process does not work for this.
Speaker B: Yes. I like the explanation that AI is an amplifier and it will amplify the good and the bad. And I think the challenge of the bad is so much that it's offsetting some of the good. If I remember right, looking back at some of AI's impact, according to the research that was done last year, individual effectiveness like went way up, but software delivery instability also went up. Amazing thing that happens to pipelines that were mostly right. When they're mostly right and you put more through them, you got more failures and more issues to go solve, which means higher incident rates, higher deploy failures, and the rest of it. The most notable example from my point of view was a recent session we ran to exemplify a new feature.
Speaker C: We.
Speaker B: We ran a proof of concept and part of that proof of concept included this text editor as part of it. And not thinking it through, we just built it all and the agents went to work and did great things, including building a text editor. It's happy to do whatever we feed it and whatever it sees and wherever it goes. Is that exemplifying what problems we see in the industry? I think so. I think there's a lot of. Do platform components get Used well, are existing APIs being duplicated? Are bad patterns being replicated? Because that's what we did in the past, maybe. How much are we learning from agentic development from what we've already done? Answering that question, I think is at the heart of where AI is amplifying some of those negatives.
Speaker D: Apart from amplifying negatives, which I. I mean amplifying everything, both negatives and positives. I feel that the problem might be that people driving AI and agents might not feel comfortable in the new role. Right? Because you have a typical developer in a stereotype enterprise that, hey, somebody made decisions, somebody specified exactly what you should be doing, and so on and so forth, and you just need to do as you're told. And now we're in a situation where that same person is supposed to be telling others, and by others, I mean agents what to do. And that person never did that. Right? Kind of like, I know how to follow the rules. I don't know how to come up with rules. Not my job.
Speaker B: I got to tell you, I. I see this a lot. Part of what I do in my job is I talk to a lot of different companies and customers and walking them through, or not walking them through, but I'm, um, part of their journey as they rediscover their new processes. And what are these new roles? Developers now see themselves in part as designers because you can, you can just jump out to tools and go do it. Developers see themselves in part as product people because you can. Product people see themselves in part as developers. So the roles are very much blending the people that were doing definition all along, starting all the way from top down of we're going to have this initiative and it's going to drive this kind of value all the way into. Did it work? Is being reimagined. The people that thought they knew so much now need to work differently. AI is doing a whole disruption of how the team operates. The size of the team, the roles, the handoffs, the measurement of value through. And that's where the conversation of the impact of AI to the software industry is super fun. We sit at the precipice of what will be an entirely new model, as big or bigger than Agile was back in the day, or even waterfall. The whole methodology of how we build from beginning to end is different. And now that the genie is out of the bottle, it never will get put back in. This is a new world we live in.
Speaker C: I'm thinking about what you were just saying there. You're talking about value and flow. I Have a feeling you come from a bit of a Value Stream background because that just rolls off your tongue too easily.
Speaker B: Yeah, it does. Uh, one of the things that frustrated me a long time ago was looking at people that talked about wanting to go faster. So we would sit in these cycles of analyzing software delivery and they would really optimize how fast can we get that deployment? How fast can we get from even an idea into developing something? When there were six to 18 months on the front end of that, while they're figuring out which kind of initiative should we build, they would get a, uh, feature request and it would sit for a year waiting to decide, is that the right one? Instead of just building it Been in other environments where even after it was coded, still now has to get into some really long release cycle where it takes another quarter or six months to actually get in the hands of customers. Meanwhile, where the focus was, let's speed up development because we have all these metrics about where they're at and what they're doing. Each individual. How many keystrokes are they pressing? How many lines of code? Please, I hope we never end up back in that world again. Because no one cares. How many lines of code. Let's look at value. Part of that discussion led me to part of the journey I've been on in my career. I worked at planview. I was at Plutora. I helped start the Value Stream Management Consortium. We managed delivery looking at the entirety of the life cycle of software from idea all the way into delivery itself. Borrowing from Lean Manufacturing, the basic concepts of flow. How long does it take to get something done? I m don't just mean coded. I mean from the time you thought about it into production and actually getting value. What's your typical time to get a work item completed? What is the load that has on the team? Meaning how many different kinds of work items are there at a particular time? All that led into what became the Value Stream Management Consortium that has been reformed into a new organization called Flotopia. It is out there still in the industry and highly suggest get involved.
Speaker C: So here's a question for that. When the AI coding tools showed up, did it change the flow conversation at all or did these tools just make the old bottlenecks that much more obvious?
Speaker B: Uh, yes to both actually. Did it change 100%? It changed. It changed that you could take an idea if you knew really what you were going to build and get it through to production so much faster than you've ever been before. And look at the consumption, the value realization side of that equation faster than ever before. Does it happen that way? No. Some of the old bottlenecks are still there. Back to your second point of did the bottleneck move? Yes, bigger. Uh, part of conversation, what I'm really passionate about today is about where that bottleneck moved to. The new world of what needs to happen next in this transformation is not an engineering problem. It's a product problem. It's a definition problem. It's an organization and team topology problem. It's a how do we act and behave differently so that we can keep up with the tooling that enables us to build things on a whim and decide if we built the right thing or not? What does that look like and how do we do it in a way that is sustainable at an enterprise?
Speaker C: Well, even an enterprise. Does that matter? Even if it's a small business, it should be the same thing. I should be able to get something shipped faster now. Or I have an idea this morning I should be able to be testing out by lunchtime to figure out if I want to ship it at dinner time.
Speaker B: You're 100% right. I throw the word in at an enterprise. I mean at scale. Easier for a small company that's not laden and technical debt and decades of code to try to navigate and correlate across teams. Harder for these teams that have to do this at scale, meaning the number of people, the number of technologies, the number of historical components. Components and things that all have to work together, that have a certain amount of governance and controls that must also be in place. That's where this gets super interesting.
Speaker D: Uh, I'm, um, glad that you mentioned technical depth, because one thing I don't fully understand is until recently, people were saying, hey, we know that we have stuff that we don't want. We know that we have mainframe. We shouldn't have mainframe. We know that we have this and that, whatever, we shouldn't. And now that it's extremely cheap and we are not fixing all those things because it's too expensive. We don't have time, we don't have bandwidth, we don't have money, we don't have the things needed to fix it. Now that it's extremely cheap to do those things, why do we still have them? Why do we have technical depth? As in not technical depth generated by AI, Just to be clear, technical depth generated by our past decisions or decisions that could have been good 20 years ago, but they're, um, not good anymore.
Speaker B: Which every decision we make today is probably the same boat of it's. Going to be a uh, new world in 20 years. I've posed that same question to a head of product at a healthcare company and talking about their AI transformation, His response was really interesting. Talked about that their AI journey. They decided explicitly, while they had pockets that were dev focusing on how to implement the technology, they just waved their hand and decided that the development side was going to get solved and that uh, the cost of coding would go and be much less. Where they saw the deficiency, including in these legacy platforms, is their ability to discover and manage the original requirements of these systems so that they could rework them and bring them forward. The idea in reworking them, they were left with this choice of uh, do they go out and try to discover all the original requirements by sniffing through the code, looking through behavioral patterns, looking through the data, or do they bring it forward and decide what things do they keep or roll together or so forth? They decided on the latter and in the journey they learned a lot in deciding the code was cheap. That meant what their real work was to discover how do these systems operate. It wasn't just that the people were left the company like in some cases, like the people were no longer on the planet and they needed to uh, discover things that they weren't going to discover any other way. So they had to reverse engineer all this to then bring it forward. The point is every company at scale faces this same problem. What do you do with all this legacy stuff and how do you modernize it and how do you make progress? Because that debt, if you will, is a, ah, boat anchor that slows down any future innovation that you do. You have to service that debt, modernize. It's got to be a certain amount. Back again to the value stream side. One of the key metrics looking at value stream management is what we call flow distribution. The idea is that there is a certain amount of work that you do in any given cycle that will be a uh, breakdown between innovation, debt, risk and oh gosh, I'm trying to remember this. It'll be a certain amount that is debt, risk, innovation and defects. And the combination of all these things have to be evaluated in combination to decide where you're really going to place your bets. Whatever you're using to generate the code is great, but you need to take it into account of the combination of all that so that you're meeting your goals. Why is this relevant? If you use AI and you generate only features, will there be bugs? Of course there will be bugs. If you use AI and you only address defects Will those defects continue? Yes, because you're not reducing your debt at some point. As you look at the overall stability of systems, you can find an ideal profile of how you should be banking that investment, who should be making that investment product. So flow distribution sits at a pinnacle point, A, uh, critical measuring stick for how you look at everything. You go build. A certain amount of work you should do should be to address security and vulnerabilities. A certain amount you should do addresses risk and incorporate that with innovation and defects.
Speaker D: I agree with everything you just said, except that one of the main reasons so we have debt, we have issues, we have bugs, we have all those things, uh, without a doubt, and we will never be free of them fully. But the main reason why we have them in quantities, we have them is because somebody said, okay, we are thousand people and we need 10,000 people if we want to put this to a manageable level. And then, uh, somebody else answered and said, you're not getting 10x increase in headcount. Forget about it, deal with it. And then you need to make those choices.
Speaker B: Right?
Speaker D: But now that I can effectively increase equivalent of a headcount very cheaply, I consider AI actually to be very cheap compared to human labor. Why would somebody have more than a manageable number of open issues?
Speaker B: Why not just fix them all?
Speaker D: I'm not going to say all intentionally, but put them to a m manageable level. Let's put it that way, right?
Speaker B: It's such a great question, right? Have we discovered all of our defects? Do we know what the impact is of every time we fix a defect? Is there a large defect because it's technical debt? Is it a design technical debt or architectural technical debt? Is it a platform technical debt? There's so many different ways that this comes out that has to be addressed back in the core Dora, uh, metrics themselves, which I love. What they did, the foundational piece is that what they discovered was basically with increased AI usage, there's an increase in platform instability for all the reasons we just talked about. I had a friend that bought a Tesla, thought it was super cool, he told it to park itself and it promptly parked itself and crashed in the garage. Dented up the bumper.
Speaker C: It was awesome.
Speaker B: Just feel so bad for the guy. How much do you really trust AI? 100% to say, oh yeah, 100% of the time, it's going to be just fine fixing this bug and deploying out to production. If your failure rate is such that it is low enough that you can trust it, great. But most pipelines are not there. So even just fixing a bug and changing anything still means you've got to run through that pipeline and see how well that all works.
Speaker D: I don't trust humans 100% just to be completely clear and transparent. So I don't trust people, I don't trust agents, I don't trust anybody fully. Nobody gets a blank check but me. Not trusting agents 100% does not mean that it cannot do 20%, 80%, whatever the percent is, and I'm still winning. I'm going to be extremely conservative here and say, hey, one third of my work is now done by agents. This is extremely conservative. Right?
Speaker B: Yeah.
Speaker D: Okay, so that's one third for me to dedicate on depth. Uh, unless, and I suspect that's where companies are going, is that they maybe cannot stop themselves from constantly working exclusively on new features. Oh, you can do much more now. Let's do more fe. More new features.
Speaker B: Yeah. And they should, and they should monitor the performance of software delivery in that process to figure out if they have the right mix of debt, innovation, defects and risk. And if they have the right mix, they're making the right decisions. Because you're right, nothing's perfect. Just watch the metrics and find the areas you got to improve. I, uh, want to reiterate here. I think the right people that need to understand this more now than ever before are product people.
Speaker D: Here's then a question because I'm going in that direction. Maybe I would put a wider net and say product people, architects, CTOs, stuff like that. Are they now the right people to do the stuff? My, my question is should a, uh, traditional developer extend its role into product or should product extend its role into development?
Speaker B: And should the two extend their roles into design? And should design extend their roles into product and development? And should test extend their roles into all. Yes, is the answer to all of the above. Uh, I think the team topology of today and tomorrow is one that is much smaller. Less dependencies, more self directed, more watching for outcomes and watching the whole system. Less handoffs, more integrated work than ever before. And the reason why is because AI, AI makes any developer a part product guy and vice versa. Every product guy is now part developer. Every uh, designer is part uh, and right on through. And that's the wonderful part about it. What it means is that the handoffs that we used to have aren't going to be sufficient. The time it takes to understand each other is reduced. And we all need to take into account the distribution of the kinds of work and the why between behind what we're doing, everyone needs to know, uh, and understand our flow performance, our flow distribution, our software delivery performance in conjunction with the kinds of features and innovation with the related value that we expect to achieve. With that in mind, developers understand the business and can be part definition of what the solution will be to that problem. Product people can be part design and part architect because they have the tools to talk to the code. And same with design. Uh, it's a wonderful world. Like I said, the whole world has changed and it's fantastic.
Speaker D: I completely agree about the wonderful world for myself. I'm just worried how that there is a percentage of people, whatever the percentage is, who are simply don't want to be in that world. I speak all the time with developers and products and stuff like that. I just want to write code. That's the only thing I want to do. Tell me what to write, I'm going to write it. Or, you know, product people. I'm never touching the code. Never. That's not my job.
Speaker B: Yep.
Speaker D: Are, uh, they becoming dinosaurs?
Speaker B: Yeah. Let's talk about the world of product management for a second, just for instance, because I think most people that listen to this show have heard a lot of development. But even on the product side, there's groups of product people whose entire jobs were to take customer feedback and synthesize it and create a backlog. There were heads of product that would create grandiose roadmaps and do that to achieve alignment working up and down the management side, there were people that listened to customer incidences and tickets and would correlate that data. There were business analysts that were part product people that really worked to understand the system and were part project managers. If you really resonate that I'm one of these roles, you should really listen because I think those roles are going to either go away or at a minimum, change. The new world of product management requires that product be much more close to the customer and be closer to their correlated development partner and design. The role of product must include some key skills. You got to know how to vibe code. What you're doing doesn't mean writing production ready code. It means create mockups that you can show to customers and get feedback and actual signal on. Um, if what you're going to build is worthwhile, you've got to be able to run evals on the kind of AI components that you're building so that you can decide if the quality of output's good. You have to understand how to orchestrate data and do that analysis all yourself. And meanwhile Having an umbrella hat of a product sense that you have a taste for what this product should be like. That's the new world of AI skills that AI PMs have. So now's the time to be able to retool yourself into that world or decide if you want to go play in a different area because everything else is going to get automated.
Speaker C: You were talking about quality of outputs, what about the quality of inputs? You talked about it earlier. It's critical now. Actually. It's always been critical. The problem has always been that we have humans interpreting what those specs and other items, those reams of paper, of PRDs from Word documents that were badly formatted?
Speaker B: Yes, those, those that somebody spent an inordinate amount of time creating. And then it was this great shiny document that then got put into refinement where a team broke it down into a bunch of epics and user stories. And then an architect came in and so you really sucked at that. You forgot about these 10 different things and the number of stories and tasks tripled. And then design got involved and said, well, this isn't the right thing. That reminds me of that picture. You remember the picture where the customer wanted to swing and then it goes through all these different phases and then at the very end here's what the customer actually got because it kind of all blew up. That is the problem that has to change now and is changing. I think the world of uh, product management, of just creating a static prd, that world is done. I am not saying that the PRD is dead, in fact, far from it. I think the requirement for product is to create a product definition. That is the golden thread that lives through all these different source systems that acts as a system that, that humans are have an easy time inferring what the intent was. AI doesn't know. And I, AI will happily build whatever's on paper, even if it's ambiguous. Just like the example of building a text editor, it'll just build a new one instead of being. No, you dummy, there's source, there's platform components, there's operating system components there, there's web component. Don't go rebuild stuff. Humans know that AI doesn't. In fact it confidently and happily goes and builds as much code as it thinks it needs. The point is the world of product management is the driver for where this transformation has to happen. Because it sits at the intersection of being able to take all these inputs, this context rich world, correlating them together and understanding software delivery and these different skills to build things of Value and then sit at the measurement to decide, did we build the right thing? That sounds a lot like a mini CEO. Great. That's the right thing.
Speaker D: There is one thing I strongly disagree with what you said. Really strongly. You said humans know it. I worked with humans who built stuff that nobody wants because it was in a requirement. Maybe we know different humans. I mean, some don't. I got the answer, but this makes no sense. Yeah, but it's in requirements.
Speaker B: You know, you're so spot on. One of the. Look, I work for Allstacks. We're building a product targeting product managers and its goal is to help answer exactly that question. A PRD includes the intent, not just the feature. What are you trying to do? The user stories contain the acceptance criteria, for example. Okay, when we're done, we want it to be able to have this set of functionality. And here's how we'll verify or prove it. The design includes the intent of the user interaction. What happens on that journey we just painted before of? Uh, a PRD gets created. It's a shiny doc. It goes stale the minute it hits whatever source system it lives on. Google Drive, SharePoint, whatever. And nobody goes back to update it. Nor does anybody go back to update all the decision making that was made since then. Especially when you think of how many times were you in a cycle where you created some kind of feature and then because of all the explosion of extra requirements and things and decisions that the feature morphs so dramatically it doesn't even really represent what the original requirements doc was. So by the time somebody gets to it, and then a human has to ask questions like, this is what we were trying to build.
Speaker A: I'm lost.
Speaker B: We're back to the swing problem. One of the reasons we're building the product we're building is to have an answer for exactly that question. Why, in this world of AI, do we have to be limited to just documents that are static? Why can't we ask a digital twin of the feature, maybe a digital twin of the product manager? Why are we doing the things we do? In fact, why can't that include evidences or receipts of the various decisions or prior work? Maybe, for example, we decided to not use a platform component and build a new one. Why would we do that? Oh, I see. Every time we tweak or have any changes to this particular platform component, it's a disaster. Or it's already at scale, it can't handle anymore thinking of a logging component. That one time we had, that was a disaster. It was like, please don't use it. And all those decisions could be things that people could come back to a living product definition and ask. Because we have all the ability to correlate a set of contexts which includes all this work and background that went into it. So, uh, anyways, I'm excited about what we're building over at Allstacks for exactly that reason. To have an answer to some of these things that help put together software delivery, performance and a living definition of the features themselves that get put into features as they go forward. So that anyone at any point in time can say, what did we build? Why? Maybe after the time when we got started, we can look back and say, where did we drift from what we originally defined? How? What did it look like? What are the upstream downstream impacts? How could we create some kind of knowledge base for our customers so that they know what changed? Are my stakeholders in this journey? How should I evaluate success differently now from what we originally planned? How will this impact the investment? All these things are great and they're all part of what we're talking about.
Speaker C: I'm going to double down on what Victor said about the humans not being what you thought. There was a phrase that you'd said earlier that the developers understand the business. I can count percentage wise on hand of how many developers may have understood the business a little bit.
Speaker B: Uh, yep. Okay, but I'll give.
Speaker C: There are some right there. There are those. If you read the Phoenix project, there was Brent. Brent knew everything. So you're going to have those Brents of the world. I'm going to extend it. I have been around enough vendors that product people are coming and going almost on a yearly basis. The chances of them understanding the business is also very low now. And what you just sort of laid out before with the product that you all are building at Allstacks now, we can have AI help us do that, but does it really need to help us if we properly define how the agents should react? Your digital twin again, do we need humans anymore? Initially, yes, but forecast it out. Not forecast it, but looking ahead now, we've trained the system in such a way that, okay, we have Jeff's knowledge, we have Darren's knowledge, we have Victor's knowledge, we have Mary's knowledge and all of that's there. Does it matter anymore? Do we need those four people or not? Or do we only need one of those four people?
Speaker B: Uh, that's a great, great question. And I think there are entire companies betting on that question. ThoughtWorks did a release of their new operating system for Software development. There are others linear for example comes to mind and how they're approaching this same space. I'll tell you what I see currently that when we looked at some of the token spend from some of our customers because one of the things that Allstacks has is uh, a software engineering intelligence where think of us like an observability solution across software delivery. We create all these metrics and can show you your dora, uh, metrics, your value stream flow metrics and any other kind of slice and dice dashboard that you want. But we were looking at their token spin and saw something interesting. Initially token spend was largely code, all sorts of things that was just pent up demand. Go build this stuff. What was interesting is we started to evaluate the spend after the initial push and after features that were backlogged and then after let's go fix up the pipelines and after let's bolster up some of our test harnesses and frameworks and everything else for what I would deem sort of delivery risk kinds of items for their pipelines, the majority of the spend then turned into what should we build? What's, what's it look like? Uh, how do we define it? What's the right thing? And starting to pull in data from the other side. Is it increasing daily active users, we're starting to speak this language like let's look at consumption. Did they use it? Did they click on it? Are they churning more? Are they getting value out of it? We start to have the conversation go much more cleanly to this clearly use consumption based kind of value. Did we produce the value because people are using it? And if so great. And what's interesting is we look back on the tokens is that's where people are starting to spend their money. Parallel to this is there's a, we were first kicking around the idea of building Product Studio which is another module of all stacks. We were looking at the various competitors to Allstacks. Okay, who's targeting product people right now? And there's a bunch of tools, there's some open source ones. Even GitHub has their spec kit. And then somebody spun off and did open spec and a uh, person out there did bman, a whole bunch of other tools. AWS is Kiro and there's all these different tools that are out there. What's funny is the ultimate goal of these tools was to help in the process of creating agentic requirements and specs as they go forward. The token spend that they're using on um, generating the technical steps for the code Generation is so small in comparison to all that was being spent on the requirements themselves. What do we build? Why I think that's a super interesting point when you look at the token, uh, spend is because in order to answer what should we build and what's going to get value the most, you've got to reach a lot further upstream. What's the goal of the business? Who do we target? What's our identity? You have to have a bit of product taste. You have to understand what has worked and what hasn't as you make bets. And you need to understand where the levers are that you'll look to see if you're successful and you're going to make a lot of bets. So it's a really fun process looking at all this stuff together and seeing where this bottleneck has shifted to and where the tokens are going to be spent. Especially, uh, you know, one other tangent on that it's I've. If you go to YouTube, you'll see no end of people building their own knowledge management system on their desktop to try to answer some of these questions where they're pulling in such a broad set of context of stuff. People pulling skills out everywhere to try to do some of these things. It's just like I said it, the world has changed and it's really cool to watch it.
Speaker C: I'm thinking about being able to take those outcomes and I'm thinking about, okay, great, let's say it's 80, 20, that it's design versus implementation. We're spending 80% of our time, 80% of our token spend. These are my numbers, not yours. But let's stick with it here. 80% of our token spend is for design and 20% is for implementation of that design. How many people actually read those full designs that they're having the conversations with AI about? I'm talking reading the full thing because if you've done a good spec, There's a near 0% chance you're going to read the whole thing.
Speaker B: Correct. In fact, you kind of don't wanna.
Speaker C: Victor and I talked about this because there is nothing new that anybody is creating that hasn't been created before. So all of the knowledge in some way, shape or form already exists in all of the models.
Speaker B: Yep.
Speaker C: It's just what is our as a human. How do we have that unique selling proposition there? That was my marketing phrase for the day. I mean, that's what we've got to try to search for if we're going to actually make any money off of this. Right. Now the only people I see making money are the model vendors themselves.
Speaker B: Yeah, code people, uh, people using AI to write code are making a bunch of money. There is a massive influx headed towards product people. I was just at ProductCon a few weeks back and it was standing room only. I think people are reimagining what this organization looks like and how they'll use AI in here to make that happen. And I think the whole tooling has changed, the surface of what tools people use has changed because now I have this consistent interface that I can work with. And how do we get all the data, how do we get all the context, how do I aggregate it together and get the right stuff?
Speaker C: How do we get our specs ready? In theory, used to you'd have a team of 25 people sitting around collaborating on a Word document. I didn't say Google Doc, I said a Word document that they were passing around to each other on a shared drive and producing that. Now we just have a conversation with model of the Day and we say, yeah, that looks good. Go.
Speaker B: Yeah, isn't that funny? I actually gave a talk about this as a thought exercise of with AI, why can't we use agents to get our requirements ready? What would be the features of such an agent? What kind of capabilities does it need to pull from the requirement to know that we're doing good? And what rating scale would we use to get there? The point of it is that it's within the realm of possibilities of things that you should think about building. Build one. It's awesome. And then of course I said, oh, by the way, I'm with Allstacks. We had this tool called Product Studio. We did that and if you want to try it out, come over here. But the thing that's interesting is why are we not thinking this way? There's nothing stopping anyone from starting to use agentic methods to evaluate requirements work or requirements themselves or specs or designs, or in fact it's expected we should and then include humans as part of the evaluation process. Langfuse describes this process of evals as having agents in the loop, humans in the loop, and LLM based analysis on evals. I think that's true for everything we do. The model's right. Got to keep people there as a sense of reality and as a double check to make sure that the self driving car isn't going to park itself into something dumb that you can clearly see, but it can't.
Speaker C: Thus the reason I will never own a Tesla, not for other reasons. It's like that's the one reason I just like no, I don't want to do something silly like that. You're talking about you have the platform to do all this for people that can't buy the platform today. What mhm. Are the habits they need to get into so they can get into this rhythm to where okay, now I can have Claude helping me on this or I'm using a mix of Claude and OpenAI to figure out judging each other of what's going on. Is that one of the things they can do? Or what do you say to some team that are saying no, we don't need any AI because we got this?
Speaker B: I would say to any and every team, you will not survive without AI it just isn't feasible. It's too expensive, you will be too slow. You're either going to have to disrupt or be disrupted, period. Second point is the cost of all stacks is such that you just give it a try. It's going to make you faster, better, stronger. All the six Million dollar man things. If you remember that reference from way back in the day, uh, there's no reason to not. And if you want to try it out yourself, just go to your own Claude code instance and tell it to write you up a little agent to validate your specs, validate your requirements, see how well it does and then you'll create an appetite where all of a sudden you're going to want some things. If you haven't read Andrew Carpathy's little blog about all the different things that he's doing, especially the create your own knowledge management system and put Claude on top of it to evaluate your conversations and to track the things you researched. It's a great exercise that will change the way you work. I started on that path, I don't know, some time ago. I take notes completely differently now. I operate differently. My personal to do management is completely changed. All uh, because AI side note on that, which is funny, I'm a yes guy. If I'm in a conversation, as you can tell, I, I like to talk and I say yes, I'll do that to a lot of things. And then I would promptly forget about a third of them and be like, oh, I need to remember to do that. Part of my system of AI changed that. So a lot of the meetings I'm in are recorded. And so I track those, I pull them down and I have AI every day pull out. Here's the to do's I'm supposed to do because this is what I committed to. Like oh yeah, And I go get it done. I've actually taken that the next step of where it's helping me pre create a lot of the tasks. If it's an email, it'll create the draft and so forth. So this is our new work style.
Speaker C: I'm thinking back to all stacks. Maybe not because I haven't looked at the products, but initially you're talking about, hey, we've got all the Dora things, we got all the engineering metrics, blah blah, blah, blah blah. All the things does even matter anymore other than token spend. Seriously does.
Speaker B: No. I love your question. Gartner seems to think so because they just did their developer productivity insights platform. Magic Quadrant, of which Allstacks was in, was a visionary.
Speaker C: That's great.
Speaker B: But yeah, I know.
Speaker C: Bless their hearts.
Speaker B: Bless their hearts, exactly. But I've been saying that for a long time. Does developer productivity really matter anymore? I mean we want to. How do we measure productivity?
Speaker C: Did it even matter to begin with?
Speaker B: Right to that point, uh, we take the really the shortest period of time in the overall cycle and we're going to measure them more than any other organization has ever been done more telemetry data on engineering than any other work and yet it never really was the biggest problem. It was a problem, but not the biggest problem.
Speaker C: What do you think the biggest problem was then?
Speaker B: I still back believe getting people aligned and the definition side and breaking this down so that teams operate more autonomously versus trying to achieve some big standard where everything operates at a fixed rate across a fixed kind of boundary of time, that we must have some kind of big room planning and manage your dependencies till the cows come home. I, I so hate all of that. You lose the plot. Unleash people. Unleash people. Create mini product team. I am um, content that the strongest transformation in any team isn't necessarily AI itself. It's creating a completely autonomous, self directed mini CEO over a feature who's going to whip the results that he was looking for through the team itself and making sure that the team is aligned and what they're building, they will have to use AI to get there and he will prove the results. The minute that a company turns into being more of a like a venture capitalist and operates in that model where money gets doled out based upon business performance of each individual feature and team. Great side tangent. I still remember like I used to work on the office team. We built a little product we had to go to quarterly reviews with. They were our Bill G reviews and either you were on track to be a hundred Million dollar business or find a different group to work in or different company. Go make it happen. I love that model. Internal teams would compete, build different things. Okay. See who's faster, see who's realizing value, see who's gelled. But put it on product to go build that out. Put it on product to go be the head of that feature and own their destiny. Because then they're impassioned. It's their baby.
Speaker C: Values, outcomes, all the things that we should have been measuring all along, but instead we got uh, busy measuring the uh, things that we thought we could measure to begin with. Like lines of code. Give me a break.
Speaker B: Yeah.
Speaker C: If people are listening today, what's one thing you'd want them to remember from today if they don't remember anything else?
Speaker B: The one thing to take away from this is there's not some big transition that you need to go to school for. Spend three years and you're going to come away being all new. This is a different style. Start taking the tools today, start working today, do one little change every single day. Sort of atomic habits esque. Be 1% better tomorrow. 1% closer to being self driven. Understanding the goal, breaking dependencies, using tooling, keeping up with what's changing and looking ultimately at the value that's being delivered in the process. Just a little bit of improvement every cycle, every week, something a little bit better. There's no reason to, not even for personal use. Experiment, play, Tinker.
Speaker C: 20 bucks a month and then once you get started on 20, you'll head towards 100. Ask me how I know. I haven't hit the 200 yet, but I could. Yep. Tell us more about Allstacks. What is, you know, you've talked about different features. What is Allstacks? As it stands today, In September of
Speaker B: 2026, AllStacks has three core modules. We got to start in the uh, engineering, intelligence, developer productivity space. So if you want a bunch of dashboards and a bunch of metrics about anything and everything, great, we got that. There's also an MCP server so you can pull those metrics into your own. Whatever you want to visualize. Kind of engine that's interesting and required as a context graph, as a set of data so that you can do other stuff. What's that other stuff? We enrich that context graph with capitalization and investment hours, meaning we give timesheet level fidelity data to every person that's working on the team. We know what they're doing, we know their activities, we see where they're writing code, writing docs writing designs and can track their usage by correlating source systems to the things that they check in and understand how much time spend the big thing that Allstacks is doing is we in late May released a product called Product Studio that's targeting product managers and Product Studio's goal is to help create that living definition of a set of requirements and the sharded out work tree and integrate with the prototyping tools and create a set of documents that are living with a chat interface. So you can ask why a feature is built the way it was built, why it has a particular design or use whatever components it did or what customers were looking for things and have that remain and accelerate so that it produces human and agentic ready code. So it's what we like to call build ready or we're shovel ready on once you get done with things and we have a bunch of other stuff that we do in the back end from agentic risk analysis, agentic requirements readiness, workflow analysis, looking at those bounce patterns and so forth. A lot of other things are put in. The speccing process includes things like adversarial reviewers and things that include various formats of the documents that get put out and different systems that we integrate to but ultimately the whole goal is to help accelerate the process and method of building these requirements by having an expansive context that's managed by one system versus individuals trying to build this themselves. Find that product leaders are just have so much on their plate now. If developers produce that much more, their pressures are just massive to hold all this context in their head and produce more and manage the results. It's a big job so we're helping with that.
Speaker C: So you can find allstacks@allstacks.com that's a L L S T-A C K-S.com all of Jeff's contact information will be down in the episode description. One question before I let you go. You said you worked on Microsoft Office.
Speaker B: Yep.
Speaker C: Please tell me you worked on Clippy.
Speaker B: No, I uh. Sadly I did not. It was a really fun time. We shipped a product in the Microsoft Outlook box so it was such a fun time.
Speaker C: Jeff, thanks for being on today.
Speaker B: Thank you much.
Speaker A: We hope this episode was helpful to you. If you want to discuss it or ask a question, please reach out to us. Our contact information and a link to the Slack workspace are@devopsparadox.com contact if you subscribe through Apple Podcast, be sure to leave us a review there that helps other people discover this podcast. Go sign up right now@ah, devopsparadox.com to receive an email whenever we drop the latest episode. Thank you for listening to DevOps Paradox.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.