DevOps Paradox · 2026-06-17 · 56 min
Key moments - from our scoring
Substance score
47 / 100
Five dimensions, 20 points each
DevOps Paradox episodes 355 explores a counterintuitive finding: AI-assisted pull requests paradoxically slow down code review cycles, creating bottlenecks despite automation promises. Hosts Darren Pope and Victor Farsek challenge popular statistics claiming AI code introduces 1.7x more issues and discuss why these metrics miss the point - what matters is whether iterative detection and correction reduces total defects per feature delivered. Victor argues that security vulnerabilities and code duplication flagged in AI-generated code aren't AI problems but SDLC execution problems: if you have scanners and processes, deploy them consistently throughout the pipeline instead of months later. The episode tackles the real review bottleneck: teams accelerate code generation without expanding review capacity, turning reviewers into the constraint. Victor reframes this through time allocation - if AI shrinks coding from 80% to 30% of developer time, freed capacity should absorb review work rather than creating a speed mismatch. Key takeaway: AI doesn't fix broken processes; it exposes them. Companies chasing cost reduction through AI (or any technology) miss the point - the goal is qualitative delivery gains, not headcount replacement.
AI-assisted PRs are typically much larger in scope because AI generates more code at once. The review pipeline was designed for human-speed changes, not machine-generated volume. This creates a bottleneck not because of AI itself, but because teams increased code generation without proportionally increasing review capacity.
Not necessarily. The statistic only matters if bugs aren't detected or fixed. Victor argues that if you have mechanisms to detect issues - scanners, code review, testing - and feed corrections back through the SDLC iteratively, what counts is total defects per feature delivered at the end, not raw issue count during development.
Code duplication isn't inherently an AI problem - it's a code review and static analysis problem. Tools like CodeRabbit, Anthropic's tools, and multi-layer SDLC processes should catch duplicated constants and redundant functions before merge. The solution is applying existing code quality gates consistently, not blaming AI.
If AI shrinks coding from 80% to 30% of developer time, the freed capacity should be redirected to higher-value tasks like code review, design decisions, and testing - roles where taste and judgment still matter. This resolves the review bottleneck without adding headcount.
Victor strongly argues that chasing cost reduction through AI (or outsourcing) fails long-term. Companies using AI successfully focus on qualitative gains - doing things better and faster - and cost reduction is a byproduct, not the primary driver. History shows cost-chasing strategies (offshoring, low-code hype) eventually backfire.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode has a handful of genuinely non-obvious reframes - particularly Victor's argument that the raw issue-introduction rate of AI is irrelevant as long as detection-and-iteration mechanisms exist, and the 80/10/10 time-budget analysis showing freed-up review capacity. However, these are buried under significant banter, movie tangents (Memento, Groundhog Day), and repetitive 'nothing new is happening, just faster' padding that dilutes density.
I couldn't care less whether it introduces 1x or 0.5x or 50x more issues. If I have a mechanism to detect those issues and then iterate and iterate and iterate. What I care is whether at the end of that process I have less or more issues and per number of features I'm developing
with assisted coding, let's say that in the past, you would spend 80% on coding, on average. Maybe not you, but, you know, on average in a company. 80% on coding, 10% of what comes before coding and 10% what comes after coding...Now we are shrinking those 80% to, let's say, 20. Imaginary number 30. Whatever it is, doesn't that mean that you have more time for the tasks before and tasks after?
Most of the episode's frameworks are heavily recycled: autonomous cars analogy for AI adoption, offshoring comparison, 'every tech transition feels the same' thesis. The reframe of the 1.7x issues stat as a detection/iteration problem rather than an AI-quality problem is the only genuinely contrarian moment; the rest circles familiar commentary.
There is nothing happening now that has not been happening in the past. It just feels like it's happening faster.
You can make analogy with cars. I don't know if you remember when it was, let's say, five years ago, that everybody thought that, uh, autonomous, fully autonomous cars will be on our streets in a year.
There are no external guests - this is a two-host conversation. The hosts demonstrate genuine daily hands-on experience with AI coding tools (Code Rabbit, Claude Code, prompting workflows) which gives them practitioner credibility, but neither holds documented senior roles at notable organizations and the episode intro explicitly frames them as amateurs.
This is a podcast about random stuff in which we, Darren and Victor, pretend we know what we're talking about.
That's my response to cloud code after it fetches code rabbit reviews. Uh, it always says, oh, let's fix high priority ones...And my response is no, let's fix all of them. I will tell you no only in case if I disagree
The episode cites several statistics (1.7x issues, 50-60% security flaws, 4-5x longer PR review) but Darren explicitly undercuts them as 'may or may not be correct,' and no sources are named. Victor's specific Claude/Code Rabbit workflow tips add some concrete value, but data quality is self-admittedly weak and most claims remain anecdotal.
let me bring in some stats that may or may not be correct
I read one, um, I think that cursor has like five, six times income per employee. They're on like I don't know how many millions per employee
Victor's genuine pushback on Darren's AI statistics - reframing them as irrelevant without knowing detection-and-iteration context - is sharp host craft, and there are several moments of productive disagreement (code duplication, 'right stages' for AI insertion). However, many of Darren's questions are leading or vague ('Could you agree with this…'), structure is loose, and threads are frequently abandoned before reaching resolution.
you know that it has more issues because somehow you detected those issues. Right. Otherwise you wouldn't know that there are more issues. Right. Cool. So whichever mechanism you had to detect those issues, did you feed it back and did it fix them?
I disagree with that.
Computed from the transcript - who did the talking, and the words that came up most.
#355: Picture your engineering team a year from now. A coding agent doing the coding. A testing agent on tests. A security agent on security. An infrastructure agent on infrastructure. All of them wired into GitHub and Jira, all of them working right alongside the humans. Not science fiction either - Atlassian and GitHub are already shipping these features. So out come the stats everyone loves to quote. AI code introduces 1.7 times more issues. Half of it ships with security holes. Code duplication is through the roof. AI-assisted PRs take four to five times longer to review. The response to most of it: so what? If you have a way to detect the issue and feed it back, that is just the SDLC doing its job. Couldn't care less if it is 1.7x or 50x more issues - what matters is what is left at the end, per feature shipped. Security holes? You have scanners. Detect, fix, ship. The only real problem is when you skip the detection or sit on the fix for months, and that has nothing to do with AI. Here is the one stat that actually sticks: PR reviews backing up.
Transcribed and scored by The B2B Podcast Index.
Narrator: Everyone talks about summer like it's supposed to be carefree. But if this season brings up money, stress, body stress, family stress, or social stress, that's real too. Grow Therapy can help with that. Whether it's your first time in therapy or your 50th, grow makes it easier to find a therapist who fits you, not the other way around. They connect you with thousands of independent licensed therapists across the US offering both virtual and in person visits, nights and weekends. You can search by what matters like insurance, specialty, identity or availability and get started in as little as two days. And if something comes up, you can Cancel up to 24 hours in advance at no cost. There are no subscriptions, no long term commitments. You just pay per session. Grow helps you find therapy on your time. Whatever challenges you're facing. Grow Therapy is here to help. Grow accepts over 100 insurance plans, including Medicaid in some states. Sessions average about $21 with insurance and some pay as little as $0 depending on their plan. Visit growththerapy.com booknow to get started. That's growththerapy.com booknow growththerapy.com booknow availability and coverage by state and insurance plan.
Viktor Farcic: I, uh, couldn't care less whether it introduces 1x or 0.5x or 50x more issues. If I have a mechanism to detect those issues and then iterate and iterate and iterate. What I care is whether at the end of that process I have less or more issues and per number of features I'm developing.
Darren Pope: This is DevOps Paradox episode 355 why AI Coding Slows Down Code Review welcome to DevOps Paradox. This is a podcast about random stuff in which we, Darren and Victor, pretend we know what we're talking about. Most of the time, we mask our ignorance by putting the word DevOps everywhere we can and mix it with random buzzwords like kubernetes, serverless, cicd, team productivity, islands of happiness, and other fancy expressions that make us sound like we know what we're doing. Occasionally we invite guests who do know something, but we do not do that often since they might make us look incompetent. The truth is out there and there is no way we are going to find it. Yes, it's Darren reading this text and feeling embarrassed that Victor made me do it. Here are your hosts, Darren Pope and Victor Farsek. Pitch your engineering team. In 12 months, you're gonna have a coding agent doing coding. You're gonna have a testing agent doing testing. You're going to have a security agent doing security things. You're Gonna have an infrastructure agent doing other things. All this stuff's gonna be integrating with Gabissues or jira. And all these agents will be working right alongside all of the humans.
Viktor Farcic: Okay, alongside, uh, alongside. That's the important thing.
Darren Pope: Alongside, which is fine. But you know, you might be thinking, okay, great, Darren, that's 12 months from now. Well, not really. It's not science fiction. Atlassian and GitHub and others are already announcing these features in their products today.
Viktor Farcic: Well, good luck. No, actually, no, I think it's great that they're doing it. It's just that the rollout of those things should not be instant, let's put it that way. So good luck to those who roll it out instantly to everything.
Darren Pope: I can't imagine anybody would even try it today.
Viktor Farcic: Oh, there will be, there will be people who will try it and there will be people who will say, this is working amazing. And uh, if you happen to know those people, you will realize that they're working on their pet projects and not in a real company.
Darren Pope: Well, could you agree with this that the SDLC is moving from a manual handoff type scenario to something that's a little more automated and elegant and something that we've always dreamed of in the past?
Viktor Farcic: Let me correct you there. SDLC is moving from partly automated, partly manual. And the manual part, I hope, will be partly replaced by agents. I don't imagine a world in which I need an agent to execute a command to run tests. I need agents to or people. And uh, this is the manual part to analyze if the tests fail and if they didn't, I don't need either agents or people to say, okay, continue. Kind of like continue with sdlc.
Darren Pope: Yeah, I sort of get tired of reminding my agent, continue. If my only value is adding continue. Then I could have another agent that goes around and checks. Everybody says just continue if you're stuck.
Viktor Farcic: Probably the shortest skill I have right now is a single sentence that when I execute, basically writes into the context. If I approve what you just did, you can continue without asking me. Mhm.
Darren Pope: Interesting. We're going to have naysayers to this and I think it's going to be somewhere in the middle. But if you think about the whole pipe, the whole SDLC from planning all the way through to incident remediation, monitoring and incident remediation. There are no places along that pipe that an agent could not be helpful.
Viktor Farcic: 100%. 100%.
Darren Pope: But what about fully replacing a human in each of those sections? Uh, like planning. I. There's still a Human, I think, to come up with the idea. But I also don't need to write, you know, a ream's worth of paper for the prd.
Viktor Farcic: When you say replacing, like gradually replacing or you mean you're gone? This is it.
Darren Pope: Okay, let's see if this is a hot take or not. I don't think we'll ever be able to fully replace because somebody has to train the AI to begin with. So that's my job. That's the gradual. I am going to be training it. The question is, will I lose my job after the fact or not?
Viktor Farcic: I'm 100% convinced that at least in a foreseeable future, AIs will not replace taste that we provide today. They can easily replace, uh, eventually people on many technical tasks. Like is this code change correct? Yeah, I think maybe we're not there yet, but we're getting there. Right. But that's correct from a technical perspective. Like is this correct? Are we really. Should we ship this thing? Does it make sense? Uh, that I don't see. Right. So I reserve the right to be in charge of the taste.
Darren Pope: Right. But will your bosses care for your taste if they're just trying to drive more revenue?
Viktor Farcic: Oh, there are many things that bosses don't care and they should. If you want, we can enter into that discussion and I think it's directly related. With AI, I had many bosses who don't care about the things that they should care.
Darren Pope: Let's not go deep into it, but let's play it out a little bit. In the past, bosses were. And I'm referring to a pointy hair boss, not a good manager. Right. Let's sort of split these apart. So the pointy hair boss from Dilbert type level.
Viktor Farcic: Mhm.
Darren Pope: We've had bad bosses and they hammer down on the humans. Well, they hammer down on the humans to replace the humans. To have the humans replace themselves with AI. What is that boss going to do when all he has is AI? Because we've talked about this. You have to become a man. I don't think you can be a boss of a fleet of AI agents.
Viktor Farcic: Depends on a company. In theory, I can easily predict that we will have companies worth billion in not so distant future that are run by a single person or a very
Darren Pope: small number of people.
Viktor Farcic: Yes, that I can imagine. Now if your question is, hey, take a typical bank and then kind of like all these software engineers are replaced by agents, well, good luck. That's similar to and I will probably be disrespectful without wanting to be is kind of like, hey, let's uh, replace all our engineers with cheaper engineers somewhere else. And when we do that, let's find even cheaper engineers elsewhere. That never worked out. Companies did that. And many of them are now putting people back in house like offshoring, near shoring, all that jazz. Uh, and that does not mean that people else were worse or better. But it's simply kind of if you're chasing the price, then you don't go very far. I mean, you go far and then uh. But eventually you realize that you're making a mistake.
Darren Pope: If you've never watched the movie outsourced, that's that story. I'll go ahead and spoil it for you because it's not that hard to figure out. Guy working in the States, his whole job gets offshored to. I believe it was India or Pakistan. I don't remember. So forgive me for people listening. And then as the story progresses, India hits all their numbers, everything's great, and then they decide to move everything from India to. I guess the storyline was China. I can't remember what the next one was because we're saving even more money. That's not going to happen in the world of AI because I think you're going to see a point to where. Okay, let's play that game out. We were offshoring from the humans to AI.
Viktor Farcic: Mhm.
Darren Pope: And there should be some price gains if it's trained correctly. Everything. Let's assume that's true.
Viktor Farcic: Mhm.
Darren Pope: But who's AI going to train once AI gets too expensive?
Viktor Farcic: That's the slippery slope. Okay, so let's. I'm going to simplify it now. Let's offshore it to opus. That M will be reduction in the cost. Okay, why opus? Let's offshore it to Sonnet. No, no, wait, wait, wait. How about Haiku? Oh no, Llama. We can run it on a laptop, man. That's cheap. The point. And I think that AI is going to bring us huge amount and already brought us huge amount qualitative gains on many different levels. And that's what I'm chasing. I think that whomever is chasing to be cheaper with software, if that's the primary motivation, is doing it wrong. And we saw that, uh, there were plenty of those stories and many of them turned up to be messed up, completely messed up. No results or negative results or what's not. So yeah, I want AI to do things, but only because some things can be done better, faster, whatever with AI not because it's cheaper. I read one, um, I think that cursor has like five, six times income per employee. They're on like I don't know how many millions per employee. Take total revenue divided by employees and what's the figure? And they're doing it really, really well. But they're not doing it because it's uh. And they're using heavily AI, but not because it's cheaper. It's just a net result that oh, we ended up earning more money, but we are earning more money because we are doing better with fewer people. Which is fair enough. Kind of like that's fine, I have nothing against that. But not because it's cheaper.
Darren Pope: I mentioned earlier that GitHub, um, Atlassian Opsara is another. All within the early part of this year have started launching sometimes live, sometimes in preview all of their agentic things. I mean it's going to be interesting to see how it plays out and see who is some of the first to really bite. Would I try out the technical previews? Absolutely. Give it a shot. Would I put it on anything that's not Greenfield? Would I put it on revenue producing stuff? I'd be very careful.
Viktor Farcic: Wouldn't you say the same sentence or sentences you just said for anything else?
Darren Pope: Absolutely.
Viktor Farcic: Like cloud for example, when it emerged. Let's say that you're one of the first adopters. Right. Wouldn't the story be exactly the same?
Darren Pope: Yes.
Viktor Farcic: Containers, kubernetes, whichever advancements uh, you had, like Rust, I don't know anything. The story is the same. Some go crazy and say, okay, we are doing transformation and we are going to do it in a way that we have a button that we are going to click on Friday afternoon, 4 o' clock next week. That never worked again.
Darren Pope: We've been reiterating this on the episodes where it's just me and Victor talking about this. There is nothing happening now that has not been happening in the past. It just feels like it's happening faster.
Viktor Farcic: It is happening faster. I don't think it feels it is happening faster and the scope is bigger but conceptually it happened many times before.
Darren Pope: Well, let me bring in some stats that may or may not be correct. But based on my use of AI for coding, they feel about right. AI generated code introduces 1.7 more issues, 1.7 times more issues than human written code across production systems now. Okay, whatever, we'll go for that. That will make some people happy. Let's call 50 to 60% of AI generated code contains security vulnerabilities or design flaws.
Viktor Farcic: Mhm.
Darren Pope: I can guarantee that based on things I'VE already written by the time I do security scans. It's like, oh, here's the five OSP things that you missed. It's like, okay, you should have known about that before you started. Okay. Mhm. Code duplication is extremely high with AI.
Narrator: Mhm.
Darren Pope: I've seen that. I felt that. I was like, hey, mhm. Why do we do that? In fact, that happened to me last night. It's like just one place. Please. One place. AI assisted PRs take up to between four and five times longer to process to be done with. And this one, I would say this is probably true. Only 3% of developers highly trust AI generated code. 71% refuse to merge without manual review. Yes. And we could have said that for any of the three gls, four GLS and everything else that occurred in the past. Again, we keep harping on that. It's just faster. It's just faster now. But for some reason, some people think this. And I can go back to the. The four GLS people thought that was a silver bullet. People thought that low code was the silver bullet. And for some types of applications it might be. But nowadays people are saying, hey, I want to. I'll use your idea here. I believe that we could have a billion dollar revenue company with under 10 people without even blinking. I think it's possible. In fact, based on something I saw yesterday, I think it could be done with one person.
Viktor Farcic: It will be, we'll give or let's say very small number of people. But I did something that I almost never. True. And that's. I haven't interrupted you for a single moment. Should we go back to the beginning of that list?
Darren Pope: Sure.
Viktor Farcic: Okay. What was the first one?
Darren Pope: First1 was AI generated code introduces 1.7 times more issues than human written code. Have you never seen my code?
Viktor Farcic: Okay, so uh, you know that it has more issues because somehow you detected those issues. Right. Otherwise you wouldn't know that there are more issues. Right.
Darren Pope: Right.
Viktor Farcic: Cool. So whichever mechanism you had to detect those issues, did you feed it back and did it fix them?
Darren Pope: You mean apply the SDLC to the problems that we're having?
Viktor Farcic: Yeah, kind of. Because I couldn't care less whether it introduces 1x or 0.5x or 50x more issues. If I have a mechanism to detect those issues and then iterate and iterate. Reiterate. What I care is whether at the end of that process I have less or more issues. And per number of features I'm developing, which I don't know whether that's included there because if it has, I don't know how many more issues in total, that's normal because we are doing more in total. Right, but let's say that per unit of delivery or whatever it is. So, uh, why is that a problem? It's not a problem as long as you can detect those issues and feed it back kind of. Okay, so you, dear agent, and you, dear Joe, you detect the issues and then correct them, and then let's repeat that cycle. It's as dlc, there's nothing wrong with it. And now here's what would be interesting for that study. Did that study detect those issues after they were delivered to production and sitting in production for weeks? If that's what happened, then that's very bad. But that has nothing to do with AI. Even better is the next one. Security issues. For issues in general, I can say, okay, it's difficult to detect. You know, some issues cannot be detected before it goes to production because only when users start touching it, we really see how it behaves. Blah, blah, blah, blah, blah. Okay, I can buy that one. The security issues, man, are you having a scanners in your pipeline? If you are, they detect the issues and then you check it and then you upgrade your libraries or whatever, or change the code and you go live happily ever after, unlike random issues that could be hard to detect. Security issues, aren't they easy to detect? And if they are, what's the problem? Get nothing to do with AI. The problem is the only problem with security issues, assuming that you are capable of detecting them. If you say, yeah, but I don't have time to fix it, that's the only problem. Because if you don't have a mechanism to detect, how do you know that there are more security issues? And there is a third option, kind of like you're capable of detecting them, but you execute whatever process you have to detect the months later and you say, oh, we've been running things with security issues for months. That's just silly. It's cheaper to do stuff, it's cheaper to detect stuff, and it's cheaper to correct stuff. Now if you apply it's cheaper only to one of those three, then you're not going very far.
Darren Pope: Anything else you want to bash in that list?
Viktor Farcic: If you want to go through the rest of the list, I can bash all of them.
Darren Pope: Code duplication?
Viktor Farcic: No. Maybe. No. Some, uh, say again?
Darren Pope: Code duplication.
Viktor Farcic: Oh, what's wrong with code duplication? Okay, so, uh, let me ask you a question. Why don't we want to duplicate code and before you Answer. Let me give you. Actually, let, uh, me give you a reason why you do want to duplicate code, right? Because sometimes, actually you end up with bloated libraries and bloated functions and whatever else that they're supposed to serve. 50 somehow similar but not the same cases that happens. Oh yeah, we have this function with 57 different parameters because, you know, this column needs this and that column needs this, and they're somehow doing similar but not the same things. That's the argument not to duplicate. Now, what is the argument to keep it in one place? And let me guess, is it maintenance because it's so much easier to maintain one library than five more specialized functions?
Darren Pope: Could be if you're lazy.
Viktor Farcic: Yeah, but now it's cheap to maintain anything. Uh, I've been fighting that. That's one of my demons. I've been fighting with Claude for a while. Kind of like, hey, why are you doing this? There is already a functional library that does sometimes does exactly the same. Exactly. And then you're right. Kind of you need to correct. But very often there's something very similar. You just need to add this argument to add the additional if statement to that function and you can use it. And then I gave up on fighting it. Why fight it?
Darren Pope: Well, my use case was defining a constant in two places.
Viktor Farcic: Okay, that's bad, that's bad, that's bad. Now let me tell you something that I'm. I have 98% confidence that code Rabbit would pick it up. Need to tell you about it.
Darren Pope: Yeah. Not sponsored by Cobabitra.
Viktor Farcic: No, not sponsored. I mean, you can use, um. Anthropic just announced something very similar. And there are. There is code emerge there, there are many. Right? The question is there is a multi layer approach to SDLC, with or without AI. We have code reviews. Forget about AIs so that we can catch those things, because we are not. I mean, uh, no developer. I mean, uh, even if you do it yourself, are you really going to catch. Imagine that you're not the only one working on that code base, right? Are you always going to catch that? Actually there is a library to do this, what you're trying to do, but you might not know about it. You're not going to catch it yourself either. Unless you're the solo person working that project. Right? And then you know it inside out.
Darren Pope: Do you want to talk about any of the others?
Viktor Farcic: Oh, I can go as far as you want. Depends. How long do you think you want us to do this?
Darren Pope: We're fine. AI assisted PRs. Wait 4.6 times longer to be reviewed, creating a massive review bottleneck.
Viktor Farcic: So AIs are queuing the reviews or,
Darren Pope: uh, AI assisted PRs because they're larger typically, but maybe.
Viktor Farcic: Or just waiting if they're larger. I think that that's terribly wrong. We need to get used to working much smaller chunk with AI. The fact that AI can write 10,000 lines of code is wrong. I mean, it can write 10,000 lines. You should still split it in smaller tasks. Right? So that's terribly wrong. Uh, it doesn't matter whether it's AI or a human or whatever, everybody will struggle, no matter the type of intelligence, with such big amount of changes. I think that actually the other day was the first time I hit the limit with Code rabbit myself. It was a project I just started, so kind of, I could excuse myself, but it says I, uh, cannot work with this size of a pr. It's the first time it told me, no, I cannot do this, not going to happen. But let's say that we get better and we chunk things in smaller chunks, uh, which we should do with or without AI. Then the question is, okay, so is this the same problem with sdlc, that we improve one part of it and we don't improve another? And if that's the problem, then again, that's the same problem as before AI. Right. Let's put it this way. Let's hire more developers to work on code, but require that every PR needs to be reviewed and not increase the number of people reviewing PRs. What do you get? Do you get anything different?
Darren Pope: No. Get worse, actually, Probably.
Viktor Farcic: There we go. Because. No, you get. No, no, because, you know, peer reviews, if they stay the same, same number of people, no AI included in any part of the process, same number of people reviewing. That means that you're not increasing your speed at all. You're increasing the speed of a part of the process.
Darren Pope: And we've seen speeding up just one thing doesn't help us anywhere else.
Viktor Farcic: Yeah, yeah, yeah, that doesn't help, right? If there is a Joe that needs to stump this without even having idea, because that's the whole essence of his job, then, uh, you have a problem which is called Joe.
Darren Pope: Well, let's sit there for a second because there was one more. We're going to skip it because I want to stay here for a second. The review pipeline was never built to support this level of automation of code coming in. It was built for human speed. We talked a couple episodes ago about human speed versus machine speed. How do we help sort of deal with reviews. Because if I'm getting. If used to. I was helping review two or three PRs a day, and now I'm having to review 10 to 15 PRs a day, and I'm being measured on the number of PRs I review. some point, I'm gonna get pretty wasted. And you can take wasted however you want to at that point of how do I actually manage this level of input coming into me? Again, something is sped up and up, uh, the pipe for me, and now I'm becoming the bottleneck. How am I going to solve that?
Viktor Farcic: Uh, let's imagine a world where actually AI cannot help you with PR reviews. Imagine that world, which I don't think exists, but imagine it for a second. Now, with assisted coding, let's say that in the past, you would spend 80% on coding, on average. Maybe not you, but, you know, on average in a company. 80% on coding, 10% of what comes before coding and 10% what comes after coding. And in that last 10%, that's among other things, PR reviews. Now we are shrinking those 80% to, let's say, 20. Imaginary number 30. Whatever it is, doesn't that mean that you have more time for the tasks before and tasks after? Isn't your time free now to do actually more code review in a role that Fedai cannot even help you?
Darren Pope: Let's say the answer is yes. However, I'm going to, uh. I don't want to say I'm going against it, but there's a whole lot of different brain cells used in reviewing code versus writing code or doing anything else. Uh, good brain cells. But if all I'm doing is reviewing paperwork, it's almost like I'm sitting at the desk and just reviewing paperwork all day long, because that's effectively what I'm doing. It's like at some point, I'm going to. Just like you and I have talked about. How many agents can you manage at one time?
Viktor Farcic: Yeah.
Darren Pope: Two or three. Maybe not 20 or 30. So just because I'm getting 15 PRs assigned to me to review every day, that doesn't mean I'm gonna be able to get through all 15. If I'm getting 15 every day because my overnight agents are doing all the work I'm always behind, I'm never able to speed up.
Viktor Farcic: That's true. I like your good and bad news. Which ones you want first?
Darren Pope: I, uh, think the good and the bad news is the same thing. Suck it up, buttercup, or you're out of a job.
Viktor Farcic: I was going in a Different direction.
Darren Pope: Okay.
Viktor Farcic: What I was trying to say is that, yes, I think that mentally, our jobs will be much more demanding. It's unavoidable. Right. Because if you're delegating more of the tedious work, then what's left is more demanding. Mentally. That's happening. I don't know how we will adapt to that. Uh, I honestly don't know, but that's happening. So that mental load that you just said through PR reviews, but I assume applies to other things, kind of like, oh, I'm now planning five PRs instead of planning one and working on that one. Right. That's another kind of issue. Cognitive load and, you know, running three agents. What do you do when you run three agents in parallel? Actually, you are forcing your brain to think about all those things instead of typing. Right? That's going to be tough. And it is tough, at least on me. So that's bad news. Right. The good news is that when you review things, there are silly things that you need to check, and there are important things that you need to check, and you can easily delegate silly things to AI.
Darren Pope: Give me an example of silly versus important.
Viktor Farcic: Does every. Does. And this is really silly. Does every function, uh, have a comment explaining what that function is? Because we are people who cannot understand what the function does by reading the code. You don't need to be doing that.
Darren Pope: Correct?
Viktor Farcic: Right.
Darren Pope: But what's an important.
Viktor Farcic: Huh?
Darren Pope: Uh, but what's an important.
Viktor Farcic: What's an important. Are we doing it right architecturally? Right? Is this really the feature we want to deliver? Does this fulfill even what we are trying to do? Right. On a higher level, I will, uh, tell you decently. Well, I'm not going to say perfect decently. Well, and I'm talking about specialist AI or agent whether, hey, this PR is technically sound. You should maybe double check it yourself, but on a higher level. Technically, right. Not kind of, oh, does every getter, setter, star, uh, is defining itself as camel case instead of whatever else. Uh, come on. That's a waste of your talent. Whether that pr. Architecturally sane. Whether. And more importantly, whether actually that's the feature we want to deliver, whether that's something that fixes a real issue, and so on and so forth. Right. Those are, I would argue, more important questions. And now you can spend more time on those more important questions than on silly questions. You know, the typical ones. Uh, I'm going to simplify it here and also offend the whole company. You don't need to spend time on things that you would normally go to sonar to check, which is a part of PR review. Right?
Darren Pope: Yeah. There's no need to because it's a machine talking to a machine or should be.
Viktor Farcic: Yeah, you can all sonar detected 57,000 issues. Cool. Go fix it. The same thing you would say to your younger colleague. Right. You would probably delegate. You're an old guy. I am as well. We would probably delegate that. Hey Michael, can you work on those other issues? That's extreme. That's the priority number one in company. That's when you start lying because that makes that person do a better job. When you lie that person straight into their face.
Darren Pope: Let's pull it back to reality. Let's say sonar or any of the other tools found. Not 57,000 but found. Let's call it 50 to 60, which could be a reasonable number for a newer project, maybe more. But you know, it feels like. What I mean, of course you could dump it off to Michael like you were just saying. But how would that be fixed in an AI world? Because let's think about it this way. Here. Here are the tooling that the machines have. We have mcps, we have agent to agent protocol and we have skills. Right. As a human, those are the things that we can do. Provide wire up, for lack of a better term.
Viktor Farcic: Yeah.
Darren Pope: So agents can work together amongst themselves.
Viktor Farcic: Yeah.
Darren Pope: So if we've provided all those things, where's the rub going to fall there? Because okay, I've got 50 or 60 security findings that we need to remediate because we're getting ready to do our annual PCI review. Great. Off we go.
Viktor Farcic: Yeah, I can argue that you shouldn't be doing that either. I try to focus on the work that I was doing. Not kind of. If you all try to keep let's say security medium and high at bay at all times, then there is no yearly review that matters. But go on.
Darren Pope: Well, here's where I think that agent agents can be useful and it sort of makes sense. Structured, well scoped tasks, by the way. Same thing for a human. Give a human a structured, well scoped task, more than likely you're going to get the right answer out the backside. Yep, nothing new there. Multi agent teams with a coordinator. Okay, again, this is sort of the judge model or however else you want to say it to where you've got another agent looking at all the other agent teams working together. Hmm, that sounds like a team metaphor to where we have a manager or a PM working with the actual developers that we shouldn't be surprised by that. Again, if we're Modeling. If we're, if we're thinking we're have a single agent do everything. Wrong idea. If we're thinking that we need to do a one for one replacement for a human to an agent, also the wrong idea. There's something in the middle there that will make more sense.
Viktor Farcic: Many of the issues that um, you mentioned earlier in the list is that somehow we decided to skip the things that we normally would never skip now that we use AI. When you said, oh, there are security issues. Yeah, so are we skipping now the security issue detection and remediation. That's the real question. And I feel that very often we do. Very often we have the other day, the other podcast we mentioned, uh, how AWS had an outage because they just chose to skip probably many of the things that they normally do completely in favor of using Kira. Right. And that was the problem. Eventually we will get probably in a very, very different sdlc. Very, very different. But uh, let's first make what we have work better with AI. Uh, that's my suggestion. Kind of can that be the first on a point zero, but across the whole life cycle. That's the important part. I don't care that you're now writing code faster if that's, that's the only thing that is happening.
Darren Pope: So here's something about that whole life cycle, something that I don't think we do as well for humans is everything can be easily timestamped as things move through agent to agent. So we can make sure things are done in order or have been done in order. Take your pick. However way you want to look at the clock. That's something we don't do as well with humans, except in maybe war room scenarios. Then we're watching clocks like crazy to make sure that we're doing things the right way. Here's where the agent agent handoffs sort of break down. We can lose context during this handoff between agent to agent. Maybe we had too much context built in. The agent that was doing the handoff and it gets truncated on the way in.
Viktor Farcic: I feel that that's, at least at this moment. That's actually my job. I'm personal. I don't know how others are doing it. But I'm not having agent handing off to another agent handing off to another agent situation yet I'm the coordinator. That's what I'm doing right now. Earlier in this recording I mentioned my shortest skill. Let me tell you the second shortest skill I have and that's that. And you try to figure out what I'm using it for. And when I want to go through all of them one at a time, no matter how important or unimportant you think it is.
Darren Pope: Go through them one at a time, no matter how much you think important or not important. Yeah, I have no idea.
Viktor Farcic: That's my response to cloud code after it fetches code rabbit reviews. Uh, it always says, oh, let's fix high priority ones. And that's going back to your stats.
Darren Pope: Sure.
Viktor Farcic: And my response is no, let's fix all of them. I will tell you no only in case if I don't agree. First of all, let's analyze again all of them. I want your opinion on its opinion about each of them one at a time. And then I will give you the last word, which is either yes or no. And that yes or m no. That has nothing to do with the amount of work we need to do. Nothing to do with the importance. I want to fix them all. I'm going to say no only if I disagree, even though both of you think that this is an issue.
Darren Pope: I ran into that last night where it was saying this is going to take a long time. I went, yeah, fix it. And you know, 10 minutes later it was done. It's like, okay, so I guess that's also sort of telling. When we as a human say this is going to take a long time, that's probably in the days or weeks, right? That's usually how we think. Potentially months depending on how bad to a machine a long time is minutes.
Viktor Farcic: And it doesn't have the concept of time. It actually knows the, it understands the concept of time from the data. Uh, it was trained on, on Internet. So when we discuss on stack overflow how this specific thing takes three days to fix, that's the response it is giving. You actually keep. Uh, when I'm creating PRDs from the very start I put it, it gives me estimate and it's so ridiculous. It's so ridiculous and funny that I keep it there. It still gives me estimate. And every PRD that I start working on takes anything between one day and weeks. It's entertainment part. Kind of like, hey, let's cheat up. How long will this take? Five days. Cool, let's do it. Yeah.
Darren Pope: I have another scenario to where it's always reminding me, hey, we need to fix this, we need to fix this. I'm like, no, we don't need to fix this right now. We need to fix it. But it's not important and I don't want to deal with it right now because I Know what the bigger fallout is going to be after that change is made? And I'm, uh, not ready to deal with that fallout yet. Which leads me to. Let's think about. We've got a handful of agents that are working together. Agent A makes a really subtle mistake, does the handoff to B. B doesn't catch it. So C starts building on the problem, and then D starts building on the problem until E finally just blows up because we've wrapped the bad with bad with bad. But again, how many times have we done that as humans? More than I'd like to count.
Viktor Farcic: That situation is what makes me, in a way, happy that without planning, I made certain career decisions. In the past. In the past, I was specialist, right? Kind of like, oh, I was amazing at, uh, Visual Basic. And this is not net. Just fii. All right? And I was amazed I had. During my career, I was specialized in certain things. And then through random things that happen in life, I went more towards generalization. Right? Kind of like, I want to understand how the system works. And I'm not good that I'm not better than anybody at any specific thing, but kind uh, of I'm very good at understanding how the system works. And that's the skill that I feel gives me an edge right now because now I can actually do the delegation. I'm the orchestrator of. Of the magic happening over there, right? Oh, yeah. Now. Now you create a pr. Oh, now let's do this. Now let's do that. Kind of, oh, this is wrong. This is good. Right? And this is not because I understand, let's say, go better than anybody else. That's not the superpower anymore. The wide knowledge, that's the new thing.
Darren Pope: One of the other places where agent. Agent can break down. I've seen this before, too. False completion reports.
Viktor Farcic: Yep.
Darren Pope: That's like, okay, make sure everything's fully tested. That was the directive. And it's like, okay, I'm done. Okay, where are the tests? Oh, I didn't write any tests. It's like selective memory.
Viktor Farcic: Oh, even better is kind of five tests failed. But it's not because of what we are working on. That's my third skill. Kind of like, I don't care if, uh, all the tests were passing before we started working on this. I could not care less this. Me paraphrasing the skill. It's not the exact word. I could not care less what you think whether we caused it or no. We are going to fix all of that. And if it's a flaky test that has nothing to do with uh, what, what we are working on. We're going to fix it as well. Kind of like 100% of tests need to pass, period. And you're not deleting any of them. It's OK to make it pass.
Darren Pope: Yeah, not deleting or. And I've run into this a lot, ignoring the test, like tagging it with ignore test or whatever the language.
Viktor Farcic: It was a while ago when I discovered, and this is because I didn't pay attention, that it just changed. Uh, in typescript, at least the test I'm using there is describe and then, you know, description of the test and test is inside. And would that do skip 100%. Would that pass?
Darren Pope: Of course they do. Because you skip them.
Viktor Farcic: Yeah. And you skip them because it's not relevant to what you think you're doing. And by the way, you have no idea that we are in a seventh context Clear. Uh, seven time clear the context so far.
Darren Pope: If you don't understand what he just said about clearing context, welcome to. That's probably the most painful thing that we run into in working with AI agents today is you've been working with something. It's like sitting down with somebody and you do a pair programming and then you say, okay, hey, let's, let's go eat lunch or go take a break for 10 minutes. And you come back and your pair forgot everything, like took a memory loss pill and now you're having to start all over again.
Viktor Farcic: Have you seen the movie Memento?
Darren Pope: I have not.
Viktor Farcic: Oh, man. So it's like Groundhog Day, except that he needs to figure it out. And so he writes, no. You've seen Groundhog Day, right? Yes. Yeah. Okay. It's like Groundhog Day where he loses the memory over and over again. Uh, doesn't wake up the same day, but kind of keep losing the memory and needs to kind of like start recording, start tattooing himself kind of with what he discovered. That's how I feel. We work with the agents, right? Except that instead of the tools we have, uh, our own memory or agent MD files or back to databases or what's on that kind of. But every session, let's start over.
Darren Pope: And sometimes you will do want to start over. It's like, okay, I'm getting ready. This is the other thing you'll forget to clear it out. And it's still building on what you've been working on, even though you're going down a completely different path. And that, uh, can be problematic as well. So how are we supposed to Actually update our SDLC to work politely with agents. Can you think of any? I mean we can drop it in any part of the pipe, but we've seen before that, okay, we've sped up one part of the pipe. Upstream is going to be able to get into it faster, but downstream is going to get backed up.
Viktor Farcic: There are two, I feel, important parts. And in both cases I will assume that you actually automated things that are repetitive. Right. You're not using humans to execute, execute, execute tests except while developing and things like that. Right. So you automated things that are repetitive and then you have people involved in many different stages of sdlc. That's my assumption right now. Right. And you keep people involved in all of those phase. In most of those phases. It's just that they're now augmented by with AI. That's the first step. The second step is that you don't need multiple people in the sdlc. I'm strongly believing that we are moving to the world where a team managing a product is no more than three people. And this is mostly for contingency, kind of like people leave. So three people max. When you say three people, that's not much different from before. But when I say three people per product, I mean product fully product kind of. If it has a backend and a front end, that's the same team. If it has a database, that's the same thing theme. If it deploys to production, that's the same theme. Right. Kind of like full end to end, one team up to three people. That's my new norm. Right. So you're augmenting each phase that previously required your intelligence to continue being your intelligence plus augmented with AI. And you're reducing the headcount per product. Right. So if you had five people involved in sdlc, you can do it with anything. Between one and three people, there will be more drastic changes, but I feel that those are the first steps also.
Darren Pope: You sort of leaned into this earlier when you were talking there of we want to keep humans in the loop where it makes sense in critical handoff parts. Like we want to be able to have the agent do everything to get ready to deploy to production. But the human should have the final button push to say, okay, go ahead and go. Everything is good, we're good to go. There's that final gate. Autonomy is where we're headed. That's a long term goal. AWS got there a little bit faster and found out what happened. But for now, the starting point is, okay, let's just like we did in the Past in a, In a nice Jenkins pipeline. Yeah, we don't want to go to production directly. Just, yeah, the human will go through, do the checklist and say, yeah, go ahead and go. And then everything else happens again. That's no big change. It's like we just don't want everything to be autonomous today.
Viktor Farcic: You can make analogy with cars. I don't know if you remember when it was, let's say, five years ago, that everybody thought that, uh, autonomous, fully autonomous cars will be on our streets in a year. And that never happened. We did not get fully autonomous cars all around, all the globe, uh, running wild years ago. Even though everybody predicted now the reaction from people, some people was, oh, yeah, so we're never going, uh, autonomous cars are just silly. We should abandon that. That's the wrong conclusion. Or those who tried to put autonomous cars on the street that early, that was a wrong conclusion either as well. Right. So we need the middle ground. And that middle ground is, yes, we are developing this. We are testing it, we're rolling it out. You know, it starts in San Francisco. Five streets, cool. The whole city cool. Five cities, ten cities, blah, blah, blah. Now we reach New York, probably the hardest, one of the hardest, biggest challenges for autonomous driving on the planet. Right after Istanbul. Right. And a citizen India. So, yeah, we are not abandoning the idea. Autonomous agents will be here. What will be the level of autonomy? I don't know. Uh, but don't try to roll it now. Even if technology is ready, you're not.
Darren Pope: There's a bunch of other items I want to throw in.
Narrator: One more.
Darren Pope: We need to monitor the agents, not just the output from the agents. Said differently. We need to monitor the output of the human. We need to monitor the humans, not just the output of the humans. The problem is we have machine speed now instead of human speed. Most of the time, machine speed is going to be faster than human speed. The only time that won't be true is if the human did RM-RF slash, then the machine took over these problems. You know, we're building out these pipes. We've got to make sure that if we're really trying to inject AI into our release pipelines, into our sdlc, we've got to rehab the pipe so that it's fast enough to deal with the onslaught of whatever's coming at it, whether it's PR reviews, whether it's new code generated, whether it is new features coming in. Again, all of these things. Or maybe again, following the AWS model. Now we have an agent that has operator Level access to all of our infrastructure. We got to be paying attention. And we have to make sure that it's at speed. Because once it actually works once, the business is going to expect it to work all the time.
Viktor Farcic: Yeah. Uh, uh, let's make it 10 times.
Darren Pope: Okay. 10 times. Okay.
Viktor Farcic: Yeah, yeah. Let's say that you have 10 steps and I'm ridiculous in it now, simplifying it. 10 steps in SDLC. Right. And I'll give you two options and you tell me which one is going to result in more features delivered to production. One 10 steps. Increase the speed of each of those 10 steps by 10% or double the speed of development. What gives you better results?
Darren Pope: That's a good question. Because I could argue either one depending on what my team structure is like.
Viktor Farcic: Yeah, but you know, if development is double the speed and whatever comes after development is the same speed.
Darren Pope: Oh, we're blocked.
Viktor Farcic: You're not delivering anything faster. Right.
Darren Pope: We're just backing up at that point.
Viktor Farcic: You're just kind of piling things. Like before we were piling issues, uh, in Jira or GitHub. Now we're piling PRs. That does not help.
Darren Pope: Well, some people think it is because they're meeting their goals, they're going to get their bonuses, but the people downstream aren't because they're not meeting the newly revised goals of PR processing.
Viktor Farcic: Yeah. And that's a change that those companies should have made long before. I kind of understand that your goal is delivery of something to production. That's the only goal that matters. The fact that you discover more bugs in qa, or that you delivered more lines of code in development, or that you detected more issues in security phase is irrelevant on its own.
Darren Pope: Could you agree that a fully agentic pipeline, sdlc, is coming?
Viktor Farcic: Augmented. Yes.
Darren Pope: Okay, augmented in the middle, but let's step back. Could we actually get to a fully. This is like the autonomous car.
Viktor Farcic: Like more distant future.
Darren Pope: M. More distant future. Yes.
Viktor Farcic: Yes. Oh, yeah. Imagine a factory and you have quality control. Uh, and you're making whatever screws in a factory right now. You're not checking every screw that comes out of the machine. Right. That would be just insane. You can just as well rid of the machine and do it, uh, manually, Completely. Right. You, you're making quality controls to understand whether the process works well or no. Right. You're extrapolating data from a subset of data. That's what we do in production. Like kind of. If you're storing metrics, you're not storing it for every single. Let's say traces. Like you're not storing traces for every single request to discover whether every single request worked. No, you're summarizing them, you're grouping them, you're aggregating them, and so on and so forth and trying to extrapolate. Okay, so what is the acceptance for me? Kind of like, I don't know, 99.9% of successful requests. Kind of. That's my meta. Cool. I'm not measuring all the requests. I don't know whether really 99.9, but from those random. That I picked up and stored in a database. So my sample is 1% of the traffic. And then I'm measuring whether 99.9% of the 1% is. Is, uh, okay. And if it is, I'm doing fine. And so on and so forth. You need to extrapolate from the sample. Right. That's the only way to truly do it.
Darren Pope: Well, I think the organizations that are going to win this are not going to be the ones that try to automate everything the fastest. They're going to be the ones that try to insert agents in the right places with the right guardrails and still keeping humans in the right place.
Viktor Farcic: I disagree with that.
Darren Pope: Really.
Viktor Farcic: I think that there would be companies that insert AI in almost all the places and keep humans.
Darren Pope: Well, that's going to be problematic because if you. I'm saying AI at the right stages.
Viktor Farcic: Yeah. Right stages. But right stages are almost, uh, so right stages are all the stages that are not currently automated. And they're not automated because they're not repeatable. So, uh, if you say no AI in PR review, but AI in development, you're not getting any benefit. And if you say okay, so PR is the right place. Yeah, but security analysis is also the right place. Observability. Is that the right place? It is the right. But I would argue that no place that is maybe not no place. Majority of places that are not repetitive automated already are, uh, the right places. Not full autonomy. That's not what I'm saying. Just to be clear, I'm saying augmented with AI.
Darren Pope: Yeah. I think at some point in the future it'll go beyond augmentation and it will just be AI. Oh yeah. That's longer term. Yes. Now, for some of you listening, you're thinking, at my company, There is a 0% chance any of this will ever happen or will happen after I'm retired in 50 years. Well, let me give you a cautionary tale.
Viktor Farcic: Mhm.
Darren Pope: Because I have one. I've brought it up before. My dad Used to do sheet metal fabrication, H vac ductwork, and would create or, uh, build H vac for fabric factories, textile factories, these factories, and this was in the late 80s, 40 years ago. Build these massive hundred thousand, 200,000, 300,000 square feet facilities with weaving machines and everything else. Once it was built, once it was online, the whole thing could run with five people. Five. These were the people that went around and did maintenance on the machines. That was it. Uh, that was 40 years ago. To think that that kind of automation is not coming to knowledge work. Forgive me. You're fooling yourself. You're delusional.
Viktor Farcic: That is exception.
Darren Pope: Of course, there's always exceptions. But yes, go ahead.
Viktor Farcic: There is a big, big, huge exception. And that's if you have some kind of monopoly.
Darren Pope: Oh, absolutely.
Viktor Farcic: Right. So kind of like, oh, I'm a banking system in Argentina and I'm inventing now. I'm not trying to ridicule Argentina, uh, and actually government, just this. I convinced the government that AI uh, is not allowed any anywhere in the banking system. So actually, no competition can kill me and I'm the biggest one there. Then you, you honestly don't need it.
Darren Pope: Fair enough. But for those of you that aren't working in those levels of monopolies, I'm not saying look for a job, but either brush up your skills or start thinking about what do I want to do for my next career. So what do you think? This got really heavy really fast. At the end, however, the Slack Workspace. Go to the podcast channel and leave your comments there. We hope this episode was helpful to you. If you want to discuss it or ask a question, please reach out to us. Our contact information and a link to the Slack workspace are@devopsparadox.com contact. If you subscribe to Apple Podcast, be sure to leave us a review there that helps other people discover this podcast. Go sign up right now@devopsparadox.com to receive an email whenever we drop the latest episode. Thank you for listening to DevOps Paradox.
Narrator: Are you noticing your car insurance rate creep up? Even without tickets or claims, you're not alone. That's why there's Jerry, your proactive insurance assistant. Jerry handles the legwork by comparing quotes side by side from m over 50 top insurers. So so you can confidently hit buy. No spam calls, no hidden fees. Jerry even tracks rates and alerts you when it's best to shop. Drivers who save with Jerry could save over $1,300 a year. Don't settle for higher rates. Download the Jerry app or visit Jerry AI Libsyn today. That's J E R R Y AI Libsync.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.