
The AI Native Dev · 2026-06-25 · 1h 1m
Key moments - from our scoring
Substance score
65 / 100
Five dimensions, 20 points each
Patrick Dubois, Thomas Dubinoff, and Daniel Jones discuss the organizational and process changes enterprises must make to succeed with agentic coding. The conversation reveals that AI agents are forcing companies to finally address long-neglected software development practices - CI/CD quality, testing, coding standards, and value stream mapping. Unlike human teams where poor process could be blamed on individual laziness, agents make inefficiency immediately visible and costly through token consumption. The panel highlights that scaling agentic development requires new organizational patterns: platform teams must expand beyond infrastructure to provide AI context and skills management, while feature teams and non-technical builders (PMs, designers) must be empowered to contribute directly to codebases. DJ from Resync emphasizes the PAP framework - incrementally improving code quality while delivering features - as the balance between speed and maintainability. A key insight: when development velocity increases through agents, product discovery and QA often become bottlenecks, requiring strategic thinking and deeper customer engagement rather than just faster ticket writing.
No single team owns it; responsibility typically starts with a small incubation team (similar to how DevOps evolved) and then distributes to developer experience or platform teams, which must expand their scope to include AI context management and skills distribution. In larger organizations, specialized teams (cloud platform, AI platform, data platform) may each handle their domain rather than one monolithic platform team.
The Dora report shows that organizations need strong CI/CD pipelines, comprehensive testing, agreed-upon coding standards, and value stream mapping before agentic coding will improve productivity; without these basics, agents slow teams down rather than speed them up.
Track the number of turns (iterations) and tokens an agent consumes to complete tasks; by optimizing context, tools, and prompts, you can reduce both metrics systematically - something impossible to do with human teams.
Product discovery and backlog depth become hidden bottlenecks; teams must redirect PMs and designers from administrative ticket writing toward strategic customer research and deeper thinking, not just faster code generation.
Confidence and imposter syndrome; PMs and designers fear that code generated outside traditional development workflows will be scrutinized or rejected, even when it meets quality standards. Building trust through clear examples and encouraging direct codebase prototyping (not disconnected Figma mockups) helps overcome this barrier.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains substantive discussion on organizational patterns, process changes, and practical adoption challenges (e.g., development maturity as a precondition for agentic success, platform team responsibilities, risk-based PR review), but much of it recycles familiar software engineering concepts (CI/CD maturity, observability, testing) with an agentic lens rather than introducing genuinely novel ideas. There is useful specificity on bottlenecks and governance models, but also notable filler (introductions, podcast plugs, off-topic banter about boats and weather).
The bottlenecks that that become apparent are often deficiencies in just doing software development well.
Now that it's agents doing the work and tokens and the cost of those, everyone's gonna be really like keen on making sure the software development lifecycles was as effective and as efficient as it possibly could be.
The panel articulates some genuinely counterintuitive points - e.g., that token cost creates incentive for process discipline that human teams never had, and that observability queryability for agents is a new architectural requirement - but the core thesis (good practices enable agentic success) is largely a reiteration of DevOps-era transformation patterns. The Assembly Line tool concept and the 'nudging' workflow are mildly novel in framing, but the underlying principles (deterministic checks, feedback loops) are established software engineering practice.
when it was humans, it was like, oh well, Timmy's just lazy and that's why stuff isn't getting delivered. And they didn't focus enough on process... But now it's when it's agents, it's like, well, that's gonna cost us money.
Good engineering that we bring in there... The engineering becomes creating the the software factory or the assembly line.
Patrick Dubois ('godfather of DevOps') is a genuine industry luminary with deep transformation experience spanning multiple eras. Daniel Jones and Tamuz Dubinoff are practitioners actually shipping agentic products at scale (Resync doing AI transformation consulting; Autonomy AI building an OS for non-technical contributors). All three are actively doing the work they discuss rather than theorizing, though the session format (panel at a conference) somewhat limits depth compared to 1-on-1 conversation. The guests have credible skin in the game.
Most people know me from the DevOps transformation and the cloud native transformation.
We are a AI native transformation consultancy... we offer various services around that, and I do a lot of the training for agentic coding training.
The episode provides concrete examples: Autonomy AI's 2700-component codebase reranking, Resync's PAP framework, Adevo's 72x speedup with 8 years of work in 11 months using a third of the people, feature-flag-driven workflow (plan-merge-polish), and risk-level-based token budgeting. However, these are largely anecdotal claims without detailed metrics, timelines, or financial breakdowns. The TESL Patterns site is mentioned but not deeply explored. Some claims lack supporting numbers (e.g., 'more and more staff' on observability agents catching bugs).
We see that there's kind of the span of work that needs to be done... Adevo, uh they're doing some really great work where they've got a team that has gone full agentic... they are now going 72 times faster... 8 years worth of work and they're on track to complete that in 11 months with a third of the people.
we have agents working on 160 plus organizations... the ones that are better quality, the agent does a better job faster.
Simon Maple is a competent moderator who poses reasonable opening questions and steers the discussion topically, but there is limited sharp follow-up questioning or productive pushback. When bold claims are made (e.g., 'removing human code review entirely,' 'agents opening PRs without review'), the host rarely presses for caveats, failure cases, or concrete prerequisites. The panel format encourages panelists to add to each other's points rather than being challenged. There is notable drift into tangential discussion (UI philosophy, Facebook testing) without tight retrieval to the core topic.
Tell us a little bit about more about yourself, uh your role, and and maybe a little bit about Resync.
How many hours of sleep did you did you miss when when the DevOps role came about?
Computed from the transcript - who did the talking, and the words that came up most.
Enterprises are finally being forced to care about their software development lifecycle - not because anyone suddenly got disciplined, but because agents cost money and the waste is now visible. When it was humans, it was "Timmy's just lazy." Now it's a line item. Simon Maple sat down with Patrick Debois (the godfather of DevOps, now DevRel at Tessl), Tammuz Dubnov (co-founder and CEO of Autonomy AI), and Daniel Jones (Head of Product at re:cinq) at AI Native DevCon London for a wide-ranging panel on AI enablement - who owns it, what's breaking, and what the organisations getting it right are actually doing differently.
Transcribed and scored by The B2B Podcast Index.
Now that it's agents doing the work and tokens and the cost of those, everyone's gonna be really like keen on making sure the software development lifecycles was as effective and as efficient as it possibly could be. But when it was humans, it was like, oh well, Timmy's just lazy and that's what that's why stuff isn't getting delivered. And they didn't focus enough on process and you know value stream mapping and all those important things. Now it's when it's agents, it's like, well, that's gonna cost us money, and we can easily see that it's costing us money.
Most people know me from the DevOps transformation and the cloud native transformation. And I love every new chaotic period in the industry because that's where the learning is happening. And right now that is AI encoding. The way I think about it is I go back to pre-gen AI, and I imagine I have all the manpower and all the time in the world.
And basically with that mindset, we say, okay, now you have Gen AI. DevAcient, there is no excuse not to do everything you ever thought you'd want somebody to do, and then automate it so it happens by CI automatical. The AI Native Dev is a podcast with developers and engineering leads at the cutting edge of AI and agentic coding. Join your host, Typer Johnny, and meet Simon Maple every week as we chat with the most exciting voices in AI and tackle the biggest questions that we're facing developers today.
This episode is the AI Native Dev. Hey everyone, hope you're enjoying the episode so far. Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development. Whether that's talking about the latest tools, the most efficient workflows, or defining best practices.
But for whatever reason, many of you have yet to subscribe to the channel. If you're enjoying the podcast and want us to continue to bring you the very best content, then please do us a favor and hit that subscribe button. It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you. Alright, back to the episode.
Hey there, Simon Maple here. Welcome to another episode of the AI Native Dev. This week we're sharing another fascinating conversation that we had with some of our AI native Devon speakers. This time on the topic of AI enablement and what enterprises need to get up to speed with a genetic development.
I sat down with the godfather of DevOps, Patrick Devoir, Autonomy AIs, Thomas Dubinoff, and Daniel Jones from Resync. What I think you'll agree was a really intriguing discussion. And what I love is the way that some discussions you very often get question-answer, question-answer. But these folks were very, very passionate.
They really added to each other's points. And it was a true conversation. I hope you enjoy it. Hello, Simon Maple here, and I am at the AI Native DevCon 2026 in London.
And we've got a wonderful panel here to talk about AI enablement within organizations, how we can be most successful at scaling uh agentic development across our organizations. I've got a wonderful panel, as I mentioned. We have Daniel Jones, DJ, uh head of product at Resync. We have uh uh Tamuz Dubnov, who is the co-founder and CTO of Autonomy AI, and we have Patrick Dubois, who is the DevRel at TESL, the DevOps overlord and godfather of all things pipeline and workflows in DevOps.
Welcome everyone to the panel. How are you all doing? Not too bad. Given that there was a party on a boat yesterday.
It was. Yes, and I'm just about functioning. So that's good. Well, that's that's that's always a benefit.
Always a benefit. Even with good weather in England, right? Yeah, yeah, even with good weather. On a boat on the Thames and it was sunny, who knew?
Who knew that was a possibility? So why don't we start off with uh why don't we start off with you, DJ? And we'll let's go left to right and we'll um tell us a little bit about more about yourself, uh your role, and and maybe a little bit about Resync. Sure.
So um I am head of products at Resync, which means productizing our services. Uh we are a AI native transformation consultancy. We were big in uh the co-founders were big into cloud native transformation, wrote the O'Reilly cloud native transformation book. And what we see is that organizations are repeating the same kind of mistakes when trying to adopt a disruptive technology of just plopping the technology in, not changing anything about how they work and being surprised that they're not going faster.
So we offer various services uh around that, and I do a lot of the training for uh agentic coding training. Awesome. Uh so Thomas, co-founder and uh CTO at Autonomy AI. Uh, we built an operating system on top of your codebase uh to enable an easy platform for non-technical or semi-technical individuals to actually meaningfully contribute to the code base and the organizations, something we can do their visual iterations and actually get stuff uh ready, do backlog items, do new items, and put everything inside a brown field code base that ends up in a meaningful PR that can actually be merged.
Sounds good, Patrick? Patrick Duel. I'm from Belgium and I have gray hair, so I've lived through a bunch of transformations in my life. Um most people know me from the DevOps transformation and the cloud native transformation.
And I love every new chaotic period in the industry because that's where the learning is happening, and right now that is AI and coding. And so I'm looking not just at the tools, but also where kind of the new organizational patterns are happening as this field is forming in the industry. So that's kind of my gem right now. And every time there's a chaotic transformation, the hair gets greyer, right?
Yeah, it's hard to know whether it's the cause or the effect. Like, yeah, I know. Brilliant, brilliant. So round this table, we're gonna be chatting about all things AI enablement.
Let's kick off with an interesting question about when we want to scale AI agentic development across our organization in a way where we can control it, it is something we can govern, it is something we can uh do in a meaningful and controlled way. Who owns that? Which team, which individual, who's responsible for making sure that happens well? It's it's a good question and one that I think a lot of people are grappling with.
And I should imagine that we'll probably end up with uh new departments, new teams, new named units and movements forming. Um, I'm working with one customer who happens to have a developer experience team, which is kind of overlaps with the platform team. And so there's a an amount of kind of uh it makes sense for those folks to be doing things like building and distributing and assuring the quality of skills, for example. Platform teams are kind of interesting in that they're there to help developers by presenting a you know a raised level of abstraction, reduced cognitive load, and you can imagine that like making sure that agentic coding goes well and that you know context is managed and available to agents would fit into uh that group of people's uh responsibility.
But I've seen other uh customers as well who have a DevOps team, which probably shouldn't get too uh deep into that, whether it's an anti-pattern or not. Um but I've seen those folks, like the people that get lumbered with like you're doing the CI CD pipelines for people, like those people also uh being kind of made responsible for these things. Patrick, I'd love to go to you there. How many hours of sleep did you did you miss when when the DevOps role came about?
None. None. No, it's quite easy. The big thing is that every new change in an organization usually has a like a small incubation kickstarting team.
We had the agile team. You know, we don't talk about an agile team anymore. Like everybody's doing agile. We had the DevOps team, and and and that kind of is just a way of scaling things out.
Like you need that one team that gets duplicated in multiple teams, like, hey, can I repeat this? Can I repeat this? And then I think the pattern is indeed right now is it falls back to a central team kind of nurturing all the other teams in a way. So I think the challenge, what you were saying, like, yeah, it might end up in the developer experience platform team as a role.
Their challenge right now is they're still often focused on, oh, we'll provide the infrastructure. They're not TAI savvy tech coders. So that's where there's now a little bit of a vacuum in kind of the companies where, like, okay, you know, they're not the best fit, but they you they do control the spends, they have the models, they have the gateways. So you can see their stack kind of expanding.
Now, whether they will stay in one group as a platform, that's uncertain. What you might see is that you're gonna have the cloud platform team, the AI platform team, the data platform team, and they're all kind of platform supporting this. So it's not because there's like one platform group that they have to do it all. And that's the same thing with like feature teams.
They can't do it all because they have a specific focus. Small companies, yes, same team, that alliance. Bigger companies, they might have like different skills that they're going for. So anyway, that's kind of how I look at it.
Yeah. Um as a CTO, you know, you've got your engineering team there, they want to you know use AI as much as possible. They don't want to be slowed down through through you know extra governance and things like that. Where's the balance between you know having a team wanting to add governance, uh, add ways of working, showing a golden path, versus at this stage of AI's you know, maturity, immaturity, allowing people to kind of find their own paths and understand what's good through developers playing around and trying different things.
Sure. So for us, speed up has been getting the whole team aligned. Each person finds a new skill that's important in our workflow in some manner, and then we have a way to share it with all of them and have everybody else say use it, it self-learns. Uh, but the mechanism that's really kind of sped up our RD is getting to focus on the stuff that really matters.
And we get that because we offload the stuff that doesn't need an engineer to our actually our product team and our designer team. Uh so part of the game for us has actually been well, we have builders that are engineers, builders that are not engineer. How do we manage the non-technical opening the PRs? And how do we manage the technical side getting those PRs?
Uh and that kind of reshuffle is actually that's what has really helped us uh get through that AI native speed up. Uh, because you have to do change management internally, both on the receiving end and the the new builders. Uh and certainly when that does click, that's where you see everything running forward, and you see the whole team kind of playing together in in harmony. Yeah.
I think that's interesting because you know, when you called in the past DevOps bridging two worlds, like developers and devs coming together, I think the term that like comes up a lot now is the AI product engineer. They're like in both worlds trying to cross the silo of the bridge. Like you were working on like how to align that better as a field. And uh I think that's fascinating in the way that like when the toil or the daily things are being more and more done by AI, we can start like redirecting them.
Are we it's not about like you know how are we building the thing right, but are we building the right thing? And that's kind of where the product focus like comes stronger again. And I I think I like kind of how these roles are blending. Uh it's not for everybody, I guess, but like it is kind of a good alignment to have in an organization.
So we see that there's kind of the span of work that needs to be done, and frankly, there's the span of people that care for different parts of the work. So there's the stuff that the designers and the PMs really care for, and honestly, that's not the stuff the developers care for. So the fact we can say, great, well, those builders are now doing it. Hey, developers, UK of our like architecture scaling, like uh the sandbox environments and how they live, you can focus on that because all this uh let's say annoying work you used to have to do.
Well, now the builders that care about them, they also have the authority to make those decisions, they're the ones doing it, and hey, you can focus on the stuff that you care about and making them uh more agent tech friendly and faster developer. That's just happening in our team organically, they're already there, they're already trying to speed up, and they're already figuring it out on their own. For me, the focus of CTO is enabling the side that isn't pushing or wasn't pushing as much, which is the non-technicals.
And when you get them up to speed, suddenly you get the full speed up and you get people working on the stuff that they care about, and then that's the stuff that they're fastest on. Yeah, and and it's and speed is the most interesting thing here because I think you know, when we think about a year ago, two years ago, um, gosh, I'm trying to work out how how old uh AI in the mainstream is now. But if you think about the last couple of years, we're really thinking about let's just try and get AI usage, let's try and get AI adoption as almost like um you know, get it in the hands of developers, play around with it and see what see what works and what doesn't work.
It feels like more and more now going forward, we're thinking about how can we make it work effectively for us and how do we actually get the entire organization using best practices? So let's flip it a little bit and say, what are as organizations start rolling out um AI seriously in terms of uh, you know, as a practice across every single uh development team, what would you say the, and I'll ask you this, what would you say are the greatest bottlenecks today, uh, which really maybe it's platform teams, maybe it's other teams, but they're that they're struggling with as part of that scale and that AI enablement across other teams?
The bottlenecks that that become apparent are often deficiencies in just doing software development well. And I I think that we may probably discover in the next couple of years that we spend uh a lot of our time repeating ourselves of the things we've been saying for the last 10-15 years of is your CI CD any good? Do you have tests? Uh, do you have coding standards that people actually agree on?
So all of those things uh kind of get exposed. And the the uh Dora report showed that if you are not very mature in your software development practices, then agentic coding is likely to make you go slower. If you're doing well in development maturity, then you'll go faster. The problem is we don't know where that tipping point is.
Yeah. So there are all these things inside the development process, but um, even with successful agentic coding uh adoption, I've seen in a couple of customers and people that uh I've spoken to um on the on the podcast, um uh probably shouldn't plug our own podcast there. In the comments. You're you're on the R Native Dev, I assumed you were talking about that one.
Of course you do have your own podcast as well. Indeed. Um but we've seen uh uh people there of you suddenly speed up uh the act of creating software and then uh products are caught flat-footed, and they're like, Crikey, we didn't have enough stuff in the backlog. And that's hard to speed up because it's strategic and you need deep thinking and you need to be talking to customers, understanding their needs.
So we've definitely seen cases where software development has started to go much more quickly. Um, and then in people that do QA as an after-the-fact thing, that's being a bottleneck. I'm not sure I'd recommend that pattern generally. But yeah, uh product tends to become uh a bottleneck, and then we need to empower those people to uh to have more free time and to do the strategic thinking and the discovery work that presumably they they really enjoy and they would prefer to be doing rather than writing Jira tickets that developers aren't going to complain about.
Yeah. How about you, Petri? I think if you look at the door, you know, first thing is the other option and kind of is it like you know, cranking out new code, that was you know almost like the developer productivity lines of code for a long time. Then we learned, okay, it needs to work in production.
So we brought in the metric actually, like how many times do we need to rework it if it isn't good? Like that kind of is your defect rate, almost like that. You kind of need to be on control both in generating and kind of making sure the defects are not like are in balance there as well. And I think I what I like as an uh as an almost like new um proxy metric that like when you look at like not just the users coding with AI, but if the agents are doing the coding, you look at uh the things that like how many how many uh almost like turns does an agent need to do to do their job effectively?
And you kind of can get that number down by you know, you can select another model, you can kind of give it the right tools, and you can provide it better context. So in that way, it is not anymore like can I do better code in production, but can my agents do better code towards production? So it becomes like delegated. So you need to think out how effective they are, and that's kind of a flip that I see is now slowly moving from the def working with AI to the def instructing the agents to use AI to do it.
And that that's kind of like an interesting proxy metric, I think, to track. Uh and I think there's something both amusing and mildly depressing about this in that when we have uh software factories and that idea matures and becomes more commonplace, we're gonna have this infinitely tunable and measurable way of developing software, and we see exactly how effective it is in that way that you mentioned of looking at how many turns are used, how many tokens are used. And unlike with human teams, we're gonna be able to A-B test.
We could like have uh prompts and software factory setups that you know use one approach versus another, and we can reset their memories and try it a second time and see did that work better, did it not? And you can't really do that with software engineers and real teams because you can't like do a men in black and memory wipe them. Um but the thing that's like amusing and depressing is um that now that it's agents doing the work and tokens and the cost of those, everyone's gonna be really like keen on making sure the software development uh lifecycle was as effective and as efficient as it possibly could be.
But when it was humans, it was like, oh well, Timmy's just lazy and that's what that's why stuff isn't getting delivered. And they didn't focus enough on process and you know value stream mapping and all those important things to tune it. It was just, oh, the humans can suffer. But now it's when it's agents, it's like, well, that's gonna cost us money, and we can easily see that it's costing us money, so now all of a sudden we care.
Yeah. It's funny we face this challenge today because we we have agents working on 160 plus organizations, which widely different code bases. And we see the ones that are are better quality, the agent does a better job faster. The ones that are lower quality, we have an onboarding flow that the agents get better after a few tasks, they onboard themselves to your repo.
And we see that in the difficult projects, it takes them longer to onboard to good quality. But we cannot go to we we focus on onboarding quickly to an organization. We go top-down, so that means the the business leaders we talk to, we cannot say, hey, like a sprint to refactor so the agents can have a better time, so everything will go smoother because they don't really care about that. They just care about speed, like you mentioned.
So for us, it's been all about how do you mix in doing the user-facing work quickly while writing the natural code churn that the organization has to slowly refactor. But we have like a we call it the PAP framework. And as you do work, you you slowly refactor bit by bit, so it's easier for the next agent that comes along in that part of the code. But you do not come and say, hey, dev team, you have this task, make it easy for our agents to work there.
You have the agents doing the work in bite-size going back to the PR fatigue in a way that doesn't inflate the PRs because then people won't go over them. So it's it's a really it's a balancing game. Getting the leadership happy because features are being delivered, getting the developers happy because PRs are not too big, and getting the non-technical PMs, the designers happy because they can do the work, all while slowly pushing your agenda better and better coding standards.
So everything needs to happen uh kind of weaved together. And I'd love to kind of while while while talking about that, you you previously you mentioned about the kind of like the the technical devs, the traditional devs, and the non-technical, so maybe it's a PM or something like that. When we talk about rolling out an enablement here, do you see different types tend to see different bottlenecks between a technical dev and a non-technical dev when adopting like a more of an agentic process there?
Uh yes. Okay, so we'll go with the technical side. Uh the technical side, there's an ego that we all know. Uh we are working through that, whether it's with clients or whether it's internally, like the devs.
Uh so for us, they they interface with us, they can use our platform. Uh not not to show off, but like the PRs that come out of them versus out of the platform versus developers. Often our platforms they do a better job because they map your your coding practices much more in depth than developers and they keep much better tracks of existing uh components. Like we have clients with 2700 components, no developer will know which component to use.
We do know we do data science on it, and and we can actually re-rank and pick between them in a much more intelligent way. But you have to come over that ego of the developer viewing the PR. So we tuned it. We're a little gentle when it comes to them.
Uh and then they're very fragile online developers. I I'm not gonna say that. Did the content say that? I can't remember.
Uh and then on the PM side, it's a confidence thing. It's like, hey, I'm I'm okay opening a PR. Uh it's getting the dev team to play along. We we have some experiences where uh PMs open PRs and the dev team pushed back.
And then we asked like the team lead, look at the code, is it good? It's like, oh, this is solid code. Why did the developer push back? Because it didn't come from a developer.
Uh so when that's worked out, we basically work on a conference of the PMs saying, hey, you don't need to write a ticket, that's not the deliverable. You don't need to make an HTML prototype that is disconnected, that is not the deliverable, that's effectively a ticket in a Figma. Uh you need to actually prototype inside the code base. You need to edit existing stuff in the code base.
Hey, it's possible, hey, it's feasible, hey, effectively it's really easy. And when you work with the agent, it already knows the constraints from the code base built in when you interact with it. And yes, when you click send to dev and it goes to the developers as a PR, be proud and don't be scared. Uh and that's like a big enabler.
It's almost like an imposter syndrome, I guess, for non-technical folks to feel that maybe they're resistant to do things with confidence because they feel like they're not a dev and the devs gonna scrutinize it more almost. They're scared. They don't know how to answer technical questions. Yeah.
And that's a part of the process we're we're going through. Like, you don't need to. That's the beauty of the agents. They know the code base, they manage it for you.
You need to give us the product intent, and that's what matters. The agent delivers something visual, then it's a part of their app that is rendering and running and they can interact with it. If you do the product review or the design review, depending on your persona, you should be confident. The code works the way you want it to, the code quality is there, you don't need to worry about it.
And hey, if there's like chess run calls on the back end side, the developers will take it. So let's move in a different direction now where we're going to talk about ways in which agentic development, or rather, patterns of agentic development, how people are doing this in the real world and what we see from people externally and in what they're saying they're doing and how they're approaching it. Patrick, you recently created a really interesting tool. It's called the TESL Patterns site, Tesla.
io forward slash patterns. And it does a ton of research in the background. I think it uses the Carpathy wiki approach in the background. And it provides a bunch of patterns that people can look at and say, ah, okay, I can see that people are thinking and doing things in this way, this approach, and you know, cites various uh social uh, you know, I guess uh uh ways in which people are going into more depth to say, yes, I I did it like this on on in this infrastructure in this way.
Talk us through a little bit about that and maybe highlight a couple of the more interesting patterns that you found. Yeah. So I think one of the reasons is that when you do surveys right now, you get like an enormous spread. Like, you know, person A is savvy, the other ones are not savvy, and then the effectiveness and yes and no, that's hot.
So I figured that what if we look at the socials on how they do things and kind of capture that signal. No, that has a little bit of a problem. It could be vendors saying how they do great things and practitioners. But kind of the whole idea is that we filter out like how how some of the people are using this, and you know, if there's many people doing similar things, that boils down to kind of a pattern.
Now, one of the surprising and unsurprising things is almost like what is good for a eye is actually good for a human and the other way around as well. Like given that you're working with an agent, um, if you actually give your agent good documentation, lo and behold, it performs better. If you give them tests, then you actually know whether that's still working, yes or no. Uh so if you have observability, it can expect itself.
So that's kind of like you know, when you start thinking about like these analogies on like what is actually good and what is not, you can go up to um rituals, for example, on the team. Uh, we have a team lead and they're used to doing scrum and retros and kind of those things. But what if you would have an agent review kind of at the end of the sprint? What were they missing?
How effective were they are? But it's not also what do you do as a team lead? You set a goal for a person. So your goal sets the measurement.
So that is another account. So you find like people doing like those little ceremonies, or for example, they do uh writing um almost like context together as a pair, so you get nuances. So that's the pair programming, but now with context. So that's kind of interesting that those signals are bubbling up uh in a certain way.
Which is good because it actually goes back to what you said earlier about it's good software development practices. Everything that you mentioned there should be things that we are doing today manually, but it's about almost like can I say the word agentification? Uh you know, it's essentially making an agent take care of all of that. Yeah, yeah.
But I I think it's uh, and you already mentioned it. So if you want yes, it gets used, uh it gets a little bit getting used to as you know, what's the new world as a PM and an engineer, and kind of I don't I feel okay contributing and working in that. But if for example, if there's one advice that I would give a developer right now, moving to a gentic, is just the most important thing would be don't repeat yourself. Like don't keep telling the agent what to do, but write something down.
Uh if you're using a tool to verify what it's doing, give it the tool. So you're not having so kind of that is a very strong, like almost like mindset. Uh and it's different from oh, let's collaborate. Like, not not collaborate, like just have it do all the work.
And and that pressures you in all the correct engineering points. And I think when AI encoding came up, oh, you know, the typical is like uh this is the end of coding and there's no more engineering, and lo and behold, you know, if you really get want to get all the value of that, it is good engineering that we bring in there. So and the engineering becomes creating the the software factory or the assembly line or whatever you want to call it. I can imagine that for a lot of like I can see a bifurcation for developers of the more product-minded ones who like achieving outcomes and getting motivated by I delivered a feature to users they find valuable, they're gonna kind of become more product uh focused.
And the ones who are like, I like making nice, teeth uh tidy, neat code, they're going to be more motivated by I'm going to make the best agentic software uh process that I can. I'm gonna tweak the the dials and the knobs on how the software gets generated. So that kind of engineering mindset doesn't go away. It's just you're not applying it to the code, you're applying it to the machine that makes the code.
Exactly. Yeah, yeah. And I I think that is interesting about um we had that when we went to the cloud uh the cloud, everybody wants to like build their or rebuild their own kernel. Like, why?
Like it's it's good enough. We we don't need everybody to build their own kernel. So if there was one advice that uh you know, people come up to me uh you know over all the years of DevOps and I say, but yeah, you know what, we're special, it doesn't work for us, blah, blah, blah. Guess what?
Like, you know, the industry kind of said, like, that loop is actually universal. It can work everywhere within software development. And so my advice then to the platform team would be saying, like, okay, your number one focus is telling people that like they're not so special in what they're building for code, uh, but it's their mindset that is the difference in kind of how they approach the problems and what they bring into that mix. And it's it's really weird because everybody's repeating this, which is great.
Like every new technology has a learning phase, and we all have to go through the motion and kind of say, hey, oh yeah, now I understand. And and yes, all the stories about like I vibe coded this app, like all great, but how do we bring this into the maturity of an organization? And that's only by saying, like, well, we're gonna probably don't need to, every developer needs to build their own pipeline. That's probably not very effective.
And oh, it's actually the same thing over and over again with two variables or something. And that's kind of the mindset of the platform, the reusability across different teams that goes in. I can share a little bit about our process internally, and I guess like the RD leadership perspective I try and push. Um, the way I think about it is I go back to pre-gen AI and I imagine I had all the manpower and all the time in the world.
Somebody opens a PR, I would want to map out is it high risk, is it low risk? PR is merged. I'd wait two weeks, I want to go over the logs and see did what they do, was it successful, was it not? Is there some edge case they didn't cover?
Uh and basically with that mindset, we say, okay, now you have Gen AI, you have agents. There's no excuse not to do everything you ever thought you'd want somebody to do, and then automate it so it happens by CI automatically. Uh so for example, here, all of our PRs get labeled by risk automatically by an agent. Cool.
Uh then they go to staging, they live in staging during the QA process. Before we release, we have another agent that goes PRPR, is sorted by the risk level, and goes through our logs. Our logs are queryable for agents, and they actually go and check every single PR. Hey, what was the intent?
Hey, what did the logs look like? Did it cover what it was supposed to change or fix? Was there some edge case we missed? And it's been amazing that it's caught more and more stuff.
Like the high-risk PRs often do miss something, even though our engineers are agentic in how they work and everything, and we have QA and we have testing and we put in all this effort. But hey, when you retro it a week after it's in staging, you find more stuff. Uh and often it's like easy stuff. Like the agent knocks it out, it's like, oh, this edge case pops up once and never because we're a non-deterministic product because of Gen AI and LLMs.
But hey, now we see that and let me capture it too. Uh so that sort of mindset of if I had all the manpower in the world, what would I tell them to do? Okay, well, start having agents do it once, twice, automate it. Yeah.
For us, we put it into our release cadence, uh, and suddenly you just get better and better quality. And you move stuff that you would need to think about, or your developers would need to think of, just move it. It's just automatic. I love that you bring up the risk thing because there's the belief that the digital factories, you know, kind of will churn automatically and all the things.
But because when you know what the risk is, uh, you know, again, translating it to a management world, different risk levels require different management techniques. Like, are you micromanagering? Yes, probably the risk is very high, and then you need to do that. If you have guardrails or something that is not that like important, or you have like some way of mitigating this in the product, by all means, then we'll get the feedback later.
So that is the the levels and that you can only do that if you start like thinking about like uh risk levels of like things pushing through. We even couple it to our uh PR fatigue. So high risk, get more human resources to go over it, and we even couple it to our token budget. So low risk, we're not gonna spend a lot of effort to check it from an agentic LM spend.
Uh so just a way to prioritize everything. Sounds a lot like running a business now, right? When you talk about adding this into like you know, instrumenting this through your CI, is CI still fit for purpose in this in this flow, or are there alternatives or or or changes we need to make to our CI style process to make it better for a genetic workflows? Oh, there's like key transformations you have to do.
So one key transformation is whatever your observability platform is, you need to make sure it's easily queryable for an agent so it can verify both while it's doing a task and also retroactively an entire like I don't know, dev environment, staging environment, prod environment, you need to make it really easy for the agent. And we have skills that also tune themselves so we continuously get better, especially as our product evolves and our logs change and the way you query change.
So that's like a key part. Uh RCI is very agentic heavy. Uh even the way you push and open a PR is extremely agentic heavy. Uh and all the actions that we put in from a DevOps perspective are also like really agentic heavy.
Uh that's just a way not to have uh like you want into end, you want unit test, but you want those all those mechanisms for feedback to the agent. So the agent works in higher confidence. Uh but you gotta think of as the agent stepping in, and how are they doing more than just writing code? How are they your confidence layer that there is quality, that the risk is managed, uh, that stuff is moving faster?
We even have agents help us prioritize, like I don't know, you have 20 PRs open, which one is high importance? The agents will tell us and push the reviewer to act on it. And is it do you do you see it almost like as this uh um this you know series of actions that need to happen, or do you feel like it's you know more fluid than that where you have you have different different agents providing you with different data to make you know decisions more integrated, more collaboratively?
Okay, so I'm gonna go with your answer, which is it's fluid until I figure it out. Yeah. And then it's hard-coded because I don't want to repeat it. I want to get that mental load off.
I've I've I've got um a strong hunch that um or strong opinions really on this in that um when when you're a software developer and you're trying to work out, like you're implementing a feature, like one, just make it work, then make it work maintainably, then make it work readably, performantly, securely. You've got all these different lenses through which you have to look at a code change. I think that what we need um with agentic development is to apply all of those lenses deterministically.
All right, instead of hoping that an agent is going to figure out all of this in one pass by sticking all the things in agent's MD. Instead, we can have some kind of CI-like process. I've built a little tool called Assembly Line that does this where you you make a change with your agent or the agent makes a change by itself, and then another agent gets spun up with one deterministically defined prompt of check for dead code, check for missing test coverage, check for this, check for that, check for the other.
And to your point, you've got infinite engineering resource. So why wouldn't you look at a change through all of those lenses and then try and provide feedback as fast as possible to the agent that was doing the main kind of set of changes so it doesn't deviate too far and uh can fold that back in. So I think the like we need to, as an industry or you know, as practitioners, figure out that the flexibility and power of LLMs to work in a non-deterministic fashion is really powerful, but we don't want that all the time.
I know of some people that are like using agent skills to figure out how to deploy stuff into prod, and I'm like, just build a platform. Like we've had platforms for 10 years, they're nice and deterministic and straightforward. Maybe use agents to build your platform, but don't replace your platform with agents burning tokens reinventing the wheel. Like there's a place for the um determinism, and I think it is in that kind of figuring out, like you explore and then you figure out okay, right, that was the right place where we should have run that kind of check or uh a pass or sweep over the code base through this lens.
I I have a ton of ton of say. Okay in a second. So excuse me. Uh so we have two harnesses in my organization.
We have the harness in the product and the harness for the RD team internally. Uh so it's beautiful because what you said is stuff we have in both harnesses, meaning we do not want the agent to go through this route to exactly what we want. We want the agent to do a big step this way, and then the one that step makes sense for it statistically. And then a big step this way, again, it makes much more sense to it statistically.
You see that you're no longer fighting with the LLM. It does not want to do one, two, three steps in one. It wants to do one, wants to do two, wants to do three. You just get better quality that way.
And it's funny you call it assembly line, we call it nudging, uh, but in that sense, you get a very clear workflow uh and you no longer fight the the agent and how it wants to work. Uh what we do internally for our like RD team, the RD harness, not the product, uh, is sadly we have to factor in the human element in in just because we do. So we developed the humans. It's like the slowest part of the whole machine.
It is. Uh so we developed kind of pump, is what I call it, um, which is plan merge polish. The code is moving so fast you can't have an open PR. Meaning, if it's an open PR and I want to do QA on it, and I want to do like a product review, and I want to do a design review, that is all lovely.
By the time those are done, the code is gonna have so much conflict, everything looks different, I have to restart. I I I am gonna be so, so happy if we manage to go back to trunk-based development as an industry. Like pull request-based workflows inside enterprises are a damn silly idea. Um they make great sense in open source repositories with people who are not strategically aligned and you you don't trust, and you know, it makes sense there.
But like that, because of that speed change that you mentioned, like you've got to get the feedback to the agents working on the code as quickly as possible. And that kind of goes back to a point that I think is running through all of this, that the fundamentals of like what made good software delivery, what made good uh methodology, those haven't changed. Things like fast feedback are uh important, they always will be. Having feedback, being able to tell, like you know, the point about observability, that the agent can tell whether it's broken something in production, having all those feedback signals is was important before for humans and continues to be important for agents, if not being even more important because of the rate of change.
So I'll expand 100% agree. Uh the the issue there is also that as soon as a PR is merged, the the landscape changed, the code changed. So the next agent that starts will do a different job. Like everything impacts it, the statistics of the path the LM goes through.
Uh so for us we we called it like the the M in the pump, the merge. As soon as the PR is open, we want to get it merged as fast as possible. So hey, human reviewer, look at it right away. Get it in.
We need it to pass our CI. We do not need it to pass QA. We do not need to pass the product or the designer, everything is feature flagged. Then we go into the last P, which is polish.
Stuff is in the code, every new feature that gets developed is off of the same code. And our product manager can go and review it and they can polish it using our platform and get the UI and UX exactly right and see all the bad decisions from the user-facing perspective that the developer made. Well, the developer got the feature working, it did not get the feature polished. So that's fine.
We have other personas for it. And now they can do their work in a way that is smooth within the evolving code base, like it's evolving really quickly. Uh so that kind of feature flagging, merging quickly with feature flags enables us to do it kind of two-step. There's no one PR.
A feature should take three or four PRs, one PR from the developer, one PR from the product manager that changes functionality, one PR from the designer that tunes the UIUEX, and then maybe one last PR from the QA that saw some sort of itch case. And in terms of like the the human reviewer being the slowdown there, I had a really interesting chat with uh Ryan Lapopoli, who's a member of technical staff at uh OpenAI. And um one of the things that they're doing in some of their projects is to entirely remove uh a human reviewer.
Uh, how much does that make you go, oh gosh, that's horrible? Or how much does that make you think, oh wow, there's a real opportunity here in certain cases for us to actually allow either an AI or tooling to perform those uh reviews and give us that feedback, maybe for different levels of criticality of change. But is that something that you look at, Mike, much? Oh, all the time.
Yeah. So again, it goes back to the risk uh conversation. High risk, no way. Uh I want a human person.
And I also define like risk is also by what part of the code base they're touching, how sensitive it is. Uh but beyond it, I'm trying to automate the PR commenting and comment resolution process, uh, which means I have uh agents that have modeled every single developer on my team and their comments for the last few hundred PRs. And I know each and every persona. So as soon as a PR is open, we have a whole prepare PR, which is a bunch of agents to check it in all the ways that we've modeled it, we need to check it.
Once a PR is open, we launch different agents that model different people from our team and comment as if they were those people, and does a surprisingly good job. Then we trigger agents again to resolve those comments. Then we trigger the retro to see are there any error logs, stuff actually behaves like it wanted to. And again, if the risk is not high from the PR's nature, I'm moving towards being able to like skip the human reviewer at all.
Again, I have developers on my team. I have some developers on my team that are absolutely against this because they want the developer to own. So I'm going through the motion of massaging them and getting them comfortable with the idea. Uh but I'm definitely pushing for that.
And it depends how reversible it is as well, in terms of uh It's the cardrails, right? And it was the same thing, like when everything was being automated during the DevOps transformation, people said, like, but if I did change one thing, I can delete everything now, automate it, right? But it was the harness of test the CR CD that kind of was like the fail system of, and then the narrative was you can have a junior come in, push to your uh you know, kind of your main, and like your harness, your test harness will catch it.
Like if that's a good one. Uh and but what you saw is that there was still a difference between I'm gonna do changes in the code of the compute to I'm gonna do changes on the data scheme model. Because people felt the risk is high. So you you can see kind of similar patterns on like uh the choices that you make.
Uh but it's it's a mindset, like how far do you go? There's also a cost to kind of doing the automation. And you mentioned a lot of things like, yeah, I have an agent for this, for that. Uh, you know, putting everything in place is it's it's currently like a lot of work, right?
Like and yes, you can buy your product or your product or like everybody's product, but that that kind of is uh, you know, funnily enough, we hoped everything was be more simple, five-coded, and everything were working. And what we end up, and that's usual for any technology maturing, is that when it's matured, it's even more complex when you want to reach the next level of kind of like full uh kind of uh you know power of the new technology. So uh one of our customers at Adevo, uh they're doing some really great work where they've got a team that has gone full agentic, they've changed their development process as uh uh uh to um embrace AI in many different ways, requirements, gathering, transcription, story writing, implementation, all of those kind of things, and they are now going 72 times faster.
So they uh have done eight years worth of work uh and they're on track to complete that in 11 months with a third of the people. So they've uh done great things there. They've uh given up on human code review, um, but each individual developer, they take an epic and then they some of them use BMAD, some of them use SpecKit, some of them vibe. Before they self-merge their PR, uh they run about seven or eight uh code reviews, different tools looking at different things, and their confidence is uh they've been doing that, that they're now confident enough that they don't need uh humans to look at it.
And I think there's something interesting, I don't know how true this is. Um I mean, data is definitely like when you screw up data, your data is gone. Do we need to move to event sourcing and having a you know a replayable history of all data transformations? That's one thing.
But like generally, as a trend, we've always seen that trying to prevent bad things happen by putting a check in place is the wrong answer. Like with data and databases, we went from like having transactions to prevent inconsistencies happening. That was slow, it doesn't scale. So, what was the answer?
Eventual consistency, and we'll reconcile. After the fact. With the cloud native transformation, we found out that you know you can't have a thousand microservices and then integration test them all. You just have to deploy them and find out what happens.
So you need great observability and you need to be able to fix forward very quickly. So again, you let the bad thing happen, but you react to it more quickly. I imagine that where we might end up with uh agentic coding is something like outcome-driven development where you've got observability of the business impact you wanted a feature to have. Like, is the application doing the things the application is supposed to do?
Are users able to do what they should be able to do? And then we have software factories chucking features out, and then you look for regressions in like, okay, the number of checkouts uh on the shopping cart has has um uh has dropped and reacting and fixing forward at that point. Is that something you feel users are gonna accept? It would be interesting to see in that like do we all end up getting used to um like buttons moving locations and software being less kind of uh who was it that um coined the term?
Um I think it might have been Steve Yeeg, uh might be misattributing that. Um, but of non-deterministic idempotence of this idea of you have the same spec, same product requirements, you run it through an agent twice, and you get two slightly different things, but you enable the same user behavior. Just as we've got used to waking up and finding out that you know the Apple have decided to change like how the photo album works, and our elderly parents are now very annoyed. Um, do we end up with something like that on a more daily basis, like this rapid churn of um features changing and being slightly different, but like I can still achieve the outcome?
I don't know. Yeah, there was a testing Facebook, for example, like it has so many rules, so many features. Writing that kind of in a consistent test. And the way that they did it was they they added agents on that were using Facebook and they had their own rules.
Like, for example, like that agent should never get a friend, like whatever. Like it should not be listed visible in anything. So they had these kind of like flags out and by like all these kind of things, but they couldn't only make it work when they hooked them in on the production system. And that was a novel way of thinking of instead of the rigid test flow, almost like you track whatever behavior is going.
If there's bad behavior, no, it's not like it's a different way of looking at observability instead of your logs, but kind of that was uh an interesting way. Like a digital twin is also something I I think will have a more future is kind of predicting things, what is gonna happen, and then the risk comes back like, okay, I I have a model, I can run it through, I can see what's happening, and kind of those things. Uh, but yeah. It's funny, I I thought about Facebook too, but from a different perspective.
They have a huge infra setup for A-B testing at scale. So any developer can do whatever they want, and it'll go and get tested, A B A B testing uh facilitated again on scale. And that I think that's kind of what you were going for. How can you do it on scale and just see what wins?
Um which sounds great on paper. A lot of the organizations we talked to don't have the setup for it. Uh and then where we get some organizations that come to us and figure like a few hundreds of organizations that are I think the one I talked to last week was like just shy of a thousand people. And they said that their product is just downward spiral because everybody's like a gen tech development is moving so fast, there's so much PR fatigue.
The stakeholders, like the PMs and the designers and stuff doesn't go through them anymore because everything's moving so fast. And then like uh on the call was like the head of the front end infra, and he says, Yeah, I'm on calls. And I discover new features on the call that don't look anything like our other features, that don't even match our design system, that functionality-wise don't make sense to me. And I find them out with the client live, and then I asked the developer after the call, like, what is this?
He's like, Oh, this got merged two weeks ago. And then I asked, did anybody look at it? Did it was there a product review? If there was like A B testing that at least it'd be guarded in some sense and not just immediately impacting all of the all of their users, and again, huge organizations, it's been around for over a decade.
Uh but they have adopted agentic uh coding really quickly without amazing guardrails. And now they're kind of spiraling back. They're like, wait, wait, we gotta control it. We have a huge brand, we have users that depend on us for years.
I can't move the button and make it hidden because this is a core workflow they need to work through quickly as a part of their business needs. But it's almost like uh it's the you know, it is the UI really going to be the way we you know interact with software these days? An agent will be absolutely fine with it. Uh it's us, it's again, pesky humans, which is the problem.
And how much of this will actually change, and you actually think, no, I don't I don't need a UI on screen. I need to ask certain things, or I just need my agent to go ahead and interact in in a specific way. Um so I don't know. It's it it if I I wonder if in years buttons and things like that will just be replaced with completely different ways of us interacting with apps through.
Yeah, but we'll still be a visual-oriented creature as well, right? You we can't just say we're just typing in the words, like you know that uh there was a talk here as well from TL Draw, kind of interesting that like you know, having a your agent giving it a whiteboard, because sometimes you can express things easily in the whiteboard and like typing and and that kind of fluidity is also, I think, in the in the UX. Like sometimes it is gonna be. I I just want to ask a question, you reply back to me.
Sometimes just show me, like explain it to me. But it might be interactive and kind of going through that. I think that's also the pendulum, which was funny, like IDE, CLI, and now we're back at orchestrator review tools because we need something, like we need to see what's happening, and we can't just like do that. But the agent might crack along in their own communication and find the best way of kind of dealing internally.
But as long as we need to make a call uh shots on kind of those uh kind of risks, we need to be informed. And somebody explains it to me like what are we doing here? Uh and that's a challenge if we don't do it ourselves anymore. So we've talked a lot about good patterns, about ways in which people should should change the way they develop software, um, and to essentially directing people in groups as to how they should change.
But what is it that organizations, uh, enterprises are doing today as a part of their migration, part of their adoption, that is actually just a mistake. It's wrong, it's a it's a it's been proven that you know oh dragons live here. What should they what should they not do that they're doing today? I think I'd I'd give them the advice is that there's the still the tendency of like if we're doing a content constant conversation with our agents, we're in control.
I think there needs to be um for them to get into the next mile, they have to let go. But they it's it's almost like you have to mentor the other one instead of doing the work. And that's kind of the mental challenge that I think a lot right have right now. Like um stop doing the work yourself.
And it's it's really weird if that's what you've been doing for so many years. And what a lot of developers actually hold uh important identity cells because it's what make it's what they feel sometimes actually makes them the expert, right? Yeah, correct. It's interesting.
So I have a few different answers to this. Uh the first one is you want you want to control your AI spend. People are are releasing crazy budgets and everybody's trying to token max every single developer. Uh and if you I think the right way is give the developers the task that interests them, you will see they're way more engaged with the agent in the session where they're building it, versus give them the task that the developer's not interested, and then they will offsource all that mentor load to the agent.
Uh one, you get poor quality PRs, but two, you get way bigger token spent. Like imagine that something's out, the the disinterested developer will just say agent test it instead of testing themselves and seeing exactly what's wrong and saying, oh, this is the issue, let's fix it. They will let the agent do seven sessions of QA and the token count for the session goes crazy because the developer is not interested. And then, so that's one part.
Like, yes, people can move faster, giving them tasks they don't care about is a great way to spend your AI budget really quickly. That's one. The other one that we've seen a lot of organizations do is they want to jump in headfirst to AI native. Uh and they do that by giving cloud code to everybody.
And the cloud code has code in the name. It is great for developers. I love it. It is most definitely not good uh for non-developers, and if you give them that, you will get PRs.
One, they don't know what they're doing, they don't know how to set up an environment, and they're gonna drive your dev team crazy asking questions, even if they feel comfortable enough asking questions, or they're just not doing any meaningful work. But if they do open PRs, we see like garbage in, garbage out. Uh PRs that the VP we're on calls with like CTO and VP product, and the VP product shows off to us saying, Hey, I opened a PR yesterday, and then the CTO sitting next to him says, Yes, it was crap.
Um that also a great way to not go about it. Uh so when you do enable them, you want to get the tool that is really optimized for that user. Again, we optimize for PMs and non-developers, but giving the people the tool that's not meant for them is not the answer. Uh again, it you use your your AI budget really quickly, and you get frustrated people all around.
The PR fatigue is there on the developer side, and PMs that open really poor PRs are increasing the PR fatigue, they're getting more frustrated themselves, and you get Uber scenarios where everybody's just disappointed. I think uh in terms of what people should not do or stop doing, like the whole uh piecemeal approach to last year, 2025, was very much the year of CTOs going, well, I let people use their own tools and whatever they're comfortable with, like that is going to end up being inconsistent and you're not gonna get the organizational learning uh that you need.
So affirm mandates that like the AI bus is leaving, please get on board. Um, that's been shown in research by multitudes, I think, to uh be uh a strong uh indicator of uh success of a genetic coding adoption. Um the the counterpoint to the um token maxing is also don't go to the other extreme and set a really low limit. Um, you know, I've I've seen places where there's a hundred euros a month uh tokens uh limit, and like 100 euros you can't really do very much.
So it ends up adding cognitive load of people like, oh, should I should I use Cursor for this or should I is this the most important thing I'm gonna do this month? Um so having something in between um would uh would also be sensible. And then maybe rethinking what a test is, but you know, to invoke Beyonce, if you liked it, then you should have put a test on it. Um if you care about a thing, then you need to have something in place to assert that.
Whether that's functional tests, whether it's a uh a sweep of an agent with a particular lens looking at a concern, whether it is connecting up via MCP or agent to observability so it can figure out the thing that it just uh built and got deployed is uh now breaking. Um, making sure that there are those feedback loops in place. If you just have an agent on its own and you expect it to get everything right first time without being able to even perceive the mistakes it might make, then you're gonna get bad results.
So um making sure that you think about what you value in your code, um, that it's architecturally uh you know uh sound and um the code quality is high and the test coverage is good. You need to work out what you value and then make sure that there is something in place to prove that that is there. So it's not quite testing, but it's like thinking about the the non-functional aspects of your code as if they were tests, if that makes sense. Okay, expand slightly.
I love that point. Uh I think something that's really important is also the temporal mindset and the context hinting mindset, which means every test needs to stand the test of time, needs to be evergreen, but you also need to think of it as an opportunity for a context hint. Meaning when the test something fails, the test has enough context in whatever the assertion is, uh the comment, the error response, whatever it may be, so that the agent that gets it will know how to act.
So lazy developers just say, cool, I covered it with tests. No, like you need to cover it with tests, and you need the test to have the context hints that the future agent that will see them will know what to do with it. Like that's been really impactful in getting the agents to actually manage in sophisticated code bases. Yeah, and it's exactly the same as we were uh all saying earlier of like that's always been good development practice.
Like it's nothing worse than as a developer, you run the tests and you just get like assert true failed, and you're like, great. Or 1317. Yeah, well, what do I do with this now? So yeah, it's uh important stuff.
And thank you for quoting Beyoncé, because uh I've been trying to work out how I increase that in this podcast, but but finally someone has. So thank you very much. Um, well, we've talked about a lot, and it's been super, super valuable to have you all kind of like discussing on these points. So I appreciate appreciate all your input on that.
Um, thank you very much for joining us at AI Native DevCon. Why don't you very briefly uh talk a little bit about some of the sessions that you're each giving? Patrick, why don't you start? Yeah.
Um I'm talking about the different layers from up to a dev to a team to uh a V a platform team and a VP. What's the mindset that you need to have to go into this new agentic world? I'm not talking about generic transformation, like do a hackathon and so on, but what is really like how do you embrace the new thing? So that's my talk.
Awesome. Uh so my talk is uh about what it's like to become more of an AI native organization, where your PMs, your non-developers are actually opening PRs, and how really the metric you should be copying, actually should be tracking is the merge rate. So the the PRs that they open, don't say how many PRs do you open. Look at what the merge rate is.
That's how you will know are they opening quality stuff or are they just increasing PR for T. Yeah. Uh I was uh co-presenting with Tomas Mai uh from Adevo, and we were talking about the story of how we upskilled 120 of their developers in agentic coding. And the most important thing about that talk is you get to see me in a day glow fluorescent uh tracksuit.
So you should people should definitely check that. Beyonce end this, really is there a beyond is there a Beyoncé uh quote in that one? I don't think there is Shakira, maybe? Sadly not.
Alright, alright. Um so uh why don't we uh if if well all of these talks are recorded, so you're very, very welcome to go and we'll we'll add some we'll add some links there in the show notes. If I or my our listeners can only go and watch one talk, which one would you agree that they should No, no, let's not let's not do that. Uh all of them, all of them.
Um thank you very, very much. Really appreciate you being here and uh sharing the time with us. Thank you very much. Thanks for watching.
Thank you. I hope you enjoyed that. That was really great from my point of view. Uh would love uh to hear any feedback on that.
Please let us know. Podcast at Tesla.io and uh tune in to the next episode. Bye for now.
Hope you had as much fun listening to this episode as I had recording it. And if you want to hear more from the guests featured in this episode, we have their full talks linked in the description below. So that's it for this episode. See you next time.
The AI Native Dev is brought to you by Tesla, the package manager for Skills and Context. Your hosts are Guy Pajani and me, Simon Maple. Our producer is Tom Dowler. The AI Native Dev is not just a podcast, it's a community.
And we host a monthly meetup at the Tesla offices in Temple Land. Visit Tesla.io for community to learn more of iOS.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.