The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Leadership/Tech Leaders Intel
Tech Leaders Intel artwork

Measuring What Matters: Unlocking Engineering Productivity with AI and Metrics

Tech Leaders Intel · 2024-12-19 · 51 min

0:00--:--

Key moments - from our scoring

Substance score

45 / 100

Five dimensions, 20 points each

Insight Density10 / 20
Originality7 / 20
Guest Caliber11 / 20
Specificity & Evidence10 / 20
Conversational Craft7 / 20

Engineering productivity measurement has become mainstream as teams work across distributed locations and cloud infrastructure makes data collection easier, but poorly applied metrics create cultural resistance and team dysfunction. Gary Stevens, Director of Engineering at Trainline, and Victor Clark from Xebia explore how organizations can implement metrics frameworks like DORA (lead time, deployment frequency, failure rate, time to recovery) without falling into the trap of comparing individual contributors or imposing arbitrary targets. The key insight is that metrics should be hypothesis-driven and team-owned rather than management-imposed - teams decide which sub-metrics they control while working toward hero metrics tied to business outcomes. Both emphasize the pitfall of Goodhart's Law (when metrics become targets, they cease being effective), the importance of flow analysis across teams (not just individual throughput), and how tools that anonymize individual contributions at the team level prevent misuse. The conversation addresses the maturity gap many organizations face after adopting DORA: reading the high-performer playbook and attempting to skip steps without understanding their own practices, people, and process constraints.

Key takeaways

  • →Metrics should be team-owned and hypothesis-driven to avoid gaming, with DORA metrics (lead time, deployment frequency, failure rate, time to recovery) serving as a high-level foundation that teams build sub-metrics around.
  • →Measuring flow and wait times across teams reveals collaboration bottlenecks and prevents the anti-pattern of adding headcount to late projects, which research from Fred Brooks shows only makes projects later.
  • →Individual contributor attribution from git logs or issue trackers should be avoided; instead, restructure high-value bottleneck developers across teams to increase overall flow and reduce their cognitive load.
  • →Cultural resistance to metrics dissolves when teams understand why they're measured at the collective level, not the individual, and when metrics are clearly tied to business and user outcomes rather than arbitrary process targets.
  • →Maturity in productivity measurement requires understanding your own practices and constraints before copying high-performer playbooks - many organizations struggle with the gap between DORA awareness and actual implementation of deployment frequency, observability, and continuous delivery practices.

Guests

Gary StevensVictor Clark

Topics in this episode

DORA metricsContinuous DeliveryGoodhart's LawDeveloper productivityDeployment frequencylead time to valueengineering metricsflow analysisTrainlineXebia

Questions this episode answers

What are the four DORA metrics and why do they matter for team productivity?

The four DORA metrics are lead time to value, deployment frequency, failure rate, and time to recovery. They measure the flow of work through a team and its impact on reliability, serving as a high-level foundation that helps teams understand whether their development practices are improving both speed and stability.

How do you prevent teams from gaming productivity metrics?

Teams should help define their own metrics based on hypotheses they believe in and can influence, rather than having metrics imposed by management. Goodhart's Law states that once a metric becomes a target, it ceases to be effective, so metrics work best when teams own the reasoning behind them and understand the narrative, not just the numbers.

Should you measure individual developer productivity or team productivity?

Measure at the team level and use tools that anonymize individual contributors; teams own services and projects collectively. Individual attribution from git repositories or issue trackers enables discrimination and gaming, while team-level metrics focused on flow and collaboration reveal actual bottlenecks and opportunities for restructuring.

What's the maturity trap most organizations face after learning about DORA metrics?

Organizations read about high-performer practices and try to skip steps by simply deploying more frequently or increasing observability, without first understanding their own constraints around people, process, and technology. True maturity requires working through your own practices before copying the playbook.

How does measuring flow across teams differ from measuring individual team output?

Flow metrics track wait times and handoffs between platform teams, architecture, and frontend teams to reveal where collaboration breaks down. This prevents comparing team A versus team B output and instead shows why adding more people won't solve blocked dependencies - a key insight from Fred Brooks' research.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

10 / 20

There are occasional worthwhile observations - the 80/20 commit distribution finding, the client-as-bottleneck inversion, and the test-coverage-vs-defect correlation - but the episode is padded with lengthy, meandering exchanges and restates obvious agile/DevOps wisdom for most of its runtime. The ratio of novel ideas to filler is low for a 51-minute episode.

more than 80% of the code commits being done by only 20% of the developers
the majority of the wait time was actually due to the client and not due to the supplier

Originality

7 / 20

The episode largely recycles well-worn frameworks - DORA, Goodhart's law, Brooks' law, OKRs/GQM - without adding new angles or contrarian positions. The advice to 'start with the end in mind,' avoid comparing teams, and keep metrics small is standard consulting boilerplate with no genuinely fresh twist.

Goodhart's law has always said that once a metric becomes a target, it ceases to be an effective metric
Happy teams are more productive by their very nature

Guest Caliber

11 / 20

Gary Stevens is a credible practitioner as Director of Engineering at a real scale company (Trainline), with relevant prior experience at Compare the Market; Victor Clark is a competent consulting practitioner from Xebia with real client case data. Neither is a particularly senior operator or widely recognised expert, and Victor's consulting background tilts toward advisory rather than operator experience.

I'm Gary Stevens, I'm director of engineering at uh, Trainline. We're a European, uh, rail and travel booking, uh, platform
we did a thorough analysis of 100 developer program, uh, which was uh, well, not delivering as expected

Specificity & Evidence

10 / 20

The episode offers a handful of concrete data points - the 80/20 commit finding, build times 'hours to minutes,' and a 10% speed impact example - but most evidence is hypothetical, vague, or unnamed ('a recent assessment,' 'lots of clients'), and the AI discussion at the end is completely unsubstantiated.

more than 80% of the code commits being done only by 20% of the developers
somewhere in the region of a 10% impact on our speed to delivery

Conversational Craft

7 / 20

The host asks functional but soft, open-ended questions and never genuinely challenges either guest; affirmations like 'that was amazing' and 'that makes a lot of sense' dominate his contributions. The AI segment is explicitly rushed and superficial, and follow-up questions rarely drill into specifics or push back on unchallenged claims.

That was amazing. I don't know whether you guys realized quite how many metrics you just threw into those last statements
That makes a lot of sense

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A44%
  • Speaker B42%
  • Speaker C14%

Most-used words

teams58metrics58team50productivity31data27different22measure19important18focus17saying17developer17experience16impact16value15process15help14

Episode notes

In this episode of Tech Leaders Intel, we hosted two industry experts to discuss one of the most critical and often debated topics in software engineering: productivity and performance . Our guests are Gary Stevens , Director of Engineering at Trainline , and Viktor Clerc , an engineering leade r at Xebia, both offering insights into what it takes to build high-performing engineering teams . Gary shares his experience in balancing team motivation and productivity at Trainline, focusing on customer lifetime value and efficient software delivery. He highlights the complexity of tracking engineering productivity, especially in today’s distributed, remote-first work environments, where effective measurement can no longer rely on simple observation. Viktor expands on this by emphasizing the importance of optimizing engineering practices and using strategic assessments to overcome client-supplier challenges. His focus is on creating alignment between technology and business goals, ensuring that teams are empowered to deliver not just products, but value.

Full transcript

51 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: M welcome to Tech Leaders intel, where we discuss technology leadership, industry challenges, winning strategies and much more. If you're interested in hearing insightful tips

Speaker B: and gaining knowledge from experienced technology execs

Speaker A: that will really help drive success in your own company, then you're certainly in a right place.

Speaker C: Hello everybody and um, welcome to the latest edition of the Tech Leaders intel podcast. And today's episode focuses in on engineering productivity and performance. Team productivity is a complex and sometimes controversial subject. So I'm delighted to be joined by two guests experienced in building high performing teams and software. First off, I'm pleased to introduce Gary Stevens, Director of Engineering at Trainline, responsible for delivering on key strategic initiatives with teams across locations in Europe. Hi Gary.

Speaker B: Hi Mark. Glad to be here.

Speaker C: Thanks a lot for joining us. And um, Victor Clark, one of the team here at Xebia who's been focused on building high quality engineering teams, optimizing practices and ensuring a focus on delivering value and desired outcomes. Hi Victor. Uh, hi.

Speaker A: Uh, it's a pleasure to be here, Mark.

Speaker C: Great to have you on board. So Gary, first of all, could you tell us a little bit about yourself, your background and why this is such an important area for you?

Speaker B: Sure thing. So, uh, as you said, I'm Gary Stevens, I'm director of engineering at uh, Trainline. We're a European, uh, rail and travel booking, uh, platform. I look after teams focused around our uh, customer lifetime value, um, acquiring customers, monetizing and retaining customers as well as teams across our global platform on search, real time and traveler care. Prior to that, I was at a, uh, insurance price comparison website in the UK called Compare the Market. And the commonality between both roles has been making sure teams have got clarity around what they're doing, making sure that they are motivated by the area that they're working in and underneath all of that, measuring that their productivity is being used effectively and they're working on the right things in the right way.

Speaker C: Thanks, Gary. Victor, could you take a bit of time, tell us about yourself and I guess how you've supported customers to optimize the performance of their teams and engineering practices.

Speaker A: Yeah, thanks a lot, Mark. So, uh, in my history I've been working with different clients in different perspectives, different domains, uh, originating from a technical background. I supported them in improving their delivery, their productivity, not, uh, just from the context of, hey, are you building the best possible product, but actually is product delivery and uh, customer satisfaction actually increasing? Right. So with a lot of knowledge also within Xebia on how to uh, create uh, the best possible software, the Best possible technology. It all comes together in getting everything uh, in line to deliver that value. So the combination of focusing on how teams are using the technology, how technology is adding value, that's something I focus on. And in that uh, I engaged in coaching, uh, activities with clients, but also in more strategic software assessments where there's really uh, well it's on the edge, there's a tense client, uh, supplier relationship and actually we need to see how we can from an impartial perspective give a boost. And throughout the years I've collected wealth of experience saying, well hey, um, let's focus on what matters. And the topic of productivity, the topic of really getting the value out is something that is uh, ah, a red line in all my experience.

Speaker C: Thanks Victor. Uh, and I think that sets us up really, really nicely for a conversation this afternoon and I'm looking forward to diving into the detail. So just to tee it off, um, as the saying goes, I think most companies are becoming, well in varying degrees software companies and therefore engineering or product development is a critical function that's at the core of the business. I think that the traditional view is that measuring development or engineering productivity is hard and um, there's not many qualified to effectively measure it. But you can't improve what you don't measure. So it'd be really good to start with how you see the origins of productivity and performance, what it means really, um, and why now in particular it's so important. Gary, do you fancy kicking that one off?

Speaker B: Sure thing. So I think you've hit the nail on the head. Are saying it's a difficult and complex area. Um, productivity in other fields is something that people have talked about for a very long time. And I think as software engineering has matured and as uh, the whole space has grown in scale and size, um, we've seen the trend towards wanting to know are we focusing on the right things. And many teams just want to make sure that the thing that they are building and the way that they're doing it is optimized. And in many ways that's just an evolution of how Agile has come about and focusing on the best practices. Um, and now we live in a world where teams work remotely or over multiple locations or in person. And so it's much harder to just observe and monitor visually the kind of, the outcome, um, in terms of kind of like why now? I think there's just more tooling and more uh, optionality available. We've seen a proliferation of data platforms, a proliferation of more cloud based services where we do and store most of our work. And so it's easier for us to look at and track stats against a number of different things. And I think that's made people curious about some of this information, how it gets aggregated and how we form better tools and better opinions around what we do. Um, so it's a hugely exciting area. It's an area that's grown in size and complexity and opinions over the last couple of years. But it now feels like it's definitely mainstream, um, and something that most engineers in one way or another want to focus, want to focus on it. I think the difficult thing with the subject is, is what are the right metrics? Are there any right metrics? What is productivity? And, and I think it's, it's important to kind of be open minded around that and shape your productivity goals and your productivity philosophy around your organization, how you work, how your teams work and what, what it is you're trying to improve.

Speaker C: Yeah, that makes a lot of sense.

Speaker A: Yeah, quickly pitching in there Gary. I think that that's really important. Of course at some point you get what you measure right and we've seen numerous examples where we tend off to measure, measure the order, let's say the suboptimal things. I wouldn't say the wrong things immediately, but the things that actually are not really driving into the right direction. If I uh, take a stroll back in my memory, I recall uh, perspectives from uh, the Software Engineering Institute at Carnegie Meyer University who came up with a, a team software process, uh, perspective and also a personal software process perspective really focusing on the individual developer and the team performance. And there's been lots of uh, data out there on measurement. And then I believe it was, was Tom the Markov already decades back saying hey, uh, you cannot really fix the difference between a really high end programmer and a mediocre programmer with any method or any tools. It's also about selecting, let's say the right team. Right? So some perspective. If you just get the right um, um, engineers in an area in the right engineering team and give them freedom and autonomy and space to really start excelling, then I'm sure you will achieve better outcomes with regardless whatever process you fix on top of it. And this is where it becomes a bit well, blurry so to say. Because how do you then measure how you then objectify, uh, and discriminate between the better performing teams and the lesser performing teams?

Speaker B: I think you make a really good point there about you can't use this, you can't use any measurement as a proxy to say that is how all teams should work or that's what works for everybody at the end of the day. Teams solve different problems, they have different code bases, they have different challenges, they have different ways of working. And none of this should be about mandatory enforcing a particular way of working for any, for any one team, um, what they can be as a guide of things that have worked in particular scenarios or things that do add value or do help. And I think if we kind of look at it through that lens we gain a richer playbook of ways to solve interesting problems um, and figure out what a potential good outcome might look like and how we can move around common problems.

Speaker A: Yeah, and this is where of course um, the experiences we tend to nail uh down in some practices, let's call them test automation, continuous delivery, uh, using of agile techniques to really get a bit of the variance out and ensuring first time. Right. And also ensuring that the well, the, well the key heroes in such a delivery team can really focus on the value adding stuff. Right. And not just testing over and over again or uh, engaging in deployment difficulties. But this is kind of the, the stuff we should factor out of the equation and then you really get uh, to the speed where you can excel and deliver value because I think that most of the metrics, and I'm sure we're going to talk about that later, um, there's internally focused metrics which is, which is, which is fine but this is just showing you how smooth your machine is running and more externally focused or outcome focused metrics that focus on okay, and what does it result in. And I think way and way more we should focus on the latter than on the former. The former are of course necessary for the machine and the delivery itself. If you were to compare software delivery with conveyor belt optimization because it's also about the creativity and where the mindset can go of uh, developers in the interaction with the business. That's where we should uh, should focus on. Of course there's advances, my point being there's advances from the last couple of decades. Let's factor those out using a ah, smooth process and experience practices and then we can, then we can really get to it to the, the creativity to uh, add value.

Speaker C: So this all seems really good and really reasonable. So why do metrics have the potential to be so controversial?

Speaker B: I think any data without being properly qualified or um, contextualized can be taken out of context and the data itself doesn't give a view on whether something is good or bad. You know, it's unopinionated it's the number that it is. But I think it can be useful for teams to use as their reasoning of why a particular thing might be happening or why they're seeing a particular result. And with any data, ah, that is collected and looked at out of context, I think it can often lead to a feeling that all we're doing is kind of whittling down the narrative of what the team is doing to those numbers. Um, and that's certainly a message that I always try to avoid, always try to avoid the fact that and clarify why you're collecting it and why you care about it. In my experience, that's been to make teams care about their own individual numbers and not to use this as a comparison and clearly message that the data is not being looked at as a set of, uh, anonymized inputs and outputs that are used to compare team A versus Team B. They exist for team A to understand how team A is performing. Um, and I think if you do that and you help teams understand the way in which the metrics are gathered, that can be a little bit more empowering for the people using it. And I think it's also essential to look at what metrics you are setting within a team. And to some degree you're going to want some normalized baseline that starts as a hypothesis. I'm going to look at lead time to value. I'm going to look at, um, failure rates or deployment frequency. Um, but that's kind of your HERO metric, right? That's your end output. A team can decide what they believe the sub metrics off the back of that that tell the more complete story for them and their development. So they may choose to look at a completely different set of measures that they feel they have more control over. But improving those individual measures, they can build hypotheses and formulate, uh, thinking around how changing those smaller submetrics may lead to a bigger, uh, narrative, ah, around how they're moving the HERO metric. Um, and it's just a way that I think of really empowering people to care more about their own processes, measure the things that matter to them and look more at how they have more control and awareness of what's happening within their team.

Speaker A: Yeah, I second that. And at the same time I've come across lots of, uh, examples at my clientele on, um, hey, there's this engineering manager who's responsible for a bunch of teams and he feels like, hey, we're not getting the most out of those teams. Either those teams in isolation or the Collaboration across teams, right. So um, uh, it's kind of tempting to start feeding that comparison. Right? So why is one team delivering more, let's say output per sprint than another team? And then if you cut away all the context you typically tend to compare well John Doe of team A versus John Doe of team B and then saying well hm, he's better. And then to kind of prevent that is really uh, important to aside from the individual team based metrics to also focus on flow and an end to end perspective. Right. So in the collaboration of a, well platform teams or when um, they say relationship with architecture, front end teams or however the development department is structured, in the end the combination of that is what delivers the value. So uh, uh, what we do quite often is also analyzing the flow and the wait times and the inefficiencies within the teams but also across the teams to really spot, spot the difficulties. Right. If you talk about lead time then it's about well lead time is also wait time. So if you know the word the wait time, you know where you can improve the lead time. And then um, uh, raising also senior management saying well hey, uh, well I think it was Fred Brooks already in the 70s saying well if you add more people to a late project it will only make it later. And the tendency of them is well hey, we're stuck here, we need to have more output so we'll just add more people. And this is something I'm fighting against. And then the data we collect by also kind of extracting metrics, many of my clients are not even uh, collecting themselves can show us why this is data that shows you that adding more people here will not solve the problem. Such kind of also using metrics more in an educational perspective saying well hey, even if you're not not measuring again, uh, already starting measuring these kind of things, it can help you improve flow not by comparing the teams of really getting on towards the output. I think that's really important. And to be honest it's a bit shocking how immature the field is on this area still. So there's lots of work to do

Speaker C: here now that's understood and maturity is something we can come onto in a second. So uh, what I picked up from there is first of all not misusing the metrics, making sure they're applied in the right way and they're the right metrics, but also making sure you address any cultural resistance around that as well. And like you said, it's a balance of that quantitative and qualitative data as well. It's really important.

Speaker B: Yeah, absolutely. I think that is really important. You talked there about not uh, addressing the cultural element. And there are many platforms now that are in this space that can provide analysis and trend data and aggregate this information for you. Um, in my experience, tools that de anonymize the individual contributors in that process is a great way of, at proving that you're not manag, you're not measuring specific individuals. We instead choose to focus at a team level and a team is made conceptually of the services that team owns and the project boards that are ah, contributing data to it. So we can't drill into that platform and see um, engineer A, B or C did more or less. We see the output of the collective.

Speaker A: Exactly. On the other hand, there are examples where we did a thorough analysis of 100 developer program, uh, which was uh, well, not delivering as expected. And of course if you just mine your git repositories, mine your issue management systems and you will find all kinds of data that can be attributed to individuals, uh, and then it could be very easy to discriminate the more productive developer from the less productive developer. Uh, and then it's up to, well in this case ZBIA to say well hey, we're not going to do that. We're not going to feed that narrative, but rather that we say well hey, these are some of the key players in your organization right now we see that they are bottlenecks. So what about restructuring them, uh, assigning them to different teams to increase the flow. And then it's better for the engineering managers that want to deliver more output, but it's also better for the people in question here that are often blocked and often uh, uh, asked 20 times a day to help out and actually also want to become less of a reliance, uh, uh, relied upon. Um, so that's kind of the way to twist it. Not in a political sense, but really in the sense that it really matters to the stuff we try to achieve. There's lots of data because it's just bits and bytes. You can really attribute to what every individual is contributing. But really we need to get the way of that as quickly as possible.

Speaker C: Let's touch in on that maturity just a little bit. Um, because obviously there are varying degrees here of the types of data you can collect from basic tracking all the way through to, you know, optimization through advanced analytics and AI. And I think that maturity curve is really important. Gary, we'd spoken about this a little bit before. Do you want to talk a little bit about how you take organizations for A maturity curve or.

Speaker B: Yeah. So, you know, a few years back now, um, the DevOps Research and Assessment Group published the first of the Dora White Papers. Um, and I think that was

Speaker A: a

Speaker B: real watershed moment, I think, in taking productivity, the concept and conversation of productivity, and very much simplifying it on things that lead time to value failure rates, change rates and time to recovery. And I think that the power that had was that it was at a high enough level that people understood that what we were trying to prove was the flow, as Victor said of work through a team and the impact it had on reliability or availability. Um, and for many organizations that fit right, it was culturally relevant enough to say, I understand that this is about helping, improving the products. I think what it's been also really great at, uh, is opening the door to get people to think more about how their observability or health metrics can be used at a team to understand more. And that's been where I think we've been able to. The space has grown a little bit more and new platforms or new frameworks, whether it's space, DORA or other frameworks that can be used to do this, have had an opportunity to kind of build on that and stand on the shoulder of giants, um, in the space. And now that we've kind of ushered in that wave of thinking, teams are now much more able to focus on more granular, more specific things. Whether meetings have a big impact in flow state or whether, um, ways of working issues or traditional things. The outputs you might get from JIRA or other management tools can all be, uh, submetrics and other aggregation methods that inform the overall hero metrics of what you're trying to prove. And so I think we'll continue to see new ideas and new thinking in the space, um, and that will ultimately just kind of build, build upon the foundations that have been laid and teams will be able to find the right balance of fit for them. Um, but I think all of it has come back to just trying to position the benefit to either, uh, the business, the teams or the end user. Uh, um, has been, has been really useful at opening, uh, the door.

Speaker A: Yeah, yeah, I agree. In pitching in on that topical maturity, what I've seen quite a lot is that, uh, well, of course once the DORA report came out, we see the high performance and we know how to discriminate them from the low performers. So what do the low performers do? They read on what the high performance do and start kind of skipping a couple of chapters and also doing and kind of mimicking that and then uh, well of course uh, you should increase your observability, you should increase your deployment frequency. Okay, let's just deploy more often and then we're there. Right? That's of course not the case. So then uh, there's this kind of pendulum switch back to saying okay, what is it that we need to understand of our own practices, either individually but also at a team level, how to really collaborate and get a combination of the people, the process and technology to get to uh, more frequent deployments. And uh, that's where one of the bigger struggles have come in the last, I think, well, last eight to five years. And right now we see also because the, well, the productivity landscape is exploding, uh, uh, how teams and the developers can really use uh, more suitable technology from scratching from the shelf, that right now we're in a stage where the maturity can only further improve. But this first gap after the hype of the door, I was there, um, I've seen many teams struggle with that.

Speaker B: And one of the big things I think I often hear, uh, initially when talking about this with people is gaming the numbers. And it's something you have to be very honest about and open to it and talk about. Um, Goodhart's law has always said that once a metric becomes a target, it ceases to be an effective metric. Um, and so it really does kind of go back to getting the team to understand why the metrics they've picked to measure in their own productivity need to be metrics that they have confidence in and needs to be metrics that they believe they can influence both positively and are a good proxy for proving out the hypothesis they may have. So they may choose to game them and as you said Victor, deploy more frequently or. Well, I'll just break my stories down smaller uh, to prove that I'm delivering more story points. Um, and in some cases maybe those things will have a positive impact at some point in the world that you're doing. But ultimately if people are opting to game metrics, then perhaps the right metrics aren't being looked at and maybe the two haven't set the right, the right measure of what their productivity might be

Speaker A: or they've been imposed and then uh, they're trying to be gamed. If it's really intrinsically defined then, then it's okay. But there's, there's still quite some examples where management thinks about uh, selecting the right metrics and then uh, imposing them and either ranking teams. The thing we just discussed, which is anti pattern by, by heart, it's difficult to get it out because it also gives in a fluid process. It does give management some false sense of control. Right. Just to measure stuff, uh, on uh, a team in metrics, it's a pitfall.

Speaker B: What we're seeing a lot at the M moment I think, is the growing overlap between productivity and developer experience. And in using more qualitative measures like engagement surveys or feedback surveys, you can validate whether or not what the numbers are telling you and what the outputs are telling you marry up to, whether or not that has an impact in a perceived sense of productivity. And are they having a positive effect on how engineers feel about the work that they're doing? And does it make their lives easier? Does it make the experience of both of the products easier? Or imposing a bunch of metrics and measures, are you negatively impacting the perception that people have around their work? So it's important again to contextualize the message, to bring narrative to what your data is saying and have some way about proving what the real world impact might be.

Speaker C: That makes a lot of sense. And I think how metrics are applied often has a direct impact on uh, how and when people trying to game those metrics as well. I really want to focus in on that developer experience, um, aspect as well a little bit further on, uh, because I think that's really, really important in terms of this conversation. What it'd be great to understand a little bit more is the right way or the way that metrics are applied as well. You know, looking at the individual versus the health of the whole end to end process and then also a little bit further down there as well. You know, as we mentioned at the start, all businesses being technology business, now how is that data used or how can that data be used to make investment decisions? So really how those metrics are applied are really critical to productivity, but also to the performance of the organization and where investments are made within teams or an organizational level as well. So could we just touch in on those? I know that was quite broad, maybe starting with the individual versus the health of the end to end process.

Speaker B: Yeah, well, I think like all good data points, they can help tell a story. And Victor talked earlier about teams maybe feeling blocked or maybe feeling um, that they're not able to measure the input they want to have or their flow state is interrupted. So having a grounding and a proof point to make that point can at least give engineers some feedback around whether or not their direct interaction is having the desired effect or whether or not they have more uh, validity and Areas to say hey, the blockage in this project or interacting with this service or this team is having somewhere in the region of a 10% impact on our speed to delivery. So it helps them sort of m feel more certain and confident in their own uh, effort and impact and helps build the bridge to those conversations about how things may be improved. To your point on investment, it's great I think when you sit back from a team or across several teams and when you can identify a common, a common pain point across numerous things, whether that's CI, uh, CD times, whether that's PR processes across teams, whether that's uh, deployment speed. Having that data and knowing where technical effort or commercial investment needs to be applied to remove a common bottleneck gives you a sense of you can predict the business benefit from that. Having a theory that removing build times, reducing build times, uh, from hours to minutes, you can quantify the impact in productivity that that will have measured by your maybe your developer engagement surveys then saying people feel better about their work. You also then have a positive impact on your retention and your development of your people as well. So there can be some very easy ways to see where your biggest bets lie and where you efficiently direct your resources or commercial investment or re platforming effort to unlock some of those things.

Speaker A: Yeah, uh, fully. Great. I can only just listen pitching perhaps a ah number of other examples where we've seen this as well. So where to get the most uh, uh, bang for the buck. Right. So where to put the investment? Uh, I just recall a assessment recently completed uh where we've seen um, more than 80% of the code commits being done by only 20% of the developers which is well if you're, if you're happy or not with that doesn't uh, matter. But if you want to get more out of your delivery organization typically you should ask uh, confronting questions about what are the other 80% of developers doing? Well not so much with the code. They might be really fluent in and getting energy going and collaboration but it is something you should put on the table. And if you would add more people to such a project they might end up in the 80% doing not that much. And you think well so what's happening? So getting this real end to end flow uh, is important. And then in this case we also interviewed some of the individuals and some of those 20% key individuals doing the majority of contributions. They say well it's great but I also like to switch the page and do something different or add more value in a different way. So it actually Pretty easy win, win situation. Nobody wants to be uh, too heavily relied upon um, for the sake of flow. Um, one other example is that uh, in the collaboration with uh, different teams and also between the supplier and the client, uh, client perception was really that the supplier was not performing. And uh, in the end it turned out that the majority of the wait time was actually due to the client and not due to the supplier. Um, and this is something when, well from our more impartial perspective we can easily put on the table, well this is in none of your both interests. So really how can we collaborate more, better, more efficiently to fix that? And well there's a bunch of more examples. Um, uh, this is what it's about. So really getting, using the fine grain measures, really extrapolating directly to the real business conversation about productivity and contribution.

Speaker C: I guess it's really important when you look at those individual metrics that the way they're applied and spoken about and uh, contextualize don't lead to unhealthy competition. They don't lead to burnout as well. It's really important how do you make sure that, that. But in the context of making sure there is continuous improvement, there is continuous collaboration and there is process, how do you go about applying that an individual and at a team level it's pretty complex I guess.

Speaker B: Yeah, I think you can build consensus around that by identifying who your supporters and your champions are in your teams at the moment. And there will be those who are, have, are either passionate about this topic themselves or have observed issues that they want to add some clarity and some context around and using the proof points and demonstrating what's working for particular teams can be a great way of building support and engagement for the topic. And that coming uh, in a grassroots fashion from the ground up has way more impact than, than uh, an engineering director and sort of talking about why teams should focus on this. Um, so being able to kind of build the case studies and the vision and help teams understand the benefits some of this stuff can have can be really important. My role, to answer your other question on how do you avoid competition is to just stay authentic to that message. Say that we're not comparing individual teams and that this is a tool for teams to understand and improve their own, their own uh, productivity. And that's an easy one for me to just follow by just staying true to what we're committing to and talking to teams individually rather than saying hey team A, uh, you're not as productive as team B. Why is that? Because while we may have A metric that can be compared. It's the team's baseline, it's the team's measure. It's not an indication that that is by any means a bar that the other team has to get to. And, but in my experience when engineers have looked at some of their data, uh, the first thing they wanted to do is have something to compare it against. Just as a sense, you know, how far off from, from team B are we in team A. And hopefully that conversation should be one that say this starts with dig into the data and understand why that might be in both cases. Right? Look at the individual things that you're collecting and work out why you think those two numbers are off and what does that tell you about the differences in your development processes and then you can make the effective evaluation of whether or not those are processes you are or aren't going to keep. Um, so it can be kind of useful just as a purely holding up another benchmark against the team of people who are maybe working in some of the space. But it shouldn't be something to say just because that team has a 20% better lead, uh, time than us that we have to have. That is just a way of helping guide the conversation.

Speaker A: You might even want to say that uh, by definition the context of two different teams is different. Right? So otherwise it would just be the same team or just be a sponsor team if they want to work with the same technology stack in the same code base for the same client in the same technical ecosystem. But I think that's hardly the case. Right. There's always different perspectives there, which makes it difficult to compare. I, uh, like what you were saying about really setting it as an own baseline. Of course it's okay to look into other teams and what possible practice that these teams have been developing. Those could be reapplied or re nurtured for use in your own team. That's just a way for continuous improvement and visualizing it not for the sake of really getting to uh, a standard way of working. So the standard way of working should never be the goal. The productivity, the outcome should perhaps be the goal. Um, and not just having a uniform delivery process and be able to swap people. Right. So there's this, all kinds of nuances there. Um, I'm thinking of this example where uh, we've seen a team uh, cutting down on uh, let's say test time which led into a more feature push. So we could really measure the data that in every uh, sprint there was more production codes, uh, delivered and less uh, test Code, well, hence more uh, technical depth, hence a, uh, couple of weeks later, more rework and showing this. It's not attributed, uh, to individuals that said, well, I'm now not going to write this test or now being pushed as code is really something that is in the minds of the ecosystem across the teams that there's a culture in which we probably decide to push more on the features and rather on a sound technical harness which we are all kind of uh, licking our wounds on later on. So then, uh, this is not a productivity discussion, it's really about a cultural discussion. I think it's pretty easy. If you look at data, uh, ask these questions, I think in a matter of seconds you're talking about the real stuff rather than, hey, how can I make this developer a bit more productive? This is not the issue.

Speaker B: I think that's a great example because in identifying those issues, the team can form a more opinionated view on what is the right balance in the amount of tests that are written, its impact on lead time and its ultimate improvement on quality. And rather than just kind of saying, right, well we know we need to write uh, tests or we know we need to focus on quality, you can use it as an effective set of guardrails to work out what for you and the products you're working on, what, what's the right balance and how much time and effort needs to be spent on each of those. And if over time your, your defect rate or your, your error rate or your failed deployments increase, then maybe you know that test, test coverage needs to increase and you can communicate more effectively with what, why the, the lead time or the, the cyber time is what it is based on a narrative that you've seen a direct reduction in error rate and defect rate. And so whilst development may be 20, 30% slower as a result, quality is

Speaker A: improved by first up, first down, right, is improved, etc. So that's, that's uh, uh, the end. So that this evolutionary aspect, I think that's really pivotal, right? So monitoring those metrics over time and then trying to cautiously correlate these things, saying, well, hey, I see this. Could this be uh, that we have a bit more, more uh, defects in production because we had a bit more feature push there, there. And then it's, it's not a direct correlation. Could also be other things that are happening or just people that got ill or the, or that the team was reshuffled. Right. So lots of interesting things there.

Speaker C: That was amazing. I don't know whether you guys realized quite how Many metrics you just threw into those last statements there. But it is really, it's. I mean it does show the breadth of things that can be looked at. You know, talking everything from the classic metrics or early metrics right back from counting lines of code all the way through to some of the more modern metrics around lead time to value deployment frequency. With such a broad range of metrics out there, different objectives to different teams, different contexts within those teams as well, how do you go about the process of selecting metrics and obviously evolving those metrics over time. But what, what's your process behind set. Selecting a set of meaningful metrics that work for the team and um, for the organization.

Speaker A: You go Victor, can I, can I push in? Yeah. So what I uh, see a lot is that there's this uh, this gap between the uh, management of a IT department or a bunch of teams, uh, and uh, and then what actually the teams are doing. One of the things we are applying pretty frequently now is the um, objectives and key result approach or, or well, bit different flavor goal question metrics. So ask really what are. What is the needs and objectives that a organization would like to achieve? What are then the. Well the questions that are raised or we need to have insight and this and this and that and then translate that to the results. Either the results you are expecting from the team or the output in terms of metrics, uh, rather than reading the DORA book and saying oh all these successful organizations are measuring these and these things. So I'll just start measuring that for the sake of measuring it. So if things not be able to rely to uh, overarching goals, and this is something I think many IT teams should be in, uh, uh, should put on the table in discussing that with their management. What is it that we're expecting? Uh, is it about a market increase, is about uh, more product adoption? Is it about uh, effectiveness or, or NPS or whatever, then you will get to the right metrics. So will you start with the end of mind? Start with the goal in mind.

Speaker B: I, I think that's absolutely uh, I'd echo. Echo absolutely. That, that sentiment. You need to start with a. As uh, something you're trying to prove or at least an observation that you've had. If an organization can't form that then I would sort of question why it's trying to measure productivity. It's got to really think about what it wants to learn, what it wants to try and improve and then start to identify the ways in which you can improve it. Because without the hypothesis we have nothing to observe and we have nothing to action on. So I do think it's absolutely right that there's some, there's some fundamental questions. You need to look across the organization and try to try to form, uh, an opinion on what it is you want to improve.

Speaker A: Yeah. And quite often just adding to that. Uh, so I once made a comparison between, let's say the construction of highways, so actually physical construction world, where people know we have to talk about gravity, uh, and about planning and about ordering your goods. And then that should be there in time. You should put on piles of sand, they need to sink in, and only then you can put on the concrete or the asphalt in case of a road. Well, these kind of natural, uh, laws are not applying in, uh, in it, but there are different laws that are applying. So lots of those, those management say, well, the it can make it work. Right. So why isn't it, it's finished yet. Or so in that translation pretty quickly. You can also show how things are working by exposing the right metrics, by getting reality in and also educating and racing, uh, uh, or doing expectation management with management. Say, this is something, uh, you might want to have this project delivered by the end of the year, but our data shows it is just, it is just not finished. Right. So then deal with it or accept it, or let's take now the right business decision. Because you cannot, uh, you don't have a magic wand in which you can get this stuff fixed. You cannot get it in the real world. If there's a sinkhole in a road, you need to cut down the road, divert traffic, et cetera. But in it, I think we just generate new stuff and put it in production and it's fixed. It just doesn't happen that way.

Speaker C: Yeah, Victor, you're obviously coming at that with the benefit of not having the potholes in the roads that the UK has.

Speaker A: Uh, well, apparently, although we have our fair share here in the Netherlands as well. Right.

Speaker C: One thing around as well, I hear this. How do you avoid metric overload? Is there an optimal number I appreciate they can change over time? There's a balance between process metrics and outcome metrics and that sort of thing. But is there an optimal amount of indicators or is that not a thing?

Speaker B: I think it's important not to be too overloaded and very quickly identify which of the metrics are you learning from and which ones aren't. Are you including vanity metrics that just kind of give you a warm fuzzy feeling? Um, or are you measuring the things that really matter. And it may be that if what you're trying to measure is incredibly complex or there's a more detailed hypothesis that you want to see, whether it's the example Victor and I spoke about earlier of the management of tests delivery and quality, maybe you do need several indicators across all of that. And so the right level of data gives you the richness of uh, that story. But I think you'd always want to try and sort of have just a few sample, uh, targets that you're looking at and metrics that you're measuring to simplify and understand the challenge. You may decide over time that you can increase or decrease it, uh, based on the confidence you're getting or the usefulness of what you're getting. But yeah, I wouldn't start big, I'd start small and grow as you needed.

Speaker A: Yeah, so I think, I think doing a short retro on if you have a metrics program in place and let's say every metrics are well, part of a daily stand up or of a retrospective and then just, well, if they're infinite metrics that are all green, you should really question it whether they are still valuable to you or just for the sake of this warm feeling Gary is mentioning, I see a lot of green, so it's okay, right? This is kind of management metrics, uh, if there's metrics that are red or you're not acting upon it and probably it's also not saying enough. So those are some indicators that can really shrink down your total set of metrics to the actually the set of metrics that actually help you to steer or drive. Right. Because actually here again metrics are not, are not the goal. It's about taking decisions based on the metrics and taking the right decisions. That's the goal.

Speaker C: That's great. So I want to focus in on two last points now. Um, one's really important. It's around the developer experience and that's something we've spoken about before. It's hugely critical, not for productivity, for retention, traction, all those sorts of things. So can you tell me a little bit about the role of the impact of really driving to improve developer experience and how that can improve areas around, well, happiness as well as productivity and how they're linked.

Speaker B: I think developer experience and the perception of our experience is incredibly important because as well as being productive, we want our teams to be happy and engaged and I think as leaders that's what we focus on. Right. Happy teams are more productive by their very nature. And it also, as we mentioned earlier, is the qualitative versus the quantitative measure of whether or not these things are having a positive impact and helps us identify the things that our engineers care about, we should be putting more effort into. So I think it's a useful perspective to kind of add the more human element alongside um, the more analytical element. Um, so it can be very useful and it can capture things that the numbers can't and or it can give you insight that numbers can't. Um, engineers can tell you how a particular part of your tech stack, uh, impacts their overall enjoyment or productivity in the work just by being um, more time consuming to solve problems in or it can be um, helping understand whether or not people feel that they're more productive in that space. So there are other things that it can definitely shed a light on and that can be a real, real valuable learning.

Speaker A: Yeah, again, nothing really much to add here. There's this, lots of examples that pop up, up again about uh, some obscure technology that has been around projects, uh, for ages. We see it's a major bottleneck, but it just should be there because it was decided eight years that we should do that and we cannot retreat on that decision, which is again a fallacy because of course get rid of it as soon as possible because it's more complex, there's less people knowing it. So uh, I would say this, it right away, uh, in a proper manner. Um, and then this is exactly where then the perspective of technology and about the people that enjoy it. Uh, what are, is the tech stack direction going to? How can you incorporate those new developments into your project, uh, for the sake of developer happiness? Uh, um, and this is, and I, I would, I would conjecture that to a large extent having more flexibility and more openness and more speed in the kind of the, the process and, and, and delivery metrics highly uh, correlate with the, the increase in metrics of developer happiness. I think there's, there's a big coloration correlation there.

Speaker C: Perfect. Thanks a lot. Conscious, uh, of time and it's about time to wrap up. But I'm aware that it's probably not even a podcast these days if we don't to of AI. So could you guys give me a quick sort of sentence or two on your views around AI? I know it's a big ask to do something so big in such a small space around AI and productivity. Some of the impact that it will bring and some of the challenges and some of the considerations as well that we should be looking out for.

Speaker B: So I'll take that in two parts. Part 1 How uh, AI can be used in a developer's uh, experience to make them feel more productive and then what it can be used to help us learn from. I think we've seen tools like Copilot start to come out where tools can be effective at writing code. Um, if we're using those, inevitably the question about is that helping make our engineers better or is that helping make our engineers solve problems quicker? And there we have now a measure of how we actually answer that question. Um, in some cases we may find that while these tools do help us write things quicker, we can also now measure whether or not they lead to unintended consequences in errors or failures or issues. So there's kind of one element that having some of this tool is going to make help solve some problems. Um, can it be used in incident management? Can it be helped use to make better release processes that test better and mimic real world user behavior? There's some great things there that I think will be of real value and I'd like to see AI play a part in some of these aggregation platforms that we're using to measure productivity as well. Can it help me ask the right questions on my data? Can it help me identify the right level of trends? What insight can I get? If I'm looking through dashboards and want to understand more about the data, how can it help me? So that'd be an area I'd be really keen to see more about.

Speaker C: Thank you big sir.

Speaker A: I think so. I think of course Gary already touched upon Copilot. Uh, what I see from my perspective, interesting uh, question is about if we are using those uh, uh, more generative AI tools to create more ip. Uh, uh, whose IP is it? Right? And whose uh, and then it's not about that. Well that question but also if then stuff is wrong with that ip, who are we going to blame? Right? This is something that is uh, a bit of uncharted territory. There's some, quite some organizations also big developer organization making statements on that and also safeguarding that in contracts with our, with our clients. Um, not, not really much to do with developer productivity but in the end if it causes just failing projects or big lawsuits or etc. It does have impact. Um, this is just one thing I do like to add into the equation which is really uh, it's exciting, it's uh, it's new. You see lots of teams gain ground pretty quickly using those technologies. We are from zbia again we're training thousands of people at different clients in how to use this, this properly via Trainer, trainer program. So it's really being taken on massively. Uh, but I think right then we need to get to a kind of adoption curve to really make it, make it sustainable working. And that's I think where we are at the moment.

Speaker C: Guys, thank you. Um, what a fantastic episode. I really appreciate the way you've sort of helped me navigate some of the complexity that could be associated with this and avoid some of the controversy as well. That's uh, how there's potential for here. So that's great. Any final words and um, can you give a bit more detail of where I can find you if you'd like to reach out and get a bit more info?

Speaker B: Sure thing. So, yeah, really enjoyed that session. Mark and Victor, thank you ever so much for that. Um, and I can be found on LinkedIn or on Twitter as just Gary Stevens.

Speaker A: Yeah, from my side as well. Victor Clerk, uh, still at GBI recently, uh, got a new role, responsible for a team engaging in all uh, things Microsoft related. Uh, but equally fond of these topics that we've been discussing in this podcast. Uh, there's X for Twitter and there's LinkedIn and there's uh, all kinds of blogs that we are writing that are focusing around this topic. There's a bunch of colleagues that are also big into developer experience and that's good to uh, to keep following uh, this topic in uh, perhaps future episodes or just in terms of professional experience. Uh, this was a fun episode. Mark.

Speaker C: Thanks guys.

Speaker B: Thank you so much.

Speaker A: Thank you.

Speaker B: Now that's all for tech leaders intel today.

Speaker A: Remember you can set up regular notifications and reminders for our show by following our uh, podcast on Spreaker and all other popular streaming platforms. Thank you you for tuning in and

Speaker B: we hope to see you again soon.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why API Error Budgets Should Be Debugging BudgetsThe Developer Tools Podcast with Fexingo · on Goodhart's Law92 / 100
  • Matthias Laug from TIER on charging innovation, power of self-managed teams, and the role of technology | The Dev is in the Details #13The Dev is in the Details · on DORA metrics90 / 100
  • Beyond Dashboards: How AI Is Redefining Developer Productivity with Adeeb ValiullaShipTalk · on DORA metrics85 / 100
  • Ep. 018: Metrics That Matter - Data Points for Technology Leaders | w/ Special Guest Sebastian Kline of Nelnet Community EngagementOpen Source CXO: The Tech Leader's Podcast · on DORA metrics80 / 100
  • CI/CD with Robert ErezThe Pragmatic Engineer · on Continuous Delivery76 / 100
  • AA261 - The Business Was Dying While Every Dashboard Was GreenArguing Agile · on Goodhart's Law76 / 100

More from Tech Leaders Intel

All episodes →
  • The Value of Forgettable Customer Experiences
  • Fintech Problem-Solving: Navigating Hypothetical Frontiers
  • From Risk to Reward: The Human-Centric Approach to Business Transformation with Appian
  • Bridging the Gap: Navigating Legal and Technology Challenges in Business
  • Scaling Engineering Teams: A Playbook for Success
Explore the best B2B Leadership podcasts →
All Tech Leaders Intel episodes →