The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Dev Interrupted
Dev Interrupted artwork

How LinearB helps Kraken find hidden bottlenecks across thousands of engineers | Nik Sudan

Dev Interrupted · 2026-06-30 · 44 min

0:00--:--

Key moments - from our scoring

Substance score

52 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber12 / 20
Specificity & Evidence12 / 20
Conversational Craft7 / 20

Nik Sudan, Engineering Operations Lead at Kraken, shares practical lessons on scaling AI adoption while avoiding common pitfalls when moving from proof-of-concept to production. He emphasizes treating pilots as throwaway projects designed to validate vision rather than serve as production foundations, and highlights the critical mistake of assuming AI's speed automatically translates to production-ready systems. Sudan advocates for rigorous testing, clear stakeholder communication about pilot scope, and isolating experimental work from production environments - practices Kraken applies by maintaining separate repositories for design-led AI work while production engineers handle integration. Beyond implementation, Sudan addresses the measurement challenge most organizations overlook: tracking AI ROI through cost-per-contribution metrics balanced against quality, throughput, and stability indicators via tools like LinearB, rather than vanity metrics like adoption rates or token spend. He explains how LinearB's MCP (Model Context Protocol) server bridges data silos across GitHub, GitLab, and engineering metrics to answer complex questions about bottlenecks like review time, empowering teams to move beyond simple answers to people-and-process problems.

Key takeaways

  • →Pilots should be quick throwaway projects (days to a week) focused on proving a concept, not production foundations; avoid over-polishing or using them as architectural bases for real systems.
  • →AI spending metrics like adoption rates and token consumption don't prove value; instead balance throughput, quality, and stability metrics, with cost-per-contribution weighted by impact as the primary measure.
  • →Review time is the current bottleneck for most engineering organizations and requires data-driven analysis combining LinearB metrics with git provider data to surface people-and-process problems rather than simple technical fixes.
  • →Isolate AI experimentation from production through separate repositories and environments to prevent the temptation to promote untested PoCs to production unprepared.
  • →Automated and agent-based testing becomes critical when deploying non-deterministic AI-generated code to production, as unit test passes don't guarantee system reliability.

Guests

Nik Sudan

Topics in this episode

MCP (Model Context Protocol) serverLinearBKrakenCycle time measurementCode review metricsAI cost optimizationRework rate trackingEngineering velocityGitHub/GitLab integrationAgentic AI testing

Questions this episode answers

How should engineers transition AI pilots from proof-of-concept to production?

Treat pilots as quick, throwaway projects (days to a week) designed to prove a concept and gather feedback, not as foundations for production code. After validating the vision with stakeholders, rebuild the system with proper architecture, testing, and security practices separate from the prototype.

What metrics should companies track to measure AI's actual value versus adoption?

Balance throughput, quality, and stability metrics; the most useful measure is cost-per-contribution (AI spend divided by contribution value weighted by impact), not adoption rates or token counts. This reveals whether engineers are using AI effectively, not just frequently.

Why is review time becoming a critical bottleneck in engineering pipelines?

Review time is a people-and-process problem with many variables (team size, complexity, languages, services) that can't be solved by simple tooling changes; it requires data-driven analysis to surface the specific blockers in your organization.

How does LinearB's MCP server help identify engineering bottlenecks?

The MCP server aggregates data from LinearB metrics and git providers (GitHub, GitLab, Bitbucket) into a single context, allowing AI to analyze rich data points like review comments, file changes, and cycle time to surface root causes of bottlenecks.

Why should pilots and production work be in separate repositories?

Separation prevents the temptation to promote prototype code directly to production; it keeps non-engineers (like designers) iterating in isolated environments while engineers can later properly integrate and productionize with required architecture, caching, and system design.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode contains a handful of genuinely useful operational observations - notably the two-layer MCP data approach and the P90-over-average methodology - but these are surrounded by large stretches of generic AI commentary, chit-chat about token models, and restatements of obvious points. The insight rate is uneven.

the most useful measure is, you know, cost per contribution, AI spend divided by contribution. You know, it's a very simple one. You could compute it as spend per merged merge request or per engineer per Sprint
we prefer P90 to average for example. Right. Um, because we want to find out the slowest 10%. You know, we don't want an average because it flatters us

Originality

10 / 20

The timezone-as-dominant-bottleneck finding is a genuinely counterintuitive and evidence-backed claim that cuts against the common 'large PRs are slow' narrative. Most of the rest - data-context-insight-action, AI as a partner not replacement, Goodhart's Law - is recycled from standard engineering management discourse.

first of all, we assume that big complex merge requests were what was slowing the review time down...But we joined up the data and we looked at it and size and complexity accounted for maybe 5 to 10% the slowdown
70 to 80%. Something like that came from time zone differences, a hunch that we didn't even think about too much

Guest Caliber

12 / 20

Nik Sudan is a genuine practitioner running engineering operations at Kraken, a large-scale crypto exchange with thousands of engineers, and he speaks from real operational experience rather than theory. However, he is a LinearB customer selected for a vendor interview series, which caps independent credibility, and his title (Engineering Operations Lead) is mid-senior rather than C-suite.

review time in my view is arguably the most important part of cycle time and it is the current bottleneck for us and probably the bottleneck for many companies right now
our designers are a lot more hands on now with coding...working in a separate repository...when it's time to productionize they just hand over the repository, the files, the Linux engineers

Specificity & Evidence

12 / 20

The timezone bottleneck finding is well-quantified with approximate splits (5-10% size vs 70-80% timezone) and the cost-per-contribution metric is concretely defined. However, no absolute dollar figures for AI spend are given, timelines are absent, and most other claims remain at a conceptual level.

size and complexity accounted for maybe 5 to 10% the slowdown. Right. 80, 70 to 80%. Something like that came from time zone differences
cost per contribution, AI spend divided by contribution. You know, it's a very simple one. You could compute it as spend per merged merge request or per engineer per Sprint

Conversational Craft

7 / 20

This is a vendor-produced interview with a customer guest; the host never challenges a single claim, frequently restates the guest's answers at length before asking the next question, and steers conversation toward LinearB product mentions. A mid-episode promotional break by a third voice further underscores the format's PR nature rather than journalistic depth.

So what I hear about some of, like, the shape of how y' all think about A.I. here's what I, here's what I hear. It's like, it's almost like this um, permeable barrier
your CFO, they aren't looking at adoption rates or token counts anymore. They want to see what all of that generated code is actually delivering for the business

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B74%
  • Speaker A23%
  • Speaker C3%

Most-used words

data38engineering20production18level18engineers17team16technical15proof14doesn14different14linear14high13context13value12point12code12

Episode notes

Are you confusing a skyrocketing AI token bill with actual engineering value? This week on Dev Interrupted, Kraken's Engineering Operations Lead, Nik Sudan, joins the show to break down the harsh realities of moving agentic AI projects from pilot to production without compromising code health. He unpacks why raw AI adoption is a flawed vanity metric, detailing how his team uses tools like the LinearB MCP server to combine high-level engineering metrics with granular repository data to uncover hidden workflow bottlenecks. Finally, Nik reveals his exact playbook for translating complex data, like P90 cycle times, into a clear, business-driven narrative that secures vital buy-in from non-technical stakeholders. Life Beyond Tokenmaxxing Workshop: Watch the full replay on demand at linearb.io Follow the show:

Full transcript

44 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Today's guest is Nick Sudan, Engineering Operations Lead at Kraken. And in this session of LinearB's AI enablement interview series, Nick's going to share his insights and structural lessons about scaling AI maturity across uh, engineering teams, including their hands on experience deploying tools like the LINEARB MCP server. So Nyx, thanks so much for joining us today.

Speaker B: Thanks for having me. Uh, it's been really awesome to be here.

Speaker A: Awesome. Well, I want to go ahead and dive in about, uh, something that's really top of mind for a lot of folks that are working with AI tools is they start with this phase of just like adoption and then experimentation and then ultimately there's a mandate or a need to operationalize it and reach production. It can be really hard to ship AI powered ideas from the pilot to production in order to have like long term value. So like in your experiences at Kraken, like, what were the things that broke down during like transitioning from pilots to production? And how did y' all overcome being able to operationalize with AI?

Speaker B: Yeah, great question. So I love shipping stuff really quickly and you know, in this day and age, I love proof of concepts as well and everyone at Kraken loves it. Even before this whole AI renaissance or slopageddon, however you want to phrase it. Right. Um, but we have always been moving really fast. Right. We always value that kind of startup vibe, try and keep teams lean, you know, stuff like that. And every day we look for ways to execute faster. And AI is, it's made it extremely easy for everybody. Right. Um, including non engineers, which is where this can definitely have a greater effect than how it was before. Right. We don't want to block anyone from building out kind of pilots proof of concepts.

Speaker C: Right.

Speaker B: So transitioning from pilot proof of concept to production has always been quite tricky and I think today it's even easier now and it breaks down in a few predictable ways. So I want to call out some of the problems because I think knowing the problems will help mitigate it. But then I'll also give um, some examples of what we're trying to do to avoid it as well. So number one, spending too long on your proof of concept on your pilot. Right. It should be a quick and dirty couple of project. Quick and dirty things take you a few days or a week maximum. Right. Get something out the door is why it's called pilot, a proof of concept. You know, you're trying to get the point across. So don't overthink architecture or design. You know, engineers will spend a Lot of time on architecture, you know, designers, more visual people do design. Don't, don't um, just, just for what is the kind of core essence of what you're trying to get across? You know, like you know, with showing the BlackBerry proof of concept, you know, the um, a dumb phone just to show investors like that's what you're trying to do. Right. And just because we can build stuff in software doesn't mean that we should, you know, build something that's five times as good. Still, to keep it simple, spending weeks polishing a proof of concept and you, I think you've lost the plot. So secondly, treat your proof of concept as the foundation for your real thing. Don't do that, you know, don't treat it as the foundation. If you're building a scalable product, you will need proper architecture. Proof of concept doesn't need any of it. Right. Just goals are completely different. Promote proof of concept straight to production and you're probably going to get a lot of pain later on.

Speaker C: Right.

Speaker B: So don't be afraid to work on throwaway co projects.

Speaker A: Yeah, totally. It's about like showcasing what could be the end result more than about trying to lay every brick of the path to get 100%.

Speaker B: Yeah. And again I think a lot of engineers try to think about this and stuff and I guess the mistake is promoting that to your production product. Right. Um, and this is where I guess the third thing I would like to bring up, which I say is, I would say is more of a newer trap. So you know, now that everybody has AI, everyone could build something fast. You know, you go on X, you know, someone, someone's shooting their startup idea and then a week later they tweet, oh sorry, they post, oh, this broke. My secrets leaked to production. I guess I won't be committing them. It's like yeah, for us it's like no shit of course but, but um, you know, with AI and building something fast as well, there's a temptation to think it, you know, the speed, just the high output scales it consistently all the way to production. Right. So you can iterate quickly, get things in front of people quickly. But making something that's scalable production ready system, you're still going to need to take time with it. AI can help and will help you make it faster. But uh-huh. You take out all of the thinking, you get agents building it, it's gonna, it's gonna break. You know, pressing one button, telling, you know, turning on auto modem. Claude isn't going to ship you A money maker. You know, it might ship you something that looks like it does, but it doesn't. Right, right. Um, so I guess some examples of what we're doing at Kraken as well to prevent this as well. So testing, specifically automated testing as well. Um, and agent based testing. Right. Um, and also tried and tested, just internal testing with your real users. The bar for something being done has to be deliberately higher than you know, this worked in the demo. Right. Especially with AI. Now as a, as a engineer myself, I spend less time coding properly now. I will prompt stuff, I'll have a whole bunch of agents doing stuff and I'll verify it works and then you know, send it over. But that's, you know, my, my trust in that has dropped significantly. So relying on testing and especially, you know, you run a company with hundreds, thousands of engineers, you're not going to keep on top of it. So investing in testing works out a lot there. And AI having AI in the loop now, um, you know, where you have more non deterministic behavior. Right. Means um, uh, passing a unit test isn't the same as like a system. Right. So I think having, incorporating a genetic testing that works out well. Um, another thing, isolating your proof of concept from production. So um, we build proof of concepts in separate environments to resist that temptation to kind of. Right, let's actually turn it on. So throw away repos, you know, very light scaffolding for deployments and stuff like that. So people, you know, maybe relaxed kind of repository kind of push roles and stuff so people can just iterate quickly without dragging production's concerns and stuff. Um, our designers, so um, our designers are a lot more hands on now with coding and such. Um, and they're working on kind of building out their vision for certain projects, but they're working in a separate repository. Like we have our all our mobile M app, our web app repositories. But you know, we've given them an environment where they can quickly iterate, get something, it's removed from the production app. Um, and then when it's time to productionize they just hand over the repository, the files, the Linux engineers. And that's actually a lot easier to work with than a figma file or a PDF or an image.

Speaker C: Right, right.

Speaker B: And, and that engineer is then uh, okay, this is how it looks, let's productionize it, integrate it properly. You know, they know how the design system works, they know how to build things effectively, cache things effectively. You know, designers are good at design, so let them work, focus on the visual Part in that separate environment. And I think, um, lastly as well, be explicit about that pilot is right, um, because I think many stakeholders, many reviewers aren't clear that this is the proof of concept. You show it to them and they're going, oh, so can we expect it in production next week? It's like, no, no, no, no. This is something different, right? You need to make it very clear that this is not your final production product. Um, value in there is to prove the point, not to ship. You know, show a working product instead of writing long AI generated briefs, running endless meetings, right? Showing something is worth a thousand words, especially for people who are really busy all the time, right? Prove that you can build the thing that they want, the thing that they envision in their heads grants them a lot more confidence. And they'll give you feed really, really good feedback there and then. Because if you can, if you get them documents, if you talk to them endlessly, they'll be like, oh, I'm not really listening. This sounds all right. But you show them something, they'll give you feedback immediately. Right? So this is. And just making that very clear for them, right? You know, I think everyone who's built products, you know, you. It's like a week before shipping to production, some stakeholder comes in and it goes, actually, this isn't right. And it's like everyone panics, like, oh, God. Oh, no. Right?

Speaker A: Like, who let the. Who let them see it before we shift it?

Speaker B: Yeah, yeah, yeah. But it's like, well, why didn't they see it earlier? And it's like, why didn't they? Yeah, yeah. So and this is where the value of it comes in, right? Get the feedback there. And you know, having feedback at the end is expensive, right? It takes time. It causes, uh, regressions as well, because you're checking if you're changing massive stuff last minute, the chance of it not working in production, right? And especially with, you know, with linear B, you can measure the kind of rework rate and stuff like that. You know, you want to keep that kind of stable, right? And to kind of maintain that rework rate, you need to make sure that you kind of get things right. Have the start as soon as you can. Right, right. So just kind of making that kind of clear line. This is proving the idea and then this is how we make it real, right? That's kind of the ethos around it. Um, right, yeah, that's kind of my take on that.

Speaker A: So what I hear about some of, like, the shape of how y' all think about A.I. here's what I, here's what I hear. It's like, it's almost like this um, permeable barrier. Inside of it is all of this experimentation and rapid iteration. And inside of this world you're not creating production level, foundational COD code for what's going to be what drives the uh, the front of the sleigh for the next year. Instead you're creating the, the visions, the interfaces, the hooks, the things that you can show to non technical stakeholders, to the technical leaders and to your fellow engineers to all coalesce around what is, what, what are we trying to ship. And so you get this rapid bubble, I call it a permeable barrier because ultimately from that something has to exit and now start to enter, you know, the sdlc, the agentic, dlc, uh, whatever you want to label it as for your particular org. And so you know, you mentioned for example looking at uh, you know, when engineers are working, keeping the rework rate stable and tracking that through linear B, for example, helps you keep a handle on overall code and organizational health as things cross out of that barrier. So that's kind of what I want to, I want to zoom in on next and try to understand is when y' all are as engineers who've built the way to go from that vision, we're all on the same page to now, we're going to lay the foundation brick by brick and get there uh, in a really stable way and ship it. You uh, know how do you start to partner with uh, like your tooling to not only all get aligned but then also like to understand like is this working? Um, and are the practices that are leaving the barrier and starting to enter our adlc, um, are they effective? Like what are the kinds of ways that y' all look to that and make sure that your engineering pipeline stays healthy.

Speaker B: Everyone at Kraken is using AI right now. It's whether we like it or not, um, it has been a complete game changer when it comes to software engineering. I guess the myth of like it replacing, you know, software engineering isn't really coming about and to be honest it shouldn't because it is a partner, it is a tool. Right? Kind of, you know, going from assembly language to kind of typed code, uh, to frameworks, AI is just kind of another level of that on top. But rather than it being very deterministic, it's agentic. Right. And this is kind of how we interface with the modern world right now. And you know, everyone at, not just at Kraken, but I think everywhere Right. They're jumping on AI to win on kind of throughputs on output 10x their impact. Um, but the spend is, I think, one of the biggest, um, things here. Right. Um, to the point where it's close to or matching the salary of a developer. So how. So the question here is, you know, if, if it gets more expensive, um, you know, when, you know, with like, you know, Fable coming out now and, you know, people have been like, oh, I asked it one question and there goes all of my tokens. You know, stuff like that.

Speaker A: Yeah. I'm currently waiting for a token refresh. I totally understand.

Speaker B: I. I have not switched over to Fable yet. I saw the pop up and I'm like, I. I'm. I'm going to stay with Opus for now.

Speaker A: You know, you're missing out, Nick.

Speaker C: All right.

Speaker B: I don't need it to do that for many things. Actually, I guess that's one thing as well, which I'll get into a bit. But just effectively, what's the best model for your work as well. But yeah, cost, whether we like it or not, is the biggest thing about AI impact now and whether we, uh, get value out of it. Because at the end of the day, that's like time is money, right? Uh, right.

Speaker A: It's like the value of an engineer is like the cost of the tokens that they're consuming.

Speaker B: And so I even think AI is going to get even more expensive over time. Um, and businesses who aren't thinking about that regarding ROI will be starting to think about it now, this week in particular. Right. Rather than assuming that today's costs are, uh, this is how it is going to be in the business. So what my take is in regarding how to measure the value of using AI with your SDLC and such. So I think most leaders actually make one of these two mistakes. So the first is AI usage is, uh, or spend is proof of value. Right. So the assumption is if cost is high for an individual, for a team, we're getting our money's worth. Right. And then measuring adoption, like how many seats are being filled, percentage of engineers using AI adoption dashboards, tracking all of that stuff, they look great, but they don't tell you anything about whether the work gets better. Right.

Speaker C: Um.

Speaker B: Um, AI usage isn't a metric that is great for effectiveness. It's one signal across many, but by itself it doesn't prove anything regarding SDLC and, uh, the value of AI. So what you would actually want here and what my recommendation is and what we're starting to do at KRAKEN is you have to balance a series of metrics across throughput, quality, stability. Right. So with throughput merge request output frequency and the maturity and the size of those merge requests, all that you can see within linear V. Right, Quality. So, uh, like, you know, what's the rework rates of the changes? What about bug escape rates, you know, how deep are the reviews that actually happen? And then when it comes to stability, you know, is the thing that people ship to production actually holding up? Right. Any incidents, any problems, regressions, and all of it has to be contextualized as well. You know, what kind of projects are people working on? What are they contributing to? What's the real impact of those projects when it comes to the company and the products? Right. I think for me the most useful measure is, you know, cost per contribution, AI spend divided by contribution. You know, it's a very simple one. You could compute it as spend per merged merge request or per engineer per Sprint or something like that. Right. As people use AI more, their output should increase. But for generally valuable work where the cost per contribution stays stable, that's what you're looking for. Right. Um, and then this only works if you weigh by impact as well. Right. If you measure it raw, you're going to reward whoever ships the most trivial load changes. Right. I guess the best way to describe this with an example is this. You take two engineers who have the same output and the same impact. One spends thousands of dollars and the other hundreds. Who's the most effective engineer here? Maybe the one who's burning thousands of tokens has more raw output. Right. But the quality and impact of their changes is poor. Then the output is worthless. Right. What we want to have is people using AI effectively, you know, using tokens. Well, knowing which model to reason for, as we were talking about caching, optimizing the usage with context, knowing when to use deterministic kind of, um, thinking versus agentic AI isn't a replacement for this. It is an enhancer. So we need to make sure that measure it like that. And yeah, I don't think people out there are really thinking about this too much. So that's my take on that.

Speaker C: Andrew and I, we just wrapped up this really great session on token maxing and we had a lot of fun, you know, running through it. We got a lot of feedback, a lot of questions from the audience, and it's, it's very clear that this topic is really hitting a nerve right now. And I think it really has to do with the fact that executive conversations all over the place right now around AI has really shifted from this let's get everyone using it to now everyone's wondering how much are we spending and is it actually worth what we're spending? So if you missed the live stream, we have the full replay on Demand over at LinearB IO. And you know the reality is that your CFO, they aren't looking at adoption rates or token counts anymore. They want to see what all of that generated code is actually delivering for the business. So in this session we map out exactly where AI is shifting bottlenecks in your pipeline and how the linear Bees Apex framework helps you measure what is really valuable to your business. So we'll share a link in the show notes, but you can also head over to linearbee.IO to check out the

Speaker A: full session and you get like uh, everyone's over rotated and over focused on like how do we get in the code layer and get in the IDE and help make sure we're steered and we're all aligned on like we have uh, configs and you know, an agent MD and best practices, whatever. Like all of that stuff is super helpful. But what you've just described is how there's a higher order problem to solve to be effective. It's your entire engineering process needs a harness. In your case it's like a linear B is that harness because it gives you this through this like view into all of these different health signals and things like quality and throughput across all of the code you're shipping. So that when you are you know, asked or presented and you're looking at the numbers of like your token consumption and are we getting value of this? You can actually, you have a contextualized narrative for like why you are getting value out of this and uh, that kind of like bridge bridges into the next thing that I want to talk about which is connecting like data silos and actually using that to get really good velocity as a team. Because um, it sounds like from what you're describing you have a lot of signals that come into one place and then it helps you understand how effective um, our coding processes, how effectively are we shipping. But then what's the health of it downstream? Which is a really important problem to always have in perspective. Uh, so like one example is uh, when we first were chatting about this, you mentioned that Kraken was using MCP server from LinearB to uh, kind of contextualize the information but also combine it with these other sources and tools to create really effective uh, things that drove decision uh, making internally so how are y' all thinking about using mcp, for example, to distribute this engineering knowledge to folks?

Speaker B: So taking a couple of steps back, this boils down all boils down mcps, AI boils down to this one thing, which is the thing that AI is really good at right now, which is data analysis, right? So data analysis isn't new. People are in that field for many years now. But I think the difference now, um, and the relevancy to the question right now is AI lets anyone be an effective data analyst, right? It's the fact that, you know, now that everyone has a camera in their pocket, anyone can be a photographer. You don't need crazy equipment. Anyone could be a data analyst now, right? Good data, anyone can analyze it, but they need to be asking the right questions. You know, just because you have a camera in your pocket doesn't mean a really blurry photo of like, you know, crooked angle and stuff is it doesn't make you, makes you a photographer, but it makes you a horrible one. Oh yeah. So it's the same with data, it

Speaker A: makes you a photographer. Sure, yeah, yeah, yeah.

Speaker B: So quite lots of air quotes there for sure. Um, so step one, at least with all these data silos, the MTP is to get the data in front of you and as rich and as sensibly as can be, not over enriched because then you're going to be spending more and more with context and stuff as we talked about before, but with it at a point where every data point you think might come in handy is there, so timestamps, you know, names of people who reviewed it, comments, how many comments were left on merge requests, you know, what files were touched, anything that might explain kind of SKU or size of the change, right. So Linear B is great at uh, providing the high level stuff, um, here and then when you, so when you're using your kind of git, um, provider here, so GitHub, GitLab, Big Bucket, you know, that's where you can dive in deeper, right. And this is where the bridge comes in there. So example here we were asking probably one of the most common questions that people get. Why does review time take so long?

Speaker C: Right.

Speaker B: Um, and tools like Linear B help surface it, uh, and you have to do a bit of digging in to find out why. But it's not a simple answer when you have, you know, thousands of engineers contributing tons per day across different types of services, types of programming languages, types of like there's so many factors, right? And if you just get present like this needs to improve you need the data, right. So review time in my view is arguably the most important part of cycle time and it is the current bottleneck for us and probably the bottleneck for many companies right now. And the one thing that AI can help out with a lot because it's a lot more like agency with the way of thinking. There are many different possibilities. It's not as simple as like oh, if the deploy time is broke, that means we need to speed up the pipelines or a number of runners available or something like that. Right. Or reduce the number of dependencies. But review time that it, it's, it's a whole black box of like it could be this, it could be that. Right? A people and process problem rather than a tooling one. Right?

Speaker A: Yeah.

Speaker B: So with the structure, with the structural bit, right, as we were talking about before, um, so yeah, use linear B to aggregate all of that at a pure data level. Um, you know, so right now you're not using the AI to you know, aggregate or kind of make decisions. But the MCP is really good to get that data out. The MCP can help answer questions, but one data source alone, specific, especially as linearly is focused more on the high level stuff. Um, that's the entry point, right? High level themes, trends. And it shouldn't pull in the granular themes at this point because it's going to get expensive, it's going to get unnecessary. It gives you the opening path to kind of make that decision. Kind of go like right, this could be an avenue, this could be an avenue, right? If you dumped every low level detail and it would just muddy the signal as well. So high level trend is very important to first to identify. So in this case review time is slow. We find out the relevant merge requests, the repositories, the teams. Linear B is really good at that, right? Because we like we ingest thousands and thousands every week at that point. Then once you've got kind of the high level data then we use our uh, um, repository, uh, our uh, git provider MCP to then. And also maybe not even MCP at this point because if you plug in a CLI, which I actually prefer to MTPs because they're not as flaky, their authentications are long lived and sometimes they even have um, more capabilities. But whatever you use here in the agentic interface, this is where you can join, right? And the two layer layer approach works really well. Linear B high, you know, then your repository provider, uh, the low and then AI is the part that moves fluidly between each other and Enriches it. Right. So there's no human stitching and making the connections itself. And that's, that's the biggest time saver right now. And it also pays to be scientific in this rather than just trusting the hunch. Too many people trust AI and AI is a yes person. It's never gonna, and, and it's, it's generic as hell. So you need to be thinking critically, you need to be challenging it and you need to have hunches. You know, our hunch was wrong actually. Right. So first of all, we assume that big complex merge requests were what was slowing the review time down. That's what most people say. You know, linearb gives those signals as well. You know, review time, bad merge request, large size must be the cause. But we joined up the data and we looked at it and size and complexity accounted for maybe 5 to 10% the slowdown. Right. 80, 70 to 80%. Something like that came from time zone differences, a hunch that we didn't even think about too much. Team silos, reviewer availability, engineers in one region, opening a merger quest at a time that didn't line up with the code owners for the areas. Um, and it was really hard to analyze all that type of stuff and who are largely, who were uh, largely in lovers so that the work just sits waiting for a reviewer who wasn't online. Um, regardless of size. Right. And that's the hidden bottleneck that these silos are masking. And now that we could prove it, it stopped being a hunch, you know, became a real data backed problem that leadership became aware of. And the fix we're working on right now is to expand kind of the pool of code owners, you know, so reviews aren't gated to a single time zone. And you know, it's, it's a complicated thing, you know, people and process problem. But having the data to prove the hypothesis, like you know, being scientific about it, that's, that's the thing, that's the structural shift, you know, it moves leaders from the anecdote driven, the evidence driven. Right. Except of instead of acting on the feeling and maybe some high level metrics or I think you know, when you're having one on ones. If you're like an engineering manager, you uh, might get someone going, oh yeah, no, I've had some slowdowns working with this team on that and you might just be like, I'll pencil it in or put it for later. But that maybe that's the hunch to start digging a bit deeper. Right. Find out why that's happening, right. Form the hypothesis, test the data actually supports this and uh, be open if it doesn't. Right. So the one thing that AI doesn't do for you is it doesn't decide what you should ask. That has to come from you, presents, it helps with the analysis. But the quality of what you get out of that assessment is bounded by the quality and the context that you give it as well. Right. So that's ah, a little bit of how I use it and how I recommend kind of people use multiple sources as well. You know, use Jira. You know, that's really good for knowing the context as well. If you're using teams or Slack, you know, get the conversations in as well because sometimes it's not as simple as, you know, at the end of the day in engineering everything comes kind of code change but there are so many factors as well, right? This is just a very simple example of that. Right. Um, you know, plug in schedules, calendars, every, all sorts of little nuances, right. There's so many different connections now that are supported with mcps that uh, you'd be crazy not to use them.

Speaker A: Right. It's about, it's about contextualizing all of the data that's available to you so that you can, you can take your high level questions that you yourself as the expert in the domain in which you serve value to your end users can pose these really complex questions that before would have taken a team of data scientists and maybe a few weeks a crunch for you that now you can do really quickly by assembling these tools. And I totally know what you mean about like sometimes MCP can feel like it's in the way and for certain workflows it certainly is. And for yours, you know, you might benefit more from like a CLI based approach for a lot of the stuff you do. But one thing that MCP solves really well that another other toolings and ways of sharing context with AI haven't is distribution. Because that MCP server can be really relatively simple to put into the hands of somebody who is less technical and equip them with abilities to take their high level queries and actually get the crunch of the numbers. But therein lies the trap that you, you've very, very smartly identified and stepped around that you know, one, the AI is a yes man and two, you have to have that understanding of what you're doing to ask the right questions that ultimately extract the right stuff you want and if you drown it in data, you're not really going to get, you're going to get a um, facsimile like a mirrored uh, version of what you're asking instead of like the core truth. Right. So something I want to ask you about is for example how you equip those stakeholders to be able to understand what's going on in the engineering world through tools like the LinearB MCP server for example, because you can distribute that or otherwise make it accessible. So like uh, for example you've talked a lot about cycle time in this conversation, about why that's important to you. How do you use uh, these kinds of tools and distribution of context to get stakeholders that are non technical in the conversation and really understanding what's going on.

Speaker B: So when it comes to contextualizing, I think these metrics, these kind of things, they shouldn't come top up, they should come from the ground level. Um, and anything that's important for an engineering team should be important to that higher level audience. I think most times this isn't the case because of that contextualization, that translation, you know, into the non technical terms. And you know, this is why mcps are really good right now. So to give a bit of an example as to kind of like how we translate stuff Kraken, uh, so you know, we prefer P90 to average for example. Right. Um, because we want to find out the slowest 10%. You know, we don't want an average because it flatters us. Right. And if a team is doing well overall, it quietly absorbs the areas that are actually struggling and you never see them. And P90 surfaces that slow tail. Which is exactly what us as an engineering one when we're trying to improve, we hold a high bar for quality and technical ability. And P90 is where we want to go and hold everyone rather than hiding behind a comfortable mean. And so by doing that it's to stakeholders who hire up they might go oh, this is really slow. But it's like no, no, no, this, this is what we want to do because we, A uh, metric shouldn't like, you know, when a metric starts becoming a target, it's not a target, you know, Goodhart's law in this case. Right. So a metric should be helping you inform. So the aggregation is one thing. Um, but then also it's not just about you know, a single goal or a company wide P90 for us, different teams work in different ways on different services and different repositories with different repositories and sorry, different dependencies. Right. And some are ah, slower for entirely legitimate reasons. Um, so Comparing all of these across one global number, ultimately meaningless, right? Um, and this is where kind of this like, you know, hey, here's an average for the whole company. Like that's not how, you know, maybe if you're a really small team, but if you, but once you start to productionize things, once you start to start being a company with hundreds of thousands of engineers, you need to understand the different domain contexts as well, right? So identify the right development groups upfront, isolate them in linear B, you know, sometimes at the team level, sometimes subgroup like a front end or a backend or mobile, which scope to relevant repositories and then treat the industry average as sort of like the default benchmark, you know. And if a team comes up red against it, um, we don't just go, oh, so that seems rubbish. Like look into it. Because sometimes it might be explained like hey, our way of working just, it doesn't, it doesn't work with this stuff. And that's a fair outcome, right? And that's the context and that's what you need to identify before or the, before you kind of move over to stakeholders, non technical stakeholders, like you have to do this groundwork to contextualize it, um, you know, what is their green line, Set that goal, right? So that's the contextualization preamble. But then you have to translate it, right? And this is where you can't just give that to non technical stakeholders and such. Um, you can't just throw metrics at people and expect them to understand them. You know, I think many people, especially like at uh, C levels and staff director level, get many dashboards which are just like a bunch of numbers, red and green. Like you know, if they don't understand it, they don't trust it, you know, and that's, it's completely worthless. And if they don't trust it, um, they won't act on it. Right? So you know, what is the whole point by doing that? So just for me, describe everything in plain language. Cycle time. I won't assume people know what it is even if they're an engineer or not.

Speaker C: Right?

Speaker B: Um, because cycle time actually depending on what um, metrics aggregation do, it could be different. Some people go to merge, some people to deploy, you know, what does it mean? So um, describe it in plain language. I want. So I'll say basically look how long it takes us to ship something in front of our clients and link it to things that we care about, you know, like cycle to a higher cycle or a lower cycle time equals faster project execution. Speed and it's like okay, right, right. So like you know there's the technical term but then there's like the non technical term. And also the non technical term helps out engineers as well. Because um, translate anything to basic language um, and help it be contextualized then it, you know, it all becomes clear. Right. So what does this mean in practice? How do you do this actually? So you have to make it concrete. Product design, operation leaders, uh, this number is high for this team which means this project is going slower. Like get actually sit maybe in a call or in them with them, just walk through them, no stupid questions, bring it to that table. Many people might dismiss it as like we don't need to run through these engineering things but you know, this project is moving slower. They don't want to hear that. You need to make sure you translate it into that terms and then you say why as well. And don't just stop at the diagnosis.

Speaker C: Right.

Speaker B: Call out specific areas that may need improvement so you can give actions and insights rather than it being red light. Right. And um, doing this kind of helping with these leaders, bringing to them being that translation kind of Google Translate for them if you will. Unlocks prioritization of like engineering driven initiatives. Getting non technical stakeholders to understand the reasons behind the metrics will actually allow you to get time to work on tech debt. Right. And engineering initiatives. Get them onto a roadmap that's already full of product initiatives or we need to build this, we need to build that. Every company knows how that is. Um, so that's how you do that. If leadership understands red what it means and why, you can make the case of carving out that time. Right. And without that transition they're not going to give it to be like they don't need to work on that. We need to ship that. That's going to make money. But if you can conceptualize it translation thing. But this is money, you know, that's it. That's why product have all their KPIs and stuff. I know as engineers we don't, we hate hearing you know, KPI okay uh, all these acronyms. But you know what if we want to get them on board, we have to speak their language. Right. And then underneath all of it, you know, trust and adoption, you know, don't compare teams to each other gamifying it. That's the wrong move. And that's the first move that many people make. Right. Evaluate each team against their own benchmark. Not a leaderboard of like best versus worse. The message to each team Is like, hey, this is for you, this, these metrics are for you. Um, you know, if you're green, great, nothing to worry about. And if you're not, it's a signal to kind of step up. And the right use of the data is like what's worth learning from it, right? Is it the internal team working under the same constraints which beats the generic industry fix every time, like you know, just trying to, you know, make it contextual for each team. Um, and I think this matters most with leadership who are often quick to look at, you know, who's performing worst and then you know, cut them, you know that's wrong and you need to make sure that's translated to them as well.

Speaker C: Right?

Speaker B: Because the metric is a metric. But then the rationale as to why that's happening, a team at the bottom is usually one suffering from friction or they need engineering investment. And it needs non technical leadership to give it that time, that backing to fix it, not the chopping block. Because if you chop it, you're just going to make the problem worse. Right?

Speaker A: It highlights the opportunity because by chopping off the end of it, now there's a new end of the problem. Because you haven't address the core of why that might be read right for you. And like it's really brilliant, Nick, how you think about bringing everybody into the conversation, not only like translating all the technical jargon to the engineers, which is like what we're really, or to like the non technical stakeholders, which is like a really big driving function of linear B. Right. But also to using it to contextualize those higher level decisions and things like KPIs and OKRs to the engineers and help them map the work that they're doing to that direct impact to what we're shipping and to why that matters for our organization. And ultimately like by using a platform like LinearB to have an informed shared conversation, you can build trust. Because now like you said, those leaders can trust and understand their dashboard. They might not be as fluent as you in talking about the healths and the metrics of their, of their conversation, but you all now have like a pidgin language you can operate in. You have like a, a way to get each other's point across and then also effectively get the work done. And for a lot of engineering teams, that this is a completely missing ingredient. It was before. It is now. And now we're trying to go even faster. And as you've rightly called out, most of the problems we're bumping into are people and Process problems. You need to have everybody equipped with the same language and ability to operate on the same kind of data.

Speaker B: Right?

Speaker A: So I really love all of these call outs. I think it's a really like, smart way of understanding problems, uh, even just like things like using this MCP, uh, data layer to curate your own like P90 benchmark layer. Right? So you can understand these specific points because that was a priority for you. And the, the beauty of this is it can be arranged to be a microscope onto any kind of problem that you might think exists in your org. And if you are fluent in the data, then you're not going to have this whole, uh, you know, can I trust the hypothesis situation because it's like grounded in the truth of your engineering Org. It's all a really great kind of, uh, m. I think like a playbook for how to use linear B to like ship things at scale, but then also have conversations, uh, where everybody in that Zoom meeting or whatever, they all understand what we're talking about because for the first time we're all now fluent in like the health of our org. Really great kind of like way of putting that all together. I do want to say, you know, we've covered a lot of ground today and just before we wrap up, are there any, you know, last parting thoughts or words that you have for folks that maybe we're listening to this conversation and want to have a bite of that success as well?

Speaker B: One word, four letters data. You know, it's everything. If you're not measuring AI effectiveness, adoption cost and relating to that as well, you know, the quantitative engineering data, like code throughput, Dora, then it's all worthless, 100% worthless. And this has to happen on day one because if you're not doing it already, like, you know, it's going to be incredibly costly later on. Start it now. Like prioritize this, like after listening to this, just do it. Um, you need to be able to prove your point of data. And adopting kind of AI, adopting all these tools is not as simple as like, hey, we're going to procure a code editor or license for some framework. It's a whole new way of working, right? And you have to, you have to measure it like a new way. Right? But it's not just AI specific, right? It's a chain, like data, context, insight, action, right? That's the layer, um, each layer kind of gets built on it before. And most teams actually stop early. So, you know, raw metrics, step one, data and gestation. Good, you've got it. You Pull most of it where code lives and tools like Linear V help support this. Right. Um, core metrics, throughput cycle time, review time, et cetera. You know, this is part both teams are doing right now. Um, but the context part, this is the second part derived from the foundation. You're bringing the context sources that enrich that quantitative data. You know, could it be qualitative? You know, like, what are people saying? What projects? What are the actual projects, the company's priorities? You know, what's your expected outcome of all of this? Financial impact, you know, adoption impact metrics will tell you how fast your engine's moving, but context tells you, you know, are we going in the right direction? Right. And then with that data and context gives you insights. Right. Someone using AI, uh, more doesn't mean they're great. You know, as we said before, anything heavy usage is a signal to go and look what they're doing. Right. You know, are they actually doing it? Is everyone using it well and saying it's brilliant, but, uh, you don't see a business impact. Something's going wrong.

Speaker C: Right.

Speaker B: Uh, people hacking away their own little projects to improve their own little processes is fine, but needs to show results. Right. You know, it's so tempting to work on. Like, I'm going to revolutionize how the company works and roll it out to one person, which is yourself. No, it's not. You know, if it's not doing that, rethink it a hundred percent. And then this is where kind of the intentionally intentionality matters, right? Someone using AI more doesn't mean they're great. And I think with that, keep it tight as well, leave room for experimentation. But, uh, make sure generally, it's generally, you know, pushing your engineering organization forward by finding out what is working and what doesn't and back up all of that data, all of that quantitative data with the qualitative data. Surveys, like space surveys, are pretty good for this as well. That's what I recommend. And there are many other frameworks out there. Like, you know, Dora is something, although that's more quantitative. And then plenty of reading materials to look at the numbers tell you what's happening, but then the people tell you why. Right. So, you know, if I was advising someone, the real kind of final line here, measurement in place, day one, make sure it spans both the raw metrics and the context that enriches it. Keep it all intentional and build towards that, uh, insight and action rather than just looking at the numbers and stopping at the numbers. Right. The tools then will thrive with all of that.

Speaker A: Well, I think this has been a really great view into how y' all operationalize with AI as well as how you think about working with data. You know, data is really the king here. It's going to allow you to operate, uh, and soar at this level. So thanks for breaking it all down for us, Nick, and, uh, it was really great chatting with you. I hope to have another conversation with you in the future. And thanks again for joining us today.

Speaker B: Thank you very much.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Autonomous Software Development at Enterprise Scale: Inside a 1,000-Developer Pilot (with Blitzy) | CXOTalk #918CXOTalk · on GitHub/GitLab integration97 / 100
  • Inside the $1B-a-Day Stablecoin Market Maker for 1,500 Institutions, with B2C2's Cactus RaaziThe Fintech Blueprint · on Kraken76 / 100
  • Game Changers: How Sports Shape Seattle’s FutureThe Leadership Playbook · on Kraken74 / 100
  • S6 Ep16 - Shane Neman - Seasoned Entrepreneur and Venture CapitalistPitch Deck · on Kraken70 / 100
  • Introducing Skein and Beaker Stack: Using AI Agent Engineers to Ship SaaS FasterHow Many CTOs · on MCP (Model Context Protocol) server67 / 100
  • Science of SaaS Startups Podcast with Conor Bronsdon - LinearBScience of SaaS Startups · on LinearB62 / 100

More from Dev Interrupted

All episodes →
  • The discernment horizon, loop-driven development, and a wizard’s very defensible pond
  • Your developers are the attack surface now and vibe coding as a vulnerability | Tanya Janca
  • Microsoft’s wandering eyes, data labeling duties for senior devs at Meta, and prod is the new source code
  • Your SDLC needs a productivity context engine
  • How to harness your dragon with Fable, tech leaders turn to model routing, and coping with AI rockstars
Explore the best B2B Engineering & DevTools podcasts →
All Dev Interrupted episodes →