The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/AI to ROI
AI to ROI artwork

Big Book of AI Metrics

AI to ROI · 2026-07-02 · 34 min

0:00--:--

Key moments - from our scoring

Substance score

38 / 100

Five dimensions, 20 points each

Insight Density8 / 20
Originality6 / 20
Guest Caliber8 / 20
Specificity & Evidence12 / 20
Conversational Craft4 / 20

Ray Reich and Peter Buchanan unpack why measuring AI ROI has become urgent and structural. With enterprises deploying AI aggressively but unable to prove financial returns, the Big Book of AI Metrics fills a void left by consultant frameworks and LinkedIn pontification. The book is built on Ray's 20+ years as a recurring revenue operator and his founding of Benchmarket, a metrics and benchmarking research firm, plus his role establishing the SaaS Metric Standards Board in 2021.

The core thesis: companies measuring adoption and token consumption are confusing motion with progress. True success requires a causal chain framework linking input signals, leading indicators, operational KPIs, financial outcomes, and strategic value. The book targets two audiences - operating executives investing in AI across functions (engineering, sales, finance, supply chain) who need department-specific metrics, and B2B SaaS/AI-native company executives measuring inference costs, gross margin, and enterprise value impact. Real wins like Petrobas saving $120 million on tax processing in three weeks show what happens when baseline, scope, and financial measurement are locked in. Failures like Uber's Claude Code deployment (84% adoption but zero features shipped, budget exhausted in months) illustrate the cost of measuring activity instead of outcomes.

Key takeaways

  • →Define measurement success metrics before deploying AI, establish a baseline of current process costs, and measure continuously - adoption and token consumption are not ROI.
  • →Use a five-layer causal chain framework: input signals → leading indicators → operational KPIs → financial outcomes → strategic value, ensuring each layer connects to business impact on the income statement.
  • →The biggest failure pattern is measuring improvement against no baseline; without a documented control group and agreed baseline, claims of AI improvement are opinion, not evidence.
  • →AI spend is becoming compensation-scale operating expense (30-40% of revenue for AI product companies), requiring CFO-led governance metrics like cost per outcome and token spend as percentage of COGS.
  • →Continuous measurement prevents model drift and performance decay; deploying once and stopping measurement means you lose visibility to hallucinations, bad data, and degradation over time.

Guests

Ray Reich

Topics in this episode

ROI measurementToken ConsumptionBig Book of AI MetricsBenchmarketSaaS Metric Standards BoardAI performance metrics frameworkcausal chain frameworkinference costPetrobas (tax processing automation)Automation Anywhere

Questions this episode answers

Why is establishing a baseline before deploying AI so critical?

Without a baseline of current process cost and performance, you cannot prove AI had a causal impact on improvements. The baseline acts as your control group; without it, you have opinion, not evidence of ROI.

What are the five levels of the AI performance metrics framework?

Input signals (data impacted by AI), leading indicators (volume or efficiency changes), operational KPIs (functional metrics like win rate or inventory turns), financial outcomes (income statement/balance sheet impact), and strategic value (competitive positioning, market share, valuation multiples).

Why do companies like Petrobas succeed at AI ROI while companies like Uber fail?

Petrobas succeeded because they had narrow scope, a clear baseline (15 years of manual tax work), verifiable financial outputs ($120M in savings), and measured outcomes. Uber failed because they measured adoption (84% of engineers using Claude) and rewarded token consumption without tracking cost or quality, exhausting annual budget with zero new features shipped.

How should enterprises measure AI in operations phase?

Track attribution between AI use and specific business outcomes; measure operational KPIs linked to financial performance (not just task completion); implement continuous monitoring for model decay, hallucinations, and data quality; and establish a cadence that persists beyond initial deployment.

What makes the Big Book of AI Metrics different from existing frameworks like Gartner or McKinsey?

Most frameworks tell you metrics matter but don't define what to measure, how to ensure baseline pre/post comparison, or what good results look like. This book is built for operators and practitioners with 81 specific, actionable metrics organized by 13 functional roles and AI lifecycle stage.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

8 / 20

The episode delivers some usable framework concepts - particularly the five-level causal chain from input signals to strategic value, and the distinction between operational KPIs and financial outcomes - but a large share of runtime is promotional preamble for their own book and generic measurement advice dressed up as revelation.

the biggest gap I almost always see is that measurement stops at an operational KPI and is not linked to the financial performance outcome. That's the biggest chasm I see
AI isn't 1 to 2% of revenue the way it spend is. It might be 30, 40% of revenue

Originality

6 / 20

The core thesis - measure outcomes not activities, establish baselines before deploying - is measurement-101 advice recycled into an AI wrapper; the framing is timely but the thinking is neither contrarian nor first-principles, and 'don't confuse motion with progress' is a cliché the host himself introduces.

don't confuse motion with progress
companies that define their measurements of success, uh, metrics before they invest and deploy an AI initiative, typically will have the biggest win

Guest Caliber

8 / 20

Ray Reich has genuine practitioner credentials - 20+ years in recurring revenue roles, COO experience, multiple exits, and founding the SaaS Metric Standards Board - but the format is a co-hosted book-promotion episode, not an independent interview, which limits what credibility can deliver.

I was a recurring revenue operator for 20 plus years
I actually founded, with four industry colleagues and peers, the SaaS Metric Standards Board

Specificity & Evidence

12 / 20

The case studies carry the episode's evidentiary weight, with real company names and concrete numbers (Petrobras's $120M savings in three weeks, the insurance firm's 60-to-10-minute claims drop, Uber exhausting its annual AI token budget by April, Klarna's revenue-per-employee rising from $575K to ~$1M); however, the opening statistics are unsourced and the framework discussion stays abstract.

$120 million in tax savings in about three weeks
claims processing time dropped to 10 minutes from 60 and 83% reduction. Their underwriting cycle time dropped from three days to three minutes

Conversational Craft

4 / 20

This is transparently a co-promotional vehicle: the host lobs consistent setup questions ('What does the book give them?', 'How's it different from what's out there?'), never pushes back on any claim, and at one point Ray asks Peter to take over narrating case studies because Ray has been talking too much - revealing the scripted, rehearsed nature of the exchange.

Right. So how's it different from what's out there? Because there are lots of frameworks out there.
Peter, I know I always say this, but I'm so passionate about this topic

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A65%
  • Speaker B35%

Most-used words

metrics36book18revenue16peter15baseline15framework14measurement11process11first10case10cost10today9measurements9start9three9operating9

Episode notes

The AI to ROI team, Ray Rike and Peter Buchanan, mark the official launch of The Big Book of AI Metrics, an 180-page, 81-metric operator's reference guide built to close the gap between AI adoption and AI ROI. Twenty-seven percent of executives say AI has met their ROI expectations, enterprise AI token spend is up 13x since last year, and most companies still can't explain what they got for the investment. Ray and Peter break down why that gap exists and what to do about it.

Full transcript

34 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Foreign.

Speaker B: Welcome to AI to roi the Big Story. I am Peter Buchanan with the new plan and with me, as always, is Ray Reich, the CEO of, uh, Benchmarket. And Ray, today is a big day for us.

Speaker A: It really is, Peter. We've been building towards this for quite a while. You've been prodding me to get this thing finished. And this thing is, we released the big book of AI metrics today. And I want to talk about why we build it, who it's for, and how we think executives and leaders should use it.

Speaker B: Okay, so before we get in the book itself, let's frame, uh, why this moment matters and why the book is sort of needed. So the data related to AI implementations and ROI is not kind. 27% of senior executives say, uh, AIs met their ROI expectations. The average monthly AI token spend across enterprises has risen 13x since the beginning of last year. And most enterprises can't explain what they got for it. So it seems like enterprises and AI native product companies need some help.

Speaker A: And you didn't even mention that famous now six month old MIT report that said 95% of AI projects are driving.

Speaker B: No, I know, uh, I had that. You know, I had that in the report and I took it out because, you know, it's so last year, Ray.

Speaker A: I know it is. Well, six months is definitely a lifetime ago. You know, it's funny, there's this pattern that connects the successful AI initiatives and the failed ones. And it's almost always the same or at least grounded in a similar basis. And that is the companies that define their measurements of success, uh, metrics before they invest and deploy an AI initiative, typically will have the biggest win. So a lot of companies today, unfortunately, they're celebrating adoption and they're calling that success. So we wrote the book to close the gap between true success as measured on the income statement and balance sheet, and AI adoption.

Speaker B: Right, that's the thesis. So let's get into it now because for us it's quite exciting.

Speaker A: Yeah, it is exciting, but I think we need to hear a word from our sponsor first.

Speaker B: Oh, that's right. We, uh, like sponsors. Sponsors are good. Okay, Peter, let's dive into Ray, give listeners an overview. What exactly is the big Book of AI metrics?

Speaker A: Well, it is comprehensive, I think it's what, 180 pages and 81 different metrics. So it's comprehensive, it's very practitioner, operator, focused reference guide. And it starts with an AI metrics and measurements framework for delivering successful AI use cases. Then we organize the metrics by use case Type by function and even by stage of the AI lifecycle.

Speaker B: Right. So how's it different from what's out there? Because there are lots of frameworks out there. Gartner Frameworks, McKinsey Playbooks. What do we do that's different?

Speaker A: Well, there was a lot of pontification that goes on social media sites like LinkedIn or even conferences about here's how you should use AI, here's a framework of how you should consider AI. They tell you that measurements matter, but they don't go into any level of detail on what to measure, how to define it, how to ensure you have a baseline pre AI, uh, versus the post AI measurement and even what a good result looks like. So we built this big book of AI metrics for operators and practitioners, not consultants and influencers.

Speaker B: Right. So we've been publishing periodically the AI Metric of the Week in the AI to ROI newsletter. We've been doing it for months. Is this sort of those metrics compiled and organized or is it more?

Speaker A: Well, you know something, I thought ever since I founded benchmarkit was laser like, focus on one role. So a lot of the metrics that we kind of introduced in our newsletter, one was for US revenue leader, another one was for a product leader, a couple were for CFOs. But what I wanted to do is bring together one consolidated integrated approach and almost provide a systemic orientation of who, which metrics for which function, for which stage of, uh, maturity in your AI initiative. So it's really an integrated systems thinking type book about metrics, starting with a really practical framework.

Speaker B: Right. So you probably had a motivation for doing this because you spent your career as an operator, you in serious go to market roles, you've been a coo, you've had a CEO, you've had multiple exits, you've had benchmarket for the last six years, which does metrics and benchmarking based market research and builds Measurement framework, helps B2B software companies build measurement frameworks. So what did you see happening in AI that looked familiar? And also, if you were an operator and this appeared, the big book of Metrics appeared to you six years ago, 10 years ago, and you're a CEO, how would you use this? So let's just go through a little bit of history and then let's just say, how would this affect your life as a senior executive?

Speaker A: Okay, for those of you on your elliptical, taking a jog, taking a walk on your bike, this is probably one of the two most important parts of the episode. So I was a recurring revenue operator for 20 plus years. And one of the things I always wanted to do was benchmark how some process reengineering changes I made impacted things. And I'm like, well, cool. In the SAS industry we have some standard metrics. I can benchmark against annual recurring revenue, net revenue, retention rate, my CAC payback period, the rule of 40. But Peter, even then, when I started doing benchmarking in 2020, I quickly realized there wasn't a standard definition even of something as basic as annual recurring revenue. So that's why I actually founded, with four industry colleagues and peers, the SAS Metric Standards Board. But Peter, I didn't do that until 2021, 20 years after SAS started. So I said let's get ahead of this this time with this whole AI phenomena and basically revolution and let's define metrics today. Let's start having people debate, discuss and at least leverage them and see if we get can get some industry standardization or at least some guidelines that when we want to benchmark, our AI agent for customer service or customer support is performing against the industry standards. So that's why I did it.

Speaker B: Right. So we talk a lot about these missing definitions and you've had benchmark it here. What's happened? Because there's some parts of SAS metrics that apply to AI. What makes adding AI on top of this makes this effort more urgent in every way, both for enterprises and for product companies. Well, what's the layer cake here that gets added on top?

Speaker A: It's the namesake of this podcast in our newsletter. So, first of all, over the last almost three years, companies have progressively adopted more and more AI in their organization. The costs continue to evolve and a lot of the benefits have stopped at personal productivity gains. My salespeople can save X hours per week doing research, or we can now reduce the time it takes our sales development reps, um, to personalize email messages. Or we can develop lines of code much faster. But productivity in and of itself doesn't always translate into how much more revenue can we do, how much can we reduce operating expenses. So then we went from personal productivity to, well, let's at least start measuring usage. And usage may have been something of how many people are using that cloud license we bought, how many people are leveraging cloud code. And then we got into the last six months, how many tokens are we consuming? In fact, we even encourage people to maximize the number of tokens being used. Token maxing and, and in one case that resulted in a $500 monthly bill.

Speaker B: $500 million. Million dollar monthly bill.

Speaker A: 500 million?

Speaker B: Just a little bit more than $500.

Speaker A: Yes, 500 million. And by the way, a lot of the things we're measuring are activity based, like pull requests and development or emails delivered. And hell, that doesn't even really always increase productivity. And it's almost never directly correlated or have a causal relationship to benefits that impact the income statement.

Speaker B: Right. So I have this phrase I use with a lot of clients. Don't confuse motion with progress. So motion or activity. And a lot of the things you're talking about that are activity, they're individual activities or the things that you can definitely measure. And it looks like more things are going on, but it doesn't necessarily, it doesn't mean those measurements and metrics, that there are the right. Are the right metrics. So there's a measurement failure. Uh, do we have a measurement failure? And is that sort of a symptom of larger problems with AI? Is it something deeper? How do you really think about this?

Speaker A: Well, I think bottom line, it's structural. Uh, companies were never designed to truly understand how this technology investment AI is impacting my M income statement, my operating expenses, my cost of goods sold. Hell, sometimes they're not even measuring the existing cost of a particular process. So that's why I say it's structural. Now, initially, everyone knew AI was really strategically important. So a lot of companies said, well, we just need to get our employees familiar with it. So adoption was not a bad starting point. Right. Because it allowed them to say, hey, um, I'm at least introducing to my organization because I know it's going to have huge impact. Utilization is the next step. By the way, a lot of people who said, I'm going to buy ChatGPT for all my employees, they're not really measuring the utilization. Hey, is that $20 or $200 in the Pro version? Am I getting benefit out of it? But adoption still is not roi. Outcomes. Outcomes are ROI and outcomes that translate into better financial performance. That's true ROI that a cfo, uh, investor, and a board of directors can get behind.

Speaker B: All right, so I can definitely be a believer in that. But now let's talk about who we wrote the big book of AI metrics for. Because there are a lot of roles across the company. In fact, when we have the, the very long chapter with all the metrics, it's basically broken into 13 different roles. Right. So let's talk about who we wrote this book for and then how they use it.

Speaker A: Yeah. Once again, um, this is one of the problems I've had ever Since I founded Benchmark was can I really narrow who we can try to add value and benefit for? And once again I felt that. Peter so we have two broad kind of segments. Number one are operating executives who are investing in AI to improve the performance of their function department and company. From engineering and software development. They have very different metrics they need to measure than the go to market team. Right. Engineering may be looking at code quality, production stability, production performance, developer productivity. Right. Go to market teams. Maybe they're measuring pipeline velocity. So first of all, we had to create one for every department. Finance. Yeah, they know that pipeline velocity is important, but they're really more concerned about how does that impact how many new sales or revenue growth I can get from my existing operating investment. So that's one segment, all those operators. But the second and almost half of our audience today, Peter, are B2B software executives, traditional B2B SaaS, AI native. And this is impacting how they go to market, measure their gross margin and at the end of the day their enterprise value multiples. So they need to understand things like, hey, what is the token consumption in my product? What's my inference cost of tokens consumed to the revenue generated? What's my inference cost to the value of the outcome delivering, delivered to my customer? So they have a little bit different need. And we wrote a newsletter, maybe this was about two months ago, where the cfo. The CFO is such a critical member of our audience because AI spend is going to become, if it hasn't already, a compensation scale operating expense where human labor may be 20 to 40% of total operating expenses today. I'm sorry, not operating expenses of revenue. AI is already reaching that, especially in AI product companies. So we really need to have something for that audience. Peter.

Speaker B: Right, let's go through. Let's talk about the enterprise executive. They're deploying the use case, say AI and customer service. What does the book give them that they don't already have?

Speaker A: Well, one thing is a, uh, framework. And by the way, frameworks are great for ideation, but maybe your company is going to modify it slightly. But we have this AI performance metrics framework that really focuses on having causal relationships between all five levels. We call it causal chain. So one is input signals. These are things that maybe are being impacted by the use of AI. Maybe it's a piece of equipment, sensor, etc. Those are signals. This is. Then you have your leading indicators. Oh, uh, I'm getting 25% more signals today that are being addressed through AI. Then we have your Operational metrics. Think of these as functional KPIs, something as simple in sales. What is my win rate? What is my conversion rate from a lead to an opportunity and supply chain? It's something like what is my inventory carrying cost or inventory turns. Then fourth, we have financial outcomes. These are the things that show up on an income statement or a balance sheet. And then fifth, and it's really the holy grail. And if you think about a value pyramid, it's the top of the pyramid. It's the strategic value that your AI investments are adding. It might be your competitive positioning, maybe it's market share, it might even be your enterprise valuation multiples. But the, the underlying philosophy in this framework is the same. At the end of the day, we want to go beyond measuring, uh, events, task activities, and measure outcomes.

Speaker B: Wow. Okay, so let's walk through what good looks like for an enterprise. The enterprise is starting a new use case. What would they do first? According to the book, let's go through the, like the yellow brick road to AI, to roi.

Speaker A: Well, we have the framework which is really trying to define your measurement system for your AI investments. But at the same time, you've got existing business going on today, right?

Speaker B: Sure.

Speaker A: So what I find, unfortunately, in this world of too much data, is we may not have the right measurement and metrics in place to truly understand the current cost of the baseline process. Maybe the process is generating a lead in sales. Right. So first of all, we need to establish the baseline. Then we need to understand what metrics we're going to use to measure and benchmark that baseline. And then we need to have a framework or an instrumentation for measuring the results. Post AI implementation. And that's a continuous, virtuous cycle, Peter.

Speaker B: Right. So what changes once the use case goes from this implementation? You have the baseline, you move to your implementation, you're beginning to track these metrics. What happens? You've got your use case in the wild here. People are using it, agents are off doing agentic things. What happens when it's in operations?

Speaker A: Well, at the end of the day, if people use this framework and this kind of five layers of causal relationships from AI investment to financial outcomes. The CFO gets governance and financial management metrics, things like cost per outcome. In a post AI world, token spend as a percentage of cost of goods sold or operating expenses. And the thing we haven't talked about, and sometimes it can be a dirty word, but attribution, right. One of the key things you want with the before and after measurements is to at least have A correlation, if not a causal attribution to the AI use, to the specific business outcomes and financial performance improvements. That's the language that unlocks continued investment in AI. Because without it, these budget allocation decisions towards AI becomes more of a religious faith based discussion and not one steeped in empirical evidence.

Speaker B: Right. What do you think the most common mistake in the operations phase? You've got your stuff out in the wild here. You think you've built a good application, you put your, your metrics in play, you did your baseline, you followed it all the way through to release, you got it in operations. What mistakes? What's the biggest mistake companies make once they've released their use cases?

Speaker A: Well, I don't know if this is going to answer your question, Peter, but it really does start with do you know what your baseline is? And then what the current state is and exactly how you're going to measure it. And then the biggest gap I almost always see is that measurement stops at an operational KPI and is not linked to the financial performance outcome. That's the biggest chasm I see. And by the way, this is not just for AI programs, this is for a lot of software investments over the last 20 years. I cannot tell you, Peter, how thorough of ROI justifications I had to go through with companies to get them to pay a million dollars for software. And I can count on my hand the number of my customers actually went back in a year or two years later and validated the return on investment thesis. We can't do that with AI because AI isn't 1 to 2% of revenue the way it spend is. It might be 30, 40% of revenue. Like we said, it's compensation scale cost that you can't be subjective with.

Speaker B: Right. So they deploy, they get the initial win, they report it up the chain and then um, the measurement cadence goes dark. And six months later their model performance has drifted like the application's getting worse, usage is plateaued. There's uh. And no one noticed because no one was measuring it. So I think the lesson here is once you start, you get your metrics and you start measuring, you're doing it forever.

Speaker A: Yeah. And you just brought up an important point and that is if you don't have the baseline, someone can come back after they rolled out these AI tools and said we are doing so much better. It's like against what baseline? Against what commonly agreed upon baseline. Right. So you really can't prove that AI had a causal impact on the improvement. And in fact, I kind of say the baseline is the Control group, if you skip it and don't know how before and after does, you have an opinion, not evidence.

Speaker B: Right. And also you're going to be adding, you may plan for it in your baseline, but once you're operational, you're actually going to be start adding new metrics you weren't having before. Right. Well, you have to anticipate that.

Speaker A: What I'm seeing happening is you develop this agent and then people realize, oh, this spawns off this other process. So we create sub agents. Right, Right. So now you're not really understanding how that overall agents impacting all the sub agent interrelated processes. And we do see M model decay over time. So what how your agent was performing in June of 2026, don't be so sure that it's going to keep improving because you may get bad data, you might get noise, you might get. There still is. Nobody talks about as much hallucinations, Those all still exist. And without proper human audit reinforcement and retraining, you're not necessarily going to have a continuous improving AI framework.

Speaker B: Hey, you forgot context, Ray. We love context. You have to give those agents context.

Speaker A: We do. You need to do that up front. But that context will continue to evolve.

Speaker B: Yeah, exactly. All right, so now we're going to get specific. We've gone into the, the book, but of course we've been studying some of these successful and unsuccessful, uh, AI use case deployments. Let's just go through a couple of them and apply them to the way we've done it in the big book. So let's start with Petrobas.

Speaker A: You know, Peter, I've been doing a lot of talking during this episode. Do you mind taking on and sharing a couple of those for us?

Speaker B: I can. Um, so Petrobas is a big, um, Brazilian oil, uh, company. They are the state energy company. They process, ah, complex tax regulations and quarterly tax data manually. That's what they were doing. They did it for 15 years running. It was tax season. The tax team worked weekends long hours. They fed 150 pages. So they said, well, there's got to be a better way. So they fed 150 pages of tax regulations and three months of data into generative AI models. Um, built on automation anywhere, the product automation anywhere. They have found $120 million in tax savings in about three weeks.

Speaker A: That shows up on an income statement.

Speaker B: It does. Uh, so they filed taxes, uh, in three days they had their first weekend free tax season in 15 years. And those Brazilians love to party. So I'm sure they had a Good time. Why it worked? Well, the first thing is it was a narrow scope. They knew what they wanted. They obviously had a baseline which was uh, all of that work and they had definitely objectives. They had a text heavy rule bound workload which worked very well with AI. They had verifiable financial outputs. They had clear, uh, and before and after method. They had the Rayright baseline. It's almost like they were feeding our book. And so they project more than a billion dollars in savings as they expand this model to more areas of the tax system. So tax savings identified filing time, labor hours, all that stuff.

Speaker A: Those are measurements any CFO can get behind. But let me share another one. And the company is, um, it's an insurance operations company called and so before they were really looking at process improvements for claims processing, it took 60 minutes of human per claim. They also had a very complex underwriting review process that took three days. So they deployed actually Microsoft Copilot to 3,000 employees across 14 countries. And their measurements, it got to the third level of our AI performance metrics framework. It was operational KPI improvements. Their claims processing time dropped to 10 minutes from 60 and 83% reduction. Their underwriting cycle time dropped from three days to three minutes. Now what this case study didn't say was how that impacted either operating expenses or revenue. But what mattered to them initially was getting that operational KPI of uh, cycle time by process type reduction, their employee coverage and critically the quality of the output, the right underwriting decisions. And those benefits are going to be measured over time to see what the loss ratios are for those new customers leveraging this new underwriting review process.

Speaker B: Right. So here's one that didn't work. Uber. Uh, they have sort of an engineering department extremely cautionary case. So they deployed Claude Code to 5,000 engineers at the end of last year. And lots of things went wrong. First of all, uh, what seemed to be very good in the beginning is the usage of agentic coding features dumped from 32% of engineers in February when they first deployed it to 84% by the end of the next month. So by the end of March this year. But unfortunately at the same time they weren't tracking costs and quality of output. So by April they had exhausted their annual AI token budget. So that's definitely a failure. So none of these usage increases actually translated into an increase in consumer features shipped through, through their applications. So they measured adoption and um, they measured adoption, they rewarded token consumption. They had an internal leaderboard. But adoption isn't where the value is created and so now, of course, they put their engineers on a 1,500, um, dollars per month token budget as a result of this, which may or may not be the right thing to do. So what should have been measured? Well, complexity adjusted, code quality. Are they getting what they really want out of this particular adoption? Is their code they put in production stable, and are they producing features at a rapid rate that are both high quality and stable? So not token spend or pull requests, but the impact it actually has in delivering things to customers, in creating a better application.

Speaker A: You know, Peter, I know I always say this, but I'm so passionate about this topic. But we're already coming up to the end of our show for the week, but I want to just do one quick highlight of one more company, then we'll get into our summary, and that is Klarna. Now, Klarna made a lot of news about 18 months ago when they said that they were using Vibe coding to replace their implementation of Salesforce. And everyone said, no way, blah, blah, blah, blah. But let's look at some more macro level kind of attributes of Klarna. Um, before, um, AI kind of Q1, 23, they had a workforce of 5,000 employees and a revenue per employee was about 575,000. They very systemically deployed AI across customer service, underwriting, internal operations, and they ended up with a workforce of 3,000. So if you think about that, that's a 40% reduction. Their revenue per employee rose to nearly 1 million. Right. So that's increased from 575,000. So to me, that's, you know, not quite 100% increase, but a big increase. And the metric that mattered to them was revenue per employee. Now that showing up on your income statement, it sure showed up in their enterprise value. They also focus on some of those leading indicators, customer resolution rates. They did look at AI usage. But most importantly, Peter, uh, we got to wrap that up with is they focused on the ultimate outcome metric, revenue per employee, not on AI utilization as their core metric.

Speaker B: Right. Okay, so we're sort of at the end here and we're going to go with the big finish. So if you're an operational executive and you're investing in AI, whether you're in an enterprise or you're building an AI native product, um, what should you walk away from this episode with? What are the three things you should do this week coming out of this episode?

Speaker A: Okay, so first, and this is really, some people are going to say, ray, this is like measurements 101. But believe me, I've looked a lot of companies AI projects and they haven't done this. So number one, define the outcome metric that you're ultimately going to be able to use to justify the continued investment in your AI initiative. Not an activity metric, an outcome metric. Second, establish a baseline. Know what that process or activity, know exactly what the cost inputs are today, how much it's costing you and just start there. And then third, after you justify the AI investment, you deploy it both in a proof of concept or pilot and then in a production. Make sure you have the measurements and measurement cadence and a project plan as a formal deliverable and that you're providing status updates on exactly how that performance improvements are trending over a month, a quarter, a year.

Speaker B: Right.

Speaker A: Those are three things.

Speaker B: Yes. And the Big Book of AI Metrics has both the specific metrics and the framework for you to actually do that if you're an a AI focused executive.

Speaker A: Yeah. And honestly, this is just trying to jump start the conversation in your company. Right. It's not a strategy document, it's an operator's tool. Pick it up, find your use case. If you don't like the metrics we have in there because you have more nuance in your industry, just borrow from them and add to them because there is no such thing as a single book of AI metrics because we can't capture all those nuances.

Speaker B: Right. All right, time for us to go. So thanks everyone for joining us in this episode of AI to roi, the Big Story edition. The big Book of AI Metrics is available now on the Benchmark website. Benchmark the letters it AI. There's a link in the show notes as well and also links on both our, um, our LinkedIn pages. And there's a great article, uh, that sort of summarizes how to use this on the AI to ROI newsletter substack. So you can find it in a whole bunch of different places.

Speaker A: And for anybody who wants to talk about it, pick it apart, add to it, reach out to me on LinkedIn. Ayreich, that's just ayrike because I want this to be a living, breathing and evolving, valuable document that when we look back on five years, we can say that our measurement framework and infrastructure was a big part of our AI success story. Thank you, Peter.

Speaker B: You bet. See you next week.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Developers Hit a Wall at 4 AI AgentsThe AI Native Dev · on Token Consumption90 / 100
  • Bad Data, Dark Data, & Why Your North Star Keeps MovingMaking Sense of Martech · on ROI measurement77 / 100
  • Monetizing Expertise: Kate Strachnyi on B2B Influencer MarketingThe Business Of Marketing · on ROI measurement70 / 100
  • Episode 374 Deep Dive: Mark Jones | The Department of No Is Over - Why Cyber Has to Lead AI AdoptionKBKAST · on Token Consumption63 / 100
  • Making AI Work In Product Teams: Roundtable InsightsProduct Rebels · on ROI measurement54 / 100
  • Getting to ROI with Experiences - with Jennifer Gardner, Creative Director of Event Experiences, ServiceNowReal Creative Leadership · on ROI measurement53 / 100

More from AI to ROI

All episodes →
  • Measuring the costs, utilization, proficiency and impact of AI - with Russ Fradin, Founder and CEO, Larridin
  • Leveraging AI to Reduce Churn and Increase NRR - with Dan Harmeson, Co-Founder and Co-CEO at QuadSci
  • The AI Agent Outcome-Based Pricing Journey - with Kunal Agarwal, CFO Gorgias
  • AI to ROI: OpenAI - The Most Important AI Company in the World, and the Most Fragile
  • NVIDIA - The Full-Stack Maestro
Explore the best B2B AI & Data podcasts →
All AI to ROI episodes →