The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Customer Success/Unresolved.cx
Unresolved.cx artwork

Unresolved.cx - Resolution means something different at every company - Craig Stoss - KODIF

Unresolved.cx · 2026-06-30 · 37 min

0:00--:--

Key moments - from our scoring

Substance score

64 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality12 / 20
Guest Caliber14 / 20
Specificity & Evidence14 / 20
Conversational Craft11 / 20

Craig Stoss of KODIF, an agentic CX platform for commerce, explores a critical tension in AI-powered support: the expectation that AI should solve problems faster than humans, versus the reality that good customer experience depends on meeting each company's specific needs. KODIF builds AI agents to automate ticket resolution for support teams, and Stoss practices what he preaches by using AI heavily within his own six-person, globally distributed Solutions team. He details how they've built internal tools like an AI debugger in Slack (using Python and Claude) that cuts hours from technical debugging, and a SQL optimization tool that reduced report load times from 5 minutes to 3-4 seconds. The conversation challenges the idea of a universal resolution metric - what's "fast" for warranty requests is wrong for elderly care support. Stoss introduces the Customer Frustration Index as a measurement framework specifically for AI agents, capturing issues like loops, repeated requests, and contradicting information that wouldn't occur with human agents. He advocates treating AI agents and human agents as a team working toward the same outcome, with transparently different metrics, rather than pitting them against each other. The episode covers HubSpot, Google Suite, VS Code, Claude, Gemini, Postman, and DB Beaver as part of KODIF's stack.

Key takeaways

  • →AI agents need transparent goals and measurable ROI similar to human agents, with different metrics appropriate to their capabilities and limitations.
  • →The Customer Frustration Index - measuring loops, repeated questions, and contradicting information - reveals AI-specific pain points that don't apply to human-agent interactions.
  • →Internal AI tools like debuggers and SQL optimizers can dramatically reduce operational overhead and improve implementation timelines for your own team.
  • →Good customer experience is defined by what resolution means for each specific company and demographic, not by a universal "solve 90% of tickets" benchmark.
  • →The emerging skill set for AI-era jobs isn't replacing humans but validating inputs, crafting precise prompts, and expertly validating outputs.

Guests

Craig Stoss

Topics in this episode

AI agentsGeminiClaudeHubSpotPostmanGoogle SuiteVS CodeKODIFDB BeaverCustomer Frustration Index

Questions this episode answers

What metrics should you track differently for AI agents versus human agents?

While human and AI agents should share overall customer experience goals, AI agents require unique metrics like the Customer Frustration Index - measuring loops, repeated requests, contradicting information, and customer frustration signals - that rarely occur with human agents and directly improve AI training and prompt effectiveness.

How can small teams justify the effort to build internal AI tools?

KODIF's AI debugger in Slack (using Claude and Python) cuts hours daily from technical analysis, and their SQL optimization tool reduced report load times from 5 minutes to 3-4 seconds, demonstrating rapid ROI for tools that reduce repetitive technical work.

Should AI agents be expected to resolve 90% of support tickets?

It depends on your specific customer base and use case; KODIF targets 70%+ resolution as success because nuance varies by integrations, and the real measure is whether AI provides solutions as fast, convenient, and empathetically as a human would.

How should support leaders think about AI and job security for human agents?

Position AI agents and human agents as a team with transparently different roles and metrics; AI handles high-volume repetitive work (3am warranty requests) while humans focus on empathy-driven issues, reducing anxiety and clarifying that AI augments rather than replaces.

What tech stack enables internal AI automation for CX teams?

KODIF uses HubSpot, Google Suite, VS Code, Claude, Gemini, Postman, DB Beaver, and Python scripts connecting to their database for debugging and reporting - core tools for building token-efficient AI integrations and leveraging LLMs for data formatting and SQL optimization.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode contains several substantive frameworks (value-based metrics vs. service-based metrics, Customer Frustration Index, the 'assist' concept) and concrete examples (SQL debugging tool cutting report load times from 5 minutes to 3-4 seconds, the emoji-triggered Slack debugger), but significant portions consist of tangential anecdotes (parking app story, son's Scratch game) and repetitive philosophizing about AI and jobs that don't advance operator knowledge. For a B2B operations leader, the metric definitions and measurement philosophy are useful but not densely packed.

We've had a custom report that used to take 5 minutes to load got down to 3, 4 seconds because of just different efficiencies that were found in the SQL statements
I call it the Customer Frustration Index. Is the customer saying, you're not listening to me, you don't understand me. Is the AI asking them to repeat information they've already given?

Originality

12 / 20

Craig articulates some original frameworks - particularly the Customer Frustration Index and the distinction between assist/deflection/resolution, plus the idea of transparent AI agent metrics parity with human agents. However, much of the broader thesis (AI won't eliminate jobs, clear prompting is key, humans need more meaningful work) echoes mainstream Silicon Valley thinking. The frameworks feel fresh for a CX context but lack deep contrarian challenge or first-principles rigor.

AI needs to be treated as agents. They need to be transparent with what their goals are, what they, what are they supposed to achieve.
resolution is, ah, actually a hard thing to measure...from our perspective, we incur the cost, even though potentially the thing wasn't resolved.

Guest Caliber

14 / 20

Craig is VP of Solutions at a venture-backed AI vendor with hands-on operational experience building internal tools (the Slack debugger, SQL optimization engine) and managing a 6-person distributed team. He brings practitioner credibility in CX ops and AI implementation, though his perspective is inherently vendor-side. He's not a category-creating operator or scaled founder, and the company appears early-stage, which limits the depth of large-scale organizational insight.

So Codif is a small startup. We do wear a lot of hats...my team handles the renewals
The AI debugger I've mentioned a couple times has been absolutely fundamental to how my team operates...that's by far the tool that's probably helped us the most.

Specificity & Evidence

14 / 20

The episode includes concrete examples: the Slack emoji debugging tool, SQL report optimization (5 min → 3-4 sec), warranty form automation, 70% resolution targets, and the Customer Frustration Index. However, many claims lack supporting numbers - e.g., how many customers use the tool, what the actual ROI was, what specific failures the Frustration Index has surfaced. The host's Intercom deflection example (60% without AI) is specific, but Craig's discussion of metrics definition remains somewhat abstract and untested.

a custom report that was created that used to take 5 minutes to load got down to 3, 4 seconds
we could probably get somewhere into that range and could that be provided faster...If we were able to solve anywhere north of 70% I would say would be a success.

Conversational Craft

11 / 20

The host (Brett) asks solid setup questions and introduces useful frameworks like the assist/deflection/resolution distinction, but rarely pushes back on Craig's claims or challenges his assumptions. When Craig meanders into personal anecdotes (parking app, son's game), the host doesn't redirect. There's no productive disagreement, no hard follow-ups on measurement nuances, and the host often validates Craig's points rather than testing them. The conversation feels collaborative and friendly but lacks the sharpness needed for rigorous business education.

Yeah, I love the term assist, actually...that's exactly what's happening, right
Craig, thank you so much. I appreciate you taking the time to speak with me today. And I love that you're pushing the boundaries

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B70%
  • Speaker A30%

Most-used words

human31customers24customer23different20experience19metrics18measure16team15value14agent13problem12resolution11example11humans11agents11world10

Episode notes

In this episode, Craig Stoss talks about how the industry needs to solve the AI Resolution Rate. It means something different for every vendor. So many Customer Experience leaders are testing with AI and rebuilding with AI as their foundation. Every one of us have to reinvent everything that we’ve known about customer experience systems and processes. You’re not alone, so let’s do this journey together. If you’re working to solve AI resolution, or you’re closer to building it than he is, he wants to hear from you. That’s what Unresolved is for. - Follow the people: - - - Resources: - Podcast episode - CX-News - CX Advisor

Full transcript

37 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Welcome to Unresolved, the podcast for Customer Experience Leaders by Customer Experience Leaders. In this podcast, we discuss how different companies view and measure their customer experiences, how they manage their teams, and we dive deeper into the unresolved issues that they're facing today. And we're specifically calling this podcast Unresolved because so many CX leaders are testing with AI. They're rebuilding with AI as the foundation. And every one of us have to reinvent everything we've ever known about customer experience, systems and processes. You're not alone. So let's do this journey together. Today we have Craig Stoss, the Vice President of solutions for codive, an agentic CX platform built for commerce that helps support teams automate ticket resolution, surface actionable insights, and continuously optimize operations. And I had the opportunity to Meet Craig in 2023 while I was speaking at the Philly Support Driven Summit. And for those that haven't been to a Support Driven conference, look up one near you and just go to it. Because you'll meet incredible, thoughtful and selfless people like Craig, and you'll stay in touch with everyone that you meet. And I promise you that, Craig. Welcome to Unresolved.

Speaker B: Uh, thank you so much for having me, Brett. Great to be here.

Speaker A: Of course. So I like to start off by creating a baseline so that folks can connect with you way. So can you tell me about your team, the team size, the structure and what you're responsible for?

Speaker B: Yeah. So Codif is a small startup. We do wear a lot of hats. So in our world, Solutions encompasses, uh, pre sales, uh, demos and creation of those demo environments, uh, includes custom integrations for signed clients. It, uh, includes the implementation team, ah, who does the actual white glove, uh, implementation of the AI services. And, and uh, it includes support, um, and success in the sense that we support our existing customers, uh, and my team handles the renewals, uh, and the other conversations around general client value along the way. So team, um, size, I mean, all of that fits within six people on my team and we're all remote. So I'm based near Toronto, Canada. Uh, I have members of my team in Mexico, Brazil, uh, Kyrgyzstan, in the United States, uh, Honduras as well.

Speaker A: When you think about customer experience, what does good customer experience look like in your world and how close do you feel like you are to it?

Speaker B: Yeah. So, you know, it's funny, as an AI vendor, there's certain, um, you know, I'm sure you've heard the phrase eat your own dog food or drink your own champagne. Right. And I think that in the AI world especially there's this sense of you're telling us to use AI, you're telling us the advantages of AI. Show us how do you use AI, Right. I think that's a really critical question. Um, and just like when you have for example a model home, if you're building a new group uh, of housing or if you're going to uh, a wedding venue, they sample the foods and the cakes that you're going to have on your wedding day, you know that's going to be the highest quality presentation, right? Because that's what they're trying to sell you on. They're not going to, they're not going to, they're not going to show you a house that doesn't have all the upgrades and the fancy bits that you could have, for example. And I think that's true in our world, right, is that if we're going to put out AI out there, it needs to be at the highest quality and executing on all levels. And so for me, good customer experience has two components. One, one is always how people care about delivering that service and making sure that the customers access that service when they need it, whether it be in a different time zone or during a critical situation. And then the second piece is specific to the AI world is I think using AI to your advantage and being able to use um, not only the AI technology that we sell, um, but I think it's using AI in general. I mean we've had for example when someone leaves the company, we've had customers say oh, you know, can you maybe use AI to scan their emails so that we, you know, you can catch up to speed on how this project is going, for example. And, and yeah, I think that's the expectation is, is that if we're, we're purporting this technology as has an roi, we should be using it. So how close are we to it? Not as close as I would like to be. I think there's a lot of cool projects. I think it's a, it's time and people, uh, but we are moving in the right direction. You know, we've just launched a uh, really cool report, uh, that is auto generated every week using AI and delivered to our customers for their success and being able to measure how their implementation is trending. Uh, we use AI to do debugging, uh, whenever there is a problem or something that doesn't look right inside a conversation. The customers can actually uh, hit an emoji in Slack and we'll actually auto debug. So we're getting there slowly. I think there's a ways to go to bed fully uh, operational in that high end customer experience I'd like to deliver.

Speaker A: Yeah, customers actually want more human experience whenever it comes to support uh, versus leveraging AI as much. And so I was curious, whenever you're thinking about the future state, would it be ideal that codif the AI agent on your site, uh, and also inside the platform itself would be able to handle 90% of uh, of conversations and solutions and actions for customers?

Speaker B: The preface of your question is really interesting, right? Because I do think there is an element of customers do want the human experience. I also would argue that that part of it is, is they just want solutions fast. And so where I draw the connection between you know, wanting a human experience and can I solve 90% of the, of the use cases is is that AI experience as good or as fast or as convenient as it is with a human? Um, and that to me is the measuring stick. So should it solve 90%? I mean in our world our stuff has you know, a lot of nuance depending on what integrations you, you have and other things. Uh, but we could probably get, get, get somewhere into that range and could that be provided faster and more conveniently and as uh, empathetically et cetera as a human again we probably could get close to that a year from now. If we were able to solve anywhere north of 70% I would say would be a success.

Speaker A: Before we dive too deep, uh, can you share what your tech stack is that your team uses day to day?

Speaker B: Yeah, so I mean codif again small company, uh, we do have HubSpot. We are a Google suite company so we, we do spend a lot of time in Google sheets and those types of tools. Vs Code and Claude. Um, I do make use of Gemini. I think Gemini has some really interesting features when it comes to uh, being able to for example access websites to be able to pull out things uh, that help us customize demos and make demos more personalized to our prospects. The main tools that my team use, beyond the scripts that we've used AI to write to um, to do some of these debugging and reporting ah are just database uh tools. Postman for APIs, DB Beaver which is a postgres uh tool um, and things like that that allow us to access our backends and be able to help debug common problems.

Speaker A: Gotcha. Great, thanks for that tech stack. So now I want to kind of move into like AI reporting metrics and you have such a great history And a story here that I want to talk about the good, the bad, the ugly with AI. So my first question is like where has AI actually helped your team?

Speaker B: I mean first and foremost I think the AI debugger I've mentioned a couple times has been absolutely fundamental to how my team operates. We built a tool uh, and as I mentioned it's built into Slack where we and or our customers simply have to add an emoji to a comment that has a conversation id, which is our unique identifier for every conversation as well as the uh, problem that was experienced. And it will go and do all this debugging for you and it will present you a report of what happened in the conversation, you know, what, what, how we understood the concern and then how we think it should be fixed. And this is cut hours a day off of our, my team's technical analysis. So that's by far the tool that's probably helped us the most. Um, and that is a combination of Python, uh, connecting to claude, connecting to our database, munging the data in token efficient ways, et cetera. Um, and it's been a great tool, all vibe coded like, all coded through a VS code uh, like tool. But then we do a lot of report generation um and uh, we do a lot of our uh, metrics. For example we have a custom report engine which we know sometimes uh, when these reports are built uh, the queries may not be as efficient as they could be. Um and so we actually have an AI tool that will export all those queries, run them through a prompt that does SQL, ah, checking for efficiencies or way to improve the performance of it. And we've seen some massive improvements there where a custom report that was created that used to take 5 minutes to load got down to 3, 4 seconds because of just different efficiencies that were found or inefficiencies that were found in the SQL statements run. So um, that's a big win for our customers obviously but we do a lot of this. Um, there's a lot of little things that we do internally and AI is encouraged. Uh, for example formatting data is a big one. Customers will send us data to index for the AI agent knowledge that we want to put into a certain format to make it better for the AI to learn. AI is great at that. LLMs are wonderful. Say hey, here's a file in this format. I want you to make it in this format. Here's an example and it's instantaneous. So yeah, lots of things like that,

Speaker A: yeah that's great, man. I can't imagine what my life would have been like 15 years ago with AI because one of my jobs as a customer success manager was I was taking a spreadsheet of, uh, store products. And we're talking, think about shoes. You have different sizes, you have different colors of the same shoe, different quantities of the same shoe, and not only getting it listed on ebay, Amazon, Facebook, Marketplace, their own website, different venues. We had to make sure that there was a mapping between all of it and it was all manual. So I stared at spreadsheet after spreadsheet. I'm just thinking I wouldn't have to do any of that now. Uh, and it's just amazing that one client would have taken me a couple of days to do. Now it just takes seconds just to say, here's, here's the mappings. Go to it. And it's just wild how different jobs are kind of going away. And I think it's in a good reason because humans shouldn't be doing this kind of a task. And uh, yet that's where we were back then. And I feel like we're in this era of we can finally let go of these tasks that we shouldn't be wasting our time on anyway. We were given a brain. God gave us a brain for a reason, and, uh, not to do the mundane tasks over and over. So now we can be creative, we can continually be creative and let the AI just kind of take over those kinds of things. And I'm very happy about that.

Speaker B: Yeah, I think the new skill is going to be the skill that's going to become in demand, uh, for, for many jobs, not, not all jobs is. And again, I'm not a big. I do not believe that AI is going to take our jobs. I, I don't, I don't think that that's. That, that, that kind of, uh, it's been every, every technological, uh, advancement has had that worry. Right. The Internet was going to take everyone's job at one point within the 90s when I first kind of started getting into it. Um, uh, so I think what you're going to see as a skill set is this idea of being very good at validating that the inputs are correct and the prompts are precise and clear, uh, and succinct for token usage, et cetera. Uh, and in the case of things like Claude, the skills are accurate, the skills are defined, uh, they're repeatable, et cetera. That's the inputs and then the expertise to validate the output, knowing the nuances of how to format a JSON file or knowing how to, uh, convert an array of data into something. Those middle bits of that, as you said, you, you would have to do as part of your work manually to correlate, you know, skews and such. That bit we don't need to worry about so much anymore. It's, it's the. Did I tell it to do the right thing? And did the thing actually happen? And can I. How do I validate it?

Speaker A: I was.

Speaker B: My son. My son's into, you know, uh, he's seven years old and he's into, uh, scratch, uh, the, the kind of the graphical coding interface where you can like.

Speaker A: Oh, yeah, I haven't had a chance to see that one yet. That's interesting.

Speaker B: Yeah, it's really cool. It's a free little app you can download. And he was playing with it, and then one night we were in the car and he was talking about this game he wanted to build. And I thought, oh, you know what, why don't we boot up, uh, cursor, uh, and we'll. We'll play with it a bit. And so we would. We. We built this game where it's kind of like a snake game, where you snake moves around and you can eat apples. And there was different things he wanted to add different colors, and it was a lot of fun. But at one point he said, I want each color snake to have a super. A, uh, superpower. And I said, well, what should this color be? What should green superpower be? And he says, well, I want it to move faster than the other snakes. And I, and I. It actually made me kind of think about how I work with AI because you can't just say to an AI I want the green snake to move faster. Like, that doesn't. AI would know what that meant. But what is faster? Is faster, 200% faster? Is it 0.1% faster? And so this is where I actually had a really interesting conversation with my kid about speaking accurately. And when you're talking to. Because I think that's a skill that's going to be a skill that his generation will require, is to be able to really speak and use clear language. And English is horrible for that, as we both know. Um, but for an AI, at least now, at least current AI technologies, you're going to really have to have that precision, because you can't just say, make the green snake move faster. It's not going to work.

Speaker A: Yeah. It's just forcing us to use our brain more and to be very clear. And concise. And on top of that, I feel like that even as managers, how can we be clear for our uh, direct reports as well? Because so many years, so many decades, people haven't been very clear and direct reports don't know how they would potentially move up. How can they get that promotion, how can they move into a managerial role? And now it's requiring more. If you can be clear and concise with AI and directions there, you can obviously be more clear and concise with what the trajectory of uh, somebody's role is. I want to switch over to metrics. I know that you're consistently thinking about value if you're tracking the right data, if you're uncovering the customer experience story for your customers, or if you're providing your customers with dashboards showing that codif brings the company value. So thinking at a high level, can you share with me what's on your plate regarding metrics and reporting right now?

Speaker B: Metrics in the AI era, uh, are really interesting to me. Right? And as you mentioned, I've been a proponent of this idea of value based metrics versus service based metrics. One of my old, my examples I always go through is average handle time supposed to be high or low. Right? And most people might say, well, low. We want to solve things fast. But if you work for certain types of companies, you know, you work with elderly care or uh, work with demographics, uh, that might have um, uh, uh, mental disabilities, you know, fast is not going to work. Right. And so I think we, we need to, to re examine the value you're providing versus in this case speed as my example. Um, so I, I have two things that I always say when people ask me about this. The first is, um, AI needs to be treated as AI agents. Need to be treated as agents. They, you need to be transparent with what their goals are, what they, what are they supposed to achieve. And yes, maybe their, their metrics are different than your human agents, but there's at least similar metrics. Like for example, if your human agents can handle 50 tickets a day, can AI handle 200? I don't know. But there has to be a number there where you can start to say like, here's my goals. Here's what an ROI looks like for my, my AI agent. And of course that includes the coaching if your human agents made a mistake or didn't know something, or there was a new product or feature that's released that you need to train them on. That's true with the AI agent. And so, um, when it comes to Metrics internally, you need to measure those, those humans and AI in some alike fashion so that you're able to make ah, really meaningful insights into is AI performing. What I would expect is, am I, am I getting the value out of it? Are my human agents working on the things that I want them to, that maybe require hire, more empathy, et cetera. And then the second, the second thing that I always talk about is AI has fundamentally different ways to uh, to tick off your customers, to annoy your customers, right? Rarely, rarely uh, would a human agent ask a customer to repeat themselves or um, you know, go into a loop of I don't understand what you're saying, I don't understand what you're saying. I, you know, use different words. That's not how humans write. But there are many AI solutions that could get into that loop. And so uh, when it comes to that kind of customer experience, there are things that you need to measure from a metric standpoint in the AI world that you probably wouldn't even, it wouldn't make any sense to measure in the human world. And I call this, uh, just as a name, I call it the Customer Frustration Index. Right? Is the customer saying, you're not listening to me, you don't understand me. Is the AI asking them to repeat information they've already given? Uh, is the agent uh, going into a loop saying I don't understand? Did the AI send the same question twice or send contradicting information? All of those things should be measured because that will help you improve the training of your AI agent, the automations that you're running and the prompts that you're using for your AI agent. I think if you measure both those things, you're going to have a much clearer vision into uh, your entirety of your CX offering. And I also, the, the one advantage that I also see is that it's going to help your human agents understand the role of AI. It's going to reduce the anxiety of oh, this thing's here to take my job. And you can say no. Like look at, fundamentally it's working on different things than you are, you're, you're, this is not about taking your job, this is about augmenting our services because we don't have uh, we don't have uh, support agents to speak Italian, or we don't have sport agents working at 3:00am Pacific Time. You know, it's, it's to show that there's a fundamental difference between human role and the AI agent role. But together you're a team achieving the same outcome.

Speaker A: Yeah, I love the perspective of that because so many people are just, maybe they're thinking it's, you gotta measure humans and AI separate, but at the end of the day, there's some things to measure separate, but for the most part you want to make sure that it's about the customer experience. It's not about AI versus human. It's literally about how do we provide the adequate customer experience, the experience that our customers, very specifically, like you said that our customers require, uh, in order to find value in the product. At the end of the day, that's exactly what customers want, is I just want to be able to log in, feel the value and measure, uh, that value and then move on with the rest of my day, not focus so much on the product. But like, literally it's just one, one step of, uh, of their day not taking up their entire day trying to figure out the solutions.

Speaker B: Yeah, exactly. And, and I think that you do that through a combination of services. Right. We have customers that come to us and say their number one use case is some people submitting, uh, warranty requests. And it's just a repetitive thing. Fill in this form. Did you know, have you registered your warranty? What's the damage? Can you upload a picture? Humans could do that. There's no reason why a human can't do that. But if it's three in the morning, ah, and humans are all sleeping, AI can do that just as readily. Right. Um, and so, yeah, it's all about that service level that you want to provide. And, um, many products offer you scheduling. Like maybe you turn off the AI for certain periods of time, you turn it on for other periods, whatever it might be. Um, I think that it's one of those things that you, uh, decide the service level you want and then obviously, as we just said, measuring it is really important to make sure that it's achieving the goals on both sides of your team.

Speaker A: Yeah, I believe humans were made for more. Uh, we were measured that humans were only using 10% of their brain. What are we using now? I feel like there is an uptick, uh, that we have to be using more or we're even still. We still have an untapped amount that we still could tap into based on the way that we think, the way that we structure our, our, our goals, the way that we structure just everything that we're working on. I feel like, uh, we're, we're definitely made for more.

Speaker B: I, I mean we could probably go into a different type of podcast and talk about how the brain functions and, and uh, in psychology and things. Last week I was in, I was traveling and I had to pay for parking and something went wrong with the parking app and I talked to a human agent through chat and the experience was fine. Again, there was nothing wrong with the service I got. The person that helped me was kind. They were definitely using macros. You could tell they were also had multiple conversations going on because every response was slightly delayed. Um, as a sport leader, I recognize that I'm sure you would too. Um, and I thought about it afterwards that, you know, AI could have handled that and would it have been handled better or worse? And then when it comes to expectations, this is the one that gets me all the time is, is my measurement of how I want AI to serve me against the human equivalent of that service. Like I was a, it was a 30 minute, 35 minute chat I had with this, this agent. I think AI probably could have done it in minutes or seconds. Um, um, because all it was was uploading a receipt that had the wrong data in it and correcting something in their back system backing system.

Speaker A: Um, um.

Speaker B: But am am I measuring it? Am I measuring the potential AI against the potential human or am I measuring the AI against my expectations? Um, um. And that was an interesting realization for me because I would have expected, you know, for my generation to be comparing it against humans. All human could, this is too slow. Human is slower than an AI could solve this or to an integration platform could solve this. But that's not what I think we're seeing right now. We're seeing a generation who is so used to instantaneous, customized personalized responses that the expectation has become that, that level. And that's fine if that's who you sell to. If you sell to a generation, if you're selling the, you know, the trendy shoes or clothing of the time, you want your experience to meet those expectations. And um, I found that very fascinating when I was thinking about who am I comparing this 35 minute, largely positive experience to? Um,

Speaker A: um.

Speaker B: And I think it would be a different answer depending on who you ask.

Speaker A: Sure, yeah. All right. So every company wants to measure things differently. I remember working with Intercom and back in 2015, 2016, just pestering them, saying, uh, I need different reporting, I need different data. And they would just say use the API, export it, get it over into, get it over into Looker. And because we're not going to build that reporting dashboard for you. And so many other companies were like that. I'm just picking on Intercom. Because that's where I was at at the time. And so. But I mean, I feel like data is probably the most difficult aspect to get close to.

Speaker B: Right.

Speaker A: And you mentioned that when there's no clear narrative around metrics, customers start building their own reports. So how do you think about AI's role in owning that story before someone else does?

Speaker B: Yeah, I mean it's, it's a problem and I think it's a problem. It's an unsolved problem across the industry. What, what does resolution mean?

Speaker A: Right.

Speaker B: If you, if you really think about what is an AI resolution and how do you measure that? Right. I mean, if, if, if I ask, uh, I mean, think about, forget about supports, forget E commerce. You go to GPT and ask it a question, whatever that question is, and I get an answer back and then I leave GPT. OpenAI doesn't really know that that was resolved. They know that I asked a question and they gave an answer to me. But I have, um, no, there's no indication that that answer was necessarily correct, that I didn't turn around and send an email to someone else like to, hey, uh, Brett, I just asked GPT this question. You know, it's, it's, it's, it's one of those things that resolution is, ah, actually a hard thing to measure. Um, and now that's the simple case. Now think about the complex case, maybe a refund or a warranty claim, right? And you go through some troubleshooting steps before you agree. Let's say you're a physical product that you're troubleshooting and you get halfway through and the person stops responding. Well, is that because that step fixed whatever the problem was and it's like, oh, I don't need a warranty, I fixed it. Or is that because they got frustrated and left? Um, is that, should that be a resolution? Right? And that's the customer. That's like our customers, how they think about it. But think about it then from the vendor standpoint, we still did something here, right? So if we're a vendor, uh, and this is, again, this is an industry problem, we still used our AI models, we still used our integration platforms to take the conversation that far. So from our perspective, we incur the cost, even though potentially the thing wasn't resolved. And so how do you, how do you charge for that? How do you talk about that? How do you explain the nuances? Um, and then the final case, before I go back to fully answering your question, is what about a situation where, you know, where the AI determines based on the workflow that you've trained it, that, hey, this situation is something I cannot handle. I need a human to help me here. Well, same thing. Is that resolved or is that contained? Uh, because the AI made the determination, you trained it to hand off to a human. It's not like the customer got frustrated, said, I want to speak to a person. Now you train the LLM to look at that and say, oh no, this has to go to a human. I'm making that determinant. Right? So the problem is because of that nuance, obviously a company, any vendor has to make some, some decisions on how the, what's their point of view on how this stuff should be m measured. And you know, you could criticize or praise any given metric from any given vendor in the space, but customers will do one of two things. If there's, if there's no narrative or the narrative doesn't agree with, with what they want, they will, they will generate their own narrative. And the problem that that incurs is now they're kind of telling an internal story to their stakeholders that's different from the story that you believe your platform was meant to tell. And that can lead to all sorts of problems, like, for example, misunderstanding the solutions the software is providing, or, uh, confusion in terms like if it says X over here and there's the same term over here, um, but those numbers are different. Well, how do you consolidate that? And it. So it hurts things like renewals. It hurts things like if you're having a qbr, you're not speaking the same language anymore. So how do you control it? I mean, I think the best way for most vendors to control it is to have a very clear point of view. Documented metrics, um, that are easily filterable, easily manipulatable to whatever the customer's specific requirements are, but that the story, at least as you tell it, is very consistent. It's consistent across the product. So if you're searching in one place, uh, for a certain metric, CSAT metric or resolution metric or whatever it is, this place handles it the same place as the other place. And then if a customer does want to go their own route, understanding, um, why they feel your metrics aren't telling the story they need, um, is there a way that they can filter your metrics to get to what they need? Natural language is a real hard problem, and human behavior is an even harder problem. Uh, and we're trying to measure both with, uh, these technologies. So it's a tough problem to solve.

Speaker A: In my mind, when I'm thinking about measuring AI. Think about like deflection. I know deflection and resolution used to be the same terms, but I've also started to hear that deflection is more of a negative term in terms of uh, you deterred the customer from wanting to use that solution and you needed, you required a human solution after that. Um, and then maybe there's another definition for it, but uh, from other conversations I've had, they started using the word deflected and uh, needing to pick it back up with humans and then assisted. So this is one that's not measured that uh, folks, it's hard to even identify but if you think about a tier one support person and they have the informational responses and it at least deterred the customer from needing to like talk to a human for at least five minutes. So the assisted portion was a five minute solution before needing actionable steps, uh, that weren't already implemented into the AI agent needing to go over to humans. So at least the assisted portion would be five minutes worth. And then you could start to calculate based on how much time was spent on providing that information prior to being escalated. But it wasn't a total loss. Uh, at least they didn't need to reiterate the things that they said. Um, and then another piece of it is, uh, I've seen AI agents where it's like half of a paragraph essentially and then here's a link to an article that nobody clicks and that seems to be what so many products are pushing or even people are not paying attention to that customers don't want that. Prior to AI being so integrated into support, uh, I was leveraging Intercom and uh, their routing workflows and it was pretty much like you couldn't click more than three times. Uh, I did not want that. I didn't want somebody to be like, uh, how many times do I have to click these, these options before I get to a human? But really it was no more than three times. And again this is archaic. But think about it this way. By allowing them to select where their issue was and then providing maybe two paragraphs worth or even like a video, an inline video, before it even made it to a human, we were able to deflect over 60% without leveraging AI at that time. Yeah, and then now you don't need a decision tree. You can go based on what it is that they're asking for. However, to this point it's much more difficult to provide the right clear context written by a human. And then even so an inline video that shows exactly the solution for what they were looking for. Are we, are we almost there? I feel like we're almost there. But we would need to provide, uh, the right context around. Here's why you would show this video and here's the bank of videos that you would use it for. Um, but yeah, so that's, that's kind of my take on, on measuring and even like how we're kind of moving along with AI in, in chat. But at the end of the day the question will end up being like, did we really frustrate them and they didn't get any value whatsoever? Did we assist them? And then we actually did give them value, but it's, it's not complete. And then did we resolve the situation incomplete and it never had to make it to a human?

Speaker B: I, I love the term assist, actually. I, that, that's a, uh, that's a term I haven't thought about in the terms of metrics. But that's exactly what's happening, right, is, you know, you mentioned the intercom routing system, right? And, and so like that's a form of assistance. Now the difference is in the AI world, that form of assistant costs money, whereas in the re. In, in, in the intercom world it was a feature of just, you know, basically a bunch of boolean. Did the, you know, does it beat this criteria? X, Y. And so the um, you know, now that, that decision costs money because it costs tokens, et cetera. And so there's a charging model that needs to accommodate for that assistance, even though that might not lead to a deflection or a resolution. And um, I don't pretend to know the right answer there. Um, again, I go back to value, right? I think that the discussion on how many conversations you remove, how many resolutions you make is obviously valuable and something you need to measure, don't get me wrong, but it's more about the value it brings to your business. Right. You know, um, and that value, sadly could be in savings of headcount, right? I mean that we should be honest about that. But it could be in other things. Like what's the opportunity cost of someone spending five to 10 minutes filling in a warranty form that AI can do automatically just by asking you your email and picking out all your various user information and putting it all together? Right, yeah, there's an opportunity cost there. And so to me, you know, part of the metrics is, yeah, resolution's important because you need to make sure that the system you bought is working. But what about the overall capabilities, uh, that you're increasing within your team, Are you better at customer success? Are you, do your agents have more time to work with, uh, the serious scenarios that require their attention? Are they less burnt out because they don't have to handle three chats, four chats simultaneously? They only have to handle the one that gets escalated? Um, are your international customers happier because your AI agent could speak 100 and some odd language instantaneously? Um, those are the things that you should be also including in those metrics. Beyond that, hard, you know, 75, 80% resolution rate. Um, yeah, a great, great point. I love the term assistance. I really do. I love that, that assist metric. We need to think more about that.

Speaker A: Craig, thank you so much. I appreciate you taking the time to speak with me today. And I love that you're pushing the boundaries on reporting and metrics and we're thinking about resolution and how to measure it, how to actually understand the way that we meld human effort and the way that we put together AI effort and how we start to measure things in the right way. You can connect with Craig on LinkedIn and that link is down in the description. And that wraps up today's episode. Subscribe to hear more incredible stories in this era of unresolved. Have a great day and a productive week.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • #291 Why Most AI Projects Fail to Deliver ROI, Sinohe Terrero, CFO and COO, EnvoyGrowCFO Show · on Claude91 / 100
  • AI Agents, False Productivity, and the Sales Team Reset with Gabe LarsenMake It Happen Mondays · on AI agents91 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on AI agents86 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100
  • How to Sell Against a Competitor Already in the BuildingSales Leadership with Fexingo · on HubSpot85 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Claude84 / 100

More from Unresolved.cx

All episodes →
  • Unresolved.cx - the development conversation managers skip on themselves - Yvette Johns - Maze65 / 100
  • Unresolved.cx - The customer experience signal AI doesn't know it's missing - Robert Cabral - Runway67 / 100
  • Unresolved.cx - The customer who reaches out when they have to63 / 100
  • Unresolved.cx - Proactively spotting the moment a customer experiences an issue
  • CX-News.com Feb 5, 2026 - Deeper AI Context, Automation Rate Metric, and Real-Time AI Suggestions
Explore the best B2B Customer Success podcasts →
All Unresolved.cx episodes →