The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Invisible Machines podcast by UX Magazine
Invisible Machines podcast by UX Magazine artwork

The Checklist Your Deck Is Missing ft. Jeff McMillan

Invisible Machines podcast by UX Magazine · 2026-06-18 · 53 min

0:00--:--

Key moments - from our scoring

Substance score

57 / 100

Five dimensions, 20 points each

Insight Density12 / 20
Originality10 / 20
Guest Caliber15 / 20
Specificity & Evidence11 / 20
Conversational Craft9 / 20

Jeff McMillan, former head of firm-wide AI at Morgan Stanley and founder of MacMillan AI, joins Josh Tyson and Rob Wilson to address the gap between agent hype and operational readiness. The conversation centers on a critical misalignment: most enterprises focus on AI applications and models (the visible layer) while neglecting the foundational work - data accessibility, semantic layers, control systems, and governance that actually enable reliable AI at scale. McMillan breaks down the stack: foundational data quality and accessibility (layer one), semantic models and RAG infrastructures (layer two), control layers embedding business rules (layer three), then models, orchestration, and applications. The episode explores why small deployments can succeed through brute force but scaling to 15,000 agents demands near-perfect data accessibility and 99%+ quality standards. A major theme is the hidden cost of unmanaged AI development - token burn on unshipped features and uncommitted backlogs, driven by the addictive nature of prompt coding. The hosts and McMillan discuss how to connect AI investment to strategic business priorities, measure true capacity gains and their application to value creation, and embed controls and evaluation practices that organizations currently lack. Critical missing elements include explicit prompt iteration (150+ cycles for production-grade systems), golden source testing, and independent monitoring layers that work like peer review across AI systems.

Key takeaways

  • →Scaling agents beyond 15 demands near-perfect data accessibility and 99%+ quality standards; brute forcing only works for 5-15 agents before the foundation becomes critical.
  • →Organizations must be explicit about prompt design and iterate 150+ times before production, backed by golden source testing and cosine similarities rather than relying on initial vibe coding.
  • →Uncontrolled token burn on backlogs that never ship is worse than not using AI at all; firms must connect AI spending to strategic business priorities and measure where freed capacity is actually applied.
  • →Effective AI governance requires independent monitoring layers - multiple models asking the same question in parallel to catch failures humans or single systems would miss.
  • →Knowledge management and process mapping are foundational prerequisites; most sophisticated businesses cannot articulate their core processes with sufficient specificity to apply AI precisely.

Guests

Jeff McMillan

Topics in this episode

Morgan StanleyKnowledge graphsPrompt iterationRAG (Retrieval Augmented Generation)Cosine similarityMacMillan AIGolden source testingData semanticsControl layersGovernance and monitoring

Questions this episode answers

How do you scale AI agents beyond a handful to hundreds or thousands?

Scaling requires near-perfect data accessibility (100%) and 99%+ data quality standards, plus a complete foundation of data semantics, control layers, and governance - not just better models. You cannot brute force 15,000 agents the way you can 5-15.

What is the most common problem preventing organizations from deploying agentic AI?

Lack of education and awareness among senior executives, combined with absence of a consistent data platform. Building a reliable data foundation typically takes much longer than five weeks.

Why is burning tokens on unshipped software features worse than not using AI?

When developers use AI on backlogs that never ship, token costs increase labor costs while productivity gains don't translate to revenue or customer value, making the organization's true economics worse than having no AI.

How many times should you iterate on prompts before production?

You may need to evaluate and iterate on prompts 150 times to achieve extraordinarily high quality output. Most organizations underestimate this requirement and confuse initial prompt success with production readiness.

What are the three strategic questions CEOs should ask about AI first?

One, what are my five to ten strategic business priorities, and how does AI enable them? Two, who could destroy my business in ten years with AI, and what moats do I need? Three, what new products, services, or markets can AI unlock given my unique position - scale, infrastructure, or knowledge base?

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

12 / 20

The episode surfaces several non-obvious operational points - most notably the distinction between capacity creation and capacity application as separate measurement challenges, the scaling cliff between 15 and 15,000 agents, and the evaluation loop methodology - but loses significant density to prolonged donut analogies, the farm anecdote, and circular philosophical meanderings that eat substantial runtime.

you not only have to know where the capacity came from, you have to know where that capacity was applied and whether or not that application of that capacity actually was driving greater revenue or greater customer service or whatever
if you're building 5 or 10 or 15, sure, you can brute force it, you can fake it, but what happens when you've got 150 agents or 1500 agents or 15,000 agents

Originality

10 / 20

A handful of genuinely fresh framings appear - 'use case zero,' the incremental seeding of human judgment, and the inversion to 'agent in the loop' - but the bulk of the advice (data foundation first, evaluation matters, human oversight needed, AI can make you dumb or smart) is well-worn enterprise AI commentary that circulates widely.

it's not going to be like one moment in time. We're going to start to seed our judgment over time and it's going to happen in very small increments
I call it Use case zero. Because I always feel like if you're going to build an optimus Robot, that use case 0 is that Optimus can build robots, not fold laundry

Guest Caliber

15 / 20

Jeff McMillan is a genuine practitioner who built firm-wide AI infrastructure at Morgan Stanley at meaningful scale, now consults and teaches at Columbia - not a thought-leader or career podcaster, and the transcript reflects real operational scar tissue rather than theory.

first on Wall street as head of firm wide AI at Morgan Stanley. Now through macmillan AI
my first project on Wall street, um, was to build out a customer information database. I was in my early 30s, uh, and knew nothing about data. And what's interesting is that problem was critical over 25 years ago

Specificity & Evidence

11 / 20

The episode offers concrete evaluation percentages, a named testing cadence (100 users, 20 questions per week), and the 15,000-agent / 99%+ data quality thresholds, which is more specific than average; however, no named company outcomes, dollar figures, or Morgan Stanley case data are shared, and most examples remain illustrative or hypothetical.

maybe the first time you run it through the model it's 80% accurate. And then it gets 85 and 90, 93, 94. And then what happens sometimes is...it goes from 89% to 74%
I need you to ask 20 questions of the system. And I need 100 people to do that every week on Monday

Conversational Craft

9 / 20

The hosts occasionally generate useful friction (the capacity-measurement question, the knowledge management follow-up) but habitually over-talk, complete the guest's sentences, and let analogies run far past their useful life; there is no meaningful pushback or challenge to any of McMillan's claims throughout the 53 minutes.

Does knowledge management kind of get you there in a way though, if you do it the right way?
But companies don't like knowledge management. It's boring. They have to fund something that they don't understand. It's hard to roi, analyze it. Um, is there a way around it?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker C57%
  • Speaker A33%
  • Speaker B10%

Most-used words

knowledge39point32agents27problem19different19information18system17data16better15back14josh13answer13agent13tokens13layer13management13

Episode notes

Everyone wants to talk about agents and models. Jeff McMillan, starts where almost nobody else does: the foundation. In this episode, Jeff McMillan, founder of McMillanAI, former Head of Firmwide AI at Morgan Stanley, and advisor on enterprise AI, maps AI as a stack: high-quality accessible data → semantic layer (knowledge graphs, RAG) → control and governance → models → orchestration → applications. The heavy lifting is in the bottom layers. Organizations that skip them can fake it for a handful of agents, but at 150 or 15,000 agents, you need near-100% accessibility and 99%-plus quality, or you’re monitoring chaos you can’t see. Josh and Robb press him on why knowledge management feels unfundable, why tribal institutional knowledge breaks when machines execute without judgment, and why evaluation (golden datasets, custom org evals, regression when models upgrade) is the work builders hate and operators can’t skip. Robb names the trap CTOs are falling into: grinding tokens on feature backlogs that never reach production or revenue.

Full transcript

53 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Foreign.

Speaker B: Welcome back to Invisible Machines. I'm Josh Tyson, joined as always by Rob Wilson, co founder and CEO of onereach AI. Rob and I are co authors of Age of Invisible Machines, the first best selling book about agentic AI. And this is the podcast where we explore how intelligence is actually built and where it's going through conversations with the people shaping it. Most enterprises are under intense board pressure to deploy AI agents. Very few can answer three questions their C suite will be faced with sooner rather than later. What data are agents actually running on? What controls prove agents are behaving? And how do you know agents are getting better, not just faster? Jeff McMillan has spent decades building that foundation, first on Wall street as head of firm wide AI at Morgan Stanley. Now through macmillan AI, he's back on Invisible Machines to address the widening gap between agent hype and operational readiness. As he sees it, you can brute force a handful of agents, but scaling to 15,000 demands near perfect data accessibility. That's why his stack starts with data semantics and controls, not models and apps. Jeff explains why evaluation and golden source testing are critical to success. And we discuss why burning tokens on software backlogs that never ship is often worse than not using AI at all. If your organization is pursuing agentic AI without a foundation in place, this conversation is the checklist your deck is probably missing. Let's hear more from Jeff McMillan. All right, Jeff, well, maybe we can start here. Um, you have a lot of experience building out canonical knowledge and kind of source of truth for large companies. And as we've talked in the past on this podcast, uh, together, it feels like that is sort of the critical first step. Right? Like, if you're going to have agents do things, you need to teach them, you need to give them knowledge to work from. And then as we were preparing for this call, we kind of had this realization that the same is sort of true for a lot of executives right now. They want to pick up tools and buy stuff, but what they really probably need first is knowledge. So it all kind of comes back to knowledge in a way.

Speaker C: Yeah. And what's interesting, first of all, guys, thanks for having me back. Um, and what I find interesting, my first project on Wall street, um, was to build out a customer information database. I was in my early 30s, uh, and knew nothing about data. And what's interesting is that problem was critical over 25 years ago and maybe even more critical now. So something would never change. Uh, but to your point, everyone that I talk to wants to talk about AI applications and models and let's be clear, that is the visible layer. And obviously that's where the hype is. But to your point, Josh, nothing happens in the world of AI without good information. And not just good information, but information that is accessible, of high quality and structured in a way that AI can truly benefit from those assets. When you talk about AI architecture, it all starts on a foundational layer of high quality, accessible information. So that's layer one. Number two is then what I'm going to describe as the semantic layer or the model or whatever you want to call it. But that's where you apply knowledge graphs and rag infrastructures. And then you have what I call your control layer, which is where you apply the business rules on what you want these systems to do in Dode. And then you have the models and then you have an orchestration layer and then you have applications. And like I said, everyone just kind of wants to jump to the top layer. And as we know, the heavy duty lifting and the non sexy stuff is in those bottom two layers. Uh, and I think maybe for good now people are starting to realize that and starting to get their act together and focusing appropriately on, uh, those issues.

Speaker A: But companies don't like knowledge management. It's boring. They have to fund something that they don't understand. It's hard to roi, analyze it. Um, is there a way around it?

Speaker C: Well, I think yes, but I think there's a tipping point, right? Like I think if you're building 5 or 10 or 15, sure, you can brute force it, you can fake it, but what happens when you've got 150 agents or 1500 agents or 15,000 agents at some point you're not going to get around this. And what's so fascinating is I've been out and about now talking to people and helping with them. And the number one problem is education awareness of senior executives. And the second problem is a lack of a consistent data platform. And people are like, well, how do I fix that? I'm like, not in five weeks. The other thing that's really funny, guys, is that when I talk to technologists and I talk to business people, there's this sort of battle going on between who's going to own the app layer. Yeah, like that where the, that's where

Speaker A: the, like that's where. That's where everything is happening, Right?

Speaker C: That's where everything's happening. And I'm like, you're fighting over crumbs.

Speaker A: Yeah, it's like gardening is about mowing the lawn or like understanding how to grow plants and botany Right. You're like, oh, let's just. I know how to mow a lawn, I know how to push a lawnmower. I'm a gardener.

Speaker C: That's right. And you and I are fighting about who gets to, who gets to push the lawnmower. But the point is, is that true transformation in a world in which you have 15,000 bots, I think that's what you have to prepare for and in which those things are going to go off the rails and they're going to do stupid things and you have to monitor them. That is not going to happen with a uh, 90% accessible data set. It needs 100% accessibility. And by the way, and you need to be in the 99 point something levels of quality, which I mean very few firms are even anywhere close to that.

Speaker A: If you think about humans, right? We have 10,000 employees in your company. It's the same problem. Look at all these people. What are they doing? How are they making decisions and choices? How are they executing? There is a, a knowledge apparatus, but it's, it's so um, institutionalized and tribal, right? That it's not something like that's that that upper management really has to spend a lot of time thinking about. Like at NASA, you have to spend a lot of time thinking about your knowledge management system. Because the difference between a crash, um, it could be just, you know, the communication between the measurement of metric to temperio. But in this scenario, so much of a company's knowledge is just passed on from person to person and it's something that the top of the org hasn't really had to think about. Right? They just sort of take it for granted that people will just train the next guy and share that information. But in this particular world, like it's different because now you're talking about a system that will just do what you ask it to do, irrespective of thinking about whether this is a smart idea or not.

Speaker C: 100%. And to your point, if you don't explicitly train your agent to do something, it has the capacity to do very stupid things. And as humans we kind of say, oh, that looks dumb. Now how do we know that it's dumb? Because we've been doing this for 20 years and we've got institutional knowledge and we have judgment. And by the way, I'm not saying you can't teach AI to do that as well, but if you don't, and

Speaker A: to be clear, you have to be explicit. Like you said, you have to be explicit.

Speaker C: And by the way, you have to be Explicit. And here's the other thing that people, they just don't do is if you want something to operate at extraordinarily high degrees of quality, one, you have to be incredibly explicit. But you might have to evaluate and iterate on those prompts 150 times before it gets the point. This piece about evaluation is such a missing component for a lot of organizations. It's like, oh, I can vibe code now. Well, sure I can, right? I mean, I could also theoretically Drive a Porsche 911 for 200 miles an hour on the racetrack, but probably wouldn't be a very good thing. And I think the illusion of skill, which, by the way, uh, I use vibe coding, it's very powerful, right? But if you want to do that, it comes with a responsibility of actually testing this thing. And because the non deterministic nature of this technology, it will do weird things on you. And you have to run it through enough golden source testing, cosine similarities, user testing. And by the way, most organizations do not have their experts waiting around for Jeff McMillan to come along and say, I need you to spend four hours this week testing. And they're like, what do you mean, testing? I got clients, right? So there's this, like, organizational problem where the business is like, well, I want these tools. I can do this thing. And one, they don't really know how to do it. They can learn, but they don't know how to do it and they don't have the capacity.

Speaker A: Yeah, Josh and I were talking, um, the other day, uh, about this, and one of the things that, you know, we were sort of, you know, came up in our discussion was like, you know, if Josh and I are going to handle a book together, you know, it'll be like, okay, you, you grab chapter five, I'll grab chapter six, right? Um, now what I don't have to worry about is, like, coming back two weeks later and being, hey, so sorry we haven't had a chance to check in. Where are you at? And he's like, oh, I, I've written 2,000 pages, right? That just will not happen. Right? So, but with an LLM M, like, it might be at 100,000 pages because, oh, shoot, I forgot. Like, it's not going to stop. And, um, and what I'm trying to say is that we, in a lot of ways, it's helpful to anthropomorphize this. Like, to understand how to think about agents, we need to understand people. Because in some ways, if I ask you to invent a different color from the Colors, you know, that's pretty much impossible for you to do unless you've seen it. You know, Josh's joke is as a kid, he kept asking, what is the color clear? And, and I think it's kind of like that. Um, and the I. And the idea behind this is, is to say, like, it's, it's great to think of the, these, these things as intelligent and we do compare to our intelligence because we have nothing else to compare to. But at some point we forget that it's not and that it does things we don't do. It will write and write and write and write, uh, forever until you stop it. So in a lot of ways, like, people have this optimal stopping built in. You know, they're lazy and they conserve calories and, and they have other things to do with their time in life. And so you don't have to worry that asking someone to start writing that you're going to have to come in and tell them when to stop.

Speaker C: Well, three points on that. Number one goes back to the data layer, because if you're going to write something, you want to make sure it's based on good information and knowledge. Right. And that you're not creating your entire thesis off of Reddit. Right. That's point number one. Point number two, you have to build embedded controls so that these systems will behave. And by the way, not perfectly, nor are we. Right. I mean, we all tell employees what they should or shouldn't do, and they don't always do it. Machines sometimes make the same mistakes. And by the way, like, I was on a thing the other day where I was talking about this concept of like, how do you embed your ethics into AI? Well, you can, you can prompt your presentation.

Speaker A: I mean, everybody does it. And Therapok has their constitution in there. Absolutely right.

Speaker C: And but by the way, like, organizations need to think about what are the ethics of their AI. I mean.

Speaker A: Yeah, their own constitution.

Speaker C: That's right. What is, and Rob and Josh and Jeff, like, what are our values and how do we want to imbue those values? And what's great about AI is you can do that. And then lastly, and I talked about this a little bit in the beginning, like, there's the governance layer, but then there's the monitoring layer. So you want to have something that says, I don't want people doing stuff. And then you may want to have a completely different model that sits outside, that's looking across that stack and is almost independently asking, does something smell Right. So you actually have like, it's almost like, Rob, you're going to do the work, and then Josh, you're going to look at Rob's work and then Josh, you're going to do the work and Rob, you're going to look at, uh, Joshua, this idea that you create independence. And by the way, you might even have two different models or three or even four that are asking the same question. And by the way, you may get different answers. But I think people need to acknowledge your point, Rob, that these things are imperfect.

Speaker A: Yeah.

Speaker C: And by the way, they're imperfect in different ways. That then we are imperfect.

Speaker A: Uh, exactly. That's the key, right? It's a different kind of intelligence. Like, it's great for a point to, to think of them in an anthropomorphic way because it helps us understand them. But then you go too far and you start to realize, like, wait, but they're different. They don't have self preservation. They want to burn as much electricity as they possibly can. They have no sense of stopping with humans. Managers are always talking about how to get them going, how to get them productive with these things. It's like, how do you stop them? When do you stop them? And that kind of comes. So, like, the point I wanted to really dive into with you, I Talked to these CTOs that are like, oh my God, like I'm so busy. I'm finally able to catch up on my backlog of all the features of all the software, uh, that I've been behind on. Really, um, software quite honestly that no one's ever going to use. Um, but now they're cranking through their backlog, right? They're like agents. I don't have time for agents. I'm getting through my backlog. Uh, agents are building the code for all the things I was supposed to build. Um, and it occurred to me at that moment, especially after talking to Joshua Gans and the microeconomics of AI, that there's really bad, then there's bad, and then there's good. Right? And really bad isn't not using AI. Really bad is burning a shit ton of tokens on things that never see the light of day or produce a dime of revenue. Better you don't use any AI, then everyone grinds on tokens within your company. But none of those tokens actually see the light of day when it comes to revenue production and how easy that is to do. Because these systems are so dopaminergic, they're so addictive that you got these developers that are cranking and producing, but they're producing stuff that's not production ready. So what happened? The company's labor bill, essentially, with tokens, just got increased, productivity got increased, but revenue didn't move, which is scary. Right? Like, that's probably what happens when you don't know what you're doing, but you're trying to lean into AI because you feel like you're falling behind.

Speaker C: So here's what I would say to that. I think it is very hard for organizations who have never experimented with AI to sit down at a board meeting and say, here's our priorities, right? Very, very difficult. And they don't know. So I think there is something to be said for burning some tokens in a controlled environment, um, where people have access to an appropriate set of tools with a set of training. And I think most organizations have to kind of go through a little bit of that for them to really understand the art of the possible. Um, so I think there is an element of what I'm going to say wasted tokens in the sense that there is some learning. Right. So I would not say that that is fundamentally bad, but I would agree with your premise, right, that what firms are not doing is maybe they get through that after six or nine months what they're not doing to a large extent. And I don't want to say all because there are some great organizations out there. Um, I like to think I worked for one of them that are basically saying, what are my strategic priorities for my business? And if I were to employ five to 10 things at a strategic level, what would they be and how would AI enable those things? That would be the first question I would ask. The second question that we need to be asking is, is if my business Is destroyed in 10 years by AI who destroyed it and how do they destroy it? And therefore what moats do I need to establish to prevent that from happening? How does AI play a role in that? Then? The final question, which nobody's asking is AI actually reduces the friction of new products, new services, new markets, new clients. And given my being a, uh, company's unique strategic position, maybe m it's its infrastructure, uh, maybe it's scale, maybe it's their knowledge base. Whatever it is, is there an opportunity to expand what I do today in ways that I've never even thought about? So who knows? In 10 years, wealth management firms are going to be selling legal and accounting services, and legal and accounting service are going to sell wealth management services, right? Because they're going to be able to create new products. And I think the missing piece is that conversation. And I'm not even here to say what that, what the result is. But my strong advice, if you're a CEO or a senior leader, like, you need to be having those conversations and let it drive the technology. Because to your point, and I'm guilty of it too, I'm sure you guys are. You know, it's two, three in the morning and you're still playing around with pipe coding.

Speaker A: Yeah. And you're adding new features because it's more fun instead of just finishing the ones you already have in there. Um, because that's painful and it has no problem offering you the next feature that you should work on.

Speaker C: It's crazy. I was doing something the other day. Um, I bought a farm and I built a tool. I'm trying to basically use AI to manage the whole thing from start to finish. And, um, I kind of was four hours in and I was pretty proud of what I had gotten to. And I said, what ideas do you have for me? And he gave me like 37 more ideas. And I said, oh, yeah, go ahead and ruin it all. Yes, just do them.

Speaker A: Exactly.

Speaker C: And guess what? Guess what? It all started breaking on me. And then I was like, then I spent another 14 hours of my life debugging and telling the thing, no, I don't want it up here, I want it down there. I mean, and I think that's exactly right. And I think we're all guilty.

Speaker B: It's addictive.

Speaker A: Uh, I don't think people understand it's dopamine. Mean, like that Chase is built. You know, whether you're chasing News feeds on TikTok or whether you're, you're grinding on new features that you think are going to be the next, the next big thing. Um, getting them out the door is painful. And it's easier to just jump on the next thing. But it's also like, so, so I was thinking about this from accounting lens. Like, we always look at labor costs. Let's look at token. Like, if you, if, if in your accounting you have like one line item, which is tokens, you're in trouble. Right. You need to understand, like, what were those tokens spent specifically on? And those have to roll up to this knowledge that we were talking about to say, like, if you have a knowledge system and that knowledge system is done well, then it understands the priorities of the business and understands what is being done. And then if you have awareness of what everybody's burning tokens on and it's being bumped up against that knowledge, you have stronger sense that you're not going to have like 10 versions of the same piece of software being produced by 10 different people.

Speaker C: Yeah, but I'll tell you, and you know this firms are, I mean, token consumption is easy to measure, but there's two problems. One, organizations have a very difficult time measuring the impact. Or let's just use, let's forget revenue and risk for a second. Let's just talk capacity. So the first problem is knowing that you've created 30% capacity because that implicitly means that you have a baseline that you know what people are doing today, which in many organizations they don't. But there's even a more complicated problem. Let's make believe that we have perfect instrumentation on what the process was before. Efficiency only comes from whether or not the capacity that you created is applied to something, value added. So you and I, all of us, can create 30% of our day. And if we go play golf with that capacity, there's no value that has been created for the organization. Right. So you not only have to know where the capacity came from, you have to know where that capacity was applied and whether or not that application of that capacity actually was driving greater revenue or greater customer service or whatever. And those, the first part is hard enough, let alone the second. Very few organizations have the sophistication to be able to really understand where that excess capacity is applied and whether it's actually generating value.

Speaker B: Does knowledge management kind of get you there in a way though, if you do it the right way? Because LLMs might have set knowledge management back a bit because it makes it so easy to just look at explicit data and, and you can throw it somewhere and summarize it and think that you've done some degree of knowledge management. But it might be tempting if you're doing that, to ignore the more difficult work which I think is mapping processes, um, and like seeking out implicit knowledge. Right? Like some of it might even require getting up from your desk and going somewhere and talking to some other person. You have people scattered everywhere, um, using more and more of these tools and maybe like not having as many conversations. So there's, there's this knowledge management challenge of finding the implicit knowledge. But once you unearth that, you're also unearthing, uh, information about processes and how they run from end to end and who touches them. And then you are finding kind of what you're talking about, right? Like the value points to where you can actually properly apply, uh, automations.

Speaker C: Right. I mean, well, you're using knowledge management in the broadest sense of that term, which, you know, I'm, you know, I'm a strong advocate of that, I guess specifically to this though, I think what we're talking about is having a deep understanding of your core processes and you understand how work moves from right to left. You understand the value added, the checks that are in place. You understand your failure and defect rates, your rework costs. In high end knowledge businesses, we don't think that way. We hire really smart kids from Columbia and Stanford and Harvard and we put them in a seat and we say work and watch me and you will learn and you will work 80 hours a week and you'll do that for four years and you will learn that tribal knowledge, right? But when I come to you and I say, you senior person, how do you run your XYZ process? Most high end sophisticated businesses have a very difficult time articulating that at a degree of specificity that you would from a consulting firm that mapped the visios. Right. And part of, I guess what I'm getting across is it is very hard for you to apply artificial intelligence in a precise, high quality output way unless you know those core processes, unless you know what your standard KPIs are. And therefore because you need those to be able to measure whether you're actually producing something that's better or worse than you had before. And most highly paid people in this world don't think that way.

Speaker A: Yeah, it's just like LLMs have EVALs, right? How does anyone know that GPT4 is better than GPT3 without EVALs, right? Someone says, oh, it's smarter. And you always talk to people who say I think cloud's smarter than. But at the end of the day these are just perceptions, right? And there has to be these evals so that we know we're making progress versus just training new models and spinning well.

Speaker C: The other, I mean, you know, not to bore people, but like you want to do this, right? One, if you're going to create, let's say a bot, let's just say, uh, a Q and a bot. Simple example, you probably want 1,100% accurate inputs and outputs and you want to run those inputs and outputs through your model every single time you do an upgrade. And maybe the first time you run it through the model it's 80% accurate. And then it gets 85 and 90, 93, 94. And then what happens sometimes is you fix something in the model that you think you fixed a prompt, or maybe you adjust some data sets or you put some tagging in and Then all of a sudden, it goes from 89% to 74% and you're like, oops, what did I do? And then you got to go back again and wire it. But the point is, unless having to your point be like, oh, I love the new anthropic model because I asked three questions of it and it gave me good answers, and then you used it and you asked about your, you know, your dinner for Friday night, and you, like, you didn't like the, uh, extra garlic in your dish, right? Like, these are very personal things, which they're not statistically significant, right?

Speaker A: No.

Speaker C: And by the way, especially because my

Speaker A: kids are using my LLM and, and polluting my memory, you know, they're like, that's a different.

Speaker C: I'll leave that. I'll leave that to you at home, Rob. But the point, though, the point I was going to make is like, what I just described is step one in the process, then step two is you put it into the wild and you say, you know what? I need you to participate in a training program here, a testing program, and every week I need you to ask 20 questions of the system. And I need 100 people to do that every week on Monday, and I need the results. And when you get something that's wrong, I want you to give me a thumbs down, and I want you to describe to me in words why you. I got a thumbs down. Because what will happen is your golden source will get in your high 90s and M. Then as soon as you give to people, guess what, it's going to drop down. And then. And then you, Rob, tell me, oh, you don't like this answer, and I fix it, and it screws up something that Josh wants to do. Right. Uh, but you have to be very methodical about this. You can't just say, oh, I gave it to 10 people. They said it feels good. It's good, Jeff. I'm like, yeah, let's go. That is not a robust evaluation. And by the way, people don't like doing that. They like to build stuff they don't like.

Speaker A: That's pain, that's not dopamine. But it's so crucial to have that. And what you're talking about, you know, I just think of as custom evals, right? Because you need them contextualized. You know, you don't. General evals aren't going to work. You need custom evals that are about your organization. You need to build a corpus that's organizational, centric. That's oagi, right? Not AGI. Organizational AGI. And, and then your point is, right, like you got to keep it static so that you have a baseline to play off of as you, as you think, as changes happen, are you improving or is it getting worse or are you just burning tokens for no reason? And this kind of goes across. It's a meta problem, right? It's there. And anything else, what if you, what if you rev an agent in some way or add a new agent, uh, to replace an old one? How do you know that it's better?

Speaker C: Well, two things, number one to your point, and this is what I recommend to people all the time, is like, you know these models, they're upgrading them every six months, sometimes every three months, right? So how do you know? I mean, and they're like, it's better. Well, is it better on my corpus or are you actually degrading? Right. Because the way I just structured my rag process, it's not, it's sort of in a weird way that it's actually acting strangely when I'm improving the model. So your improvement is not my improvement. But the other thing which I think is super important is that we never get in trouble with the core use cases, right? We know what they are, we train them. It's the 2, 3, 4 standard deviations. It's those edge cases. And again the only, there might be in some cases, five people in your entire thousand person organization who are actually capable of finding those use cases. Right? You need, if you want good, I always say if you want good AI, you need good data and really smart people.

Speaker A: Mhm.

Speaker C: Because bad data with dumb people makes really dumb AI. And that is where the, I mean, sometimes you can build stuff in days, as you know, it could take you months of testing to get to the point where you are really confident. And when we start talking about agents, I'm just talking about systems that just give you an answer, let alone systems now that are going to be acting on your behalf, passing on information to something else. Everyone is very focused on agents and which they by the way, should be. But they're missing out on this conversation we're having right now because this is where the value really lies.

Speaker A: So I call it Use case zero. Because I always feel like if you're going to build an optimus Robot, that use case 0 is that Optimus can build robots, not fold laundry. And there's no question in my mind that's what they're working on in the knowledge management space. It's knowledge that can train and learn itself. Right. It's not about like, oh, you know, a pipeline. We started, we assembled 30 people, we went through, accumulated all the knowledge, dumped it into the system, and then we keep updating it. Like use case 0 is how do we create a system that accumulates knowledge on its own and maintains its own knowledge? Because it can. Right. It's just we're not used to thinking in that way.

Speaker C: I mean, yeah, but I think we have to think about these things as a human AI, you know, collaboration.

Speaker A: Right.

Speaker C: I think we are very far away between, you know, hit the button, redo my strategy, build a new product and send it out to all the clients. We should be thinking about this agentic environment as a world that is leveraging from machines to do what machines do extraordinarily well. That is collaborating and being supervised by people, which is being collaborated and supervised by machines. Right. Like we have to think about this world, that we're handing things back and forth to each other, we're engaging, we're talking and we're using humans where humans are best and we're using machines where machines are best. And I think part of the challenge is we're thinking about agents as purely agentic machinery. Right. And the reality is in organizations it's going to be, it's going to be an interface where, you know, and I'll just use a simple example, you may have an um, an account opening process that is, that is agentic. And then at the end of that process you may have a human being come on top and then validate everything. And then you may even before they submit, have a different model that comes in and, and checking what the human does and only after all those things have happened together. Or someone may do something that feels outside of your ethical policy and that either maybe gets stopped or maybe gets escalated to a human who then looks at it and makes evaluation. So it's going to be this interplay. And in all honesty, most organizations are not really built for any of that yet.

Speaker A: I think you make a good point. This can't just be autonomous, it's got to be, uh. It's funny, like people will say human in the loop, but I almost think of it in the reverse. It's agent in the loop because at some point it's the agent saying, hey, I was asked this question. I didn't have an answer. I sought the answer. From where? Wherever, data source or ultimately it came from a person. No matter what, it always comes from a person. No matter LLM's knowledge always came from a person. There is no such thing as Information that didn't come from a person. It's just about whether it recently came from a person and which person did it come from. But there's nothing in the system that didn't come from a person at some point, and therefore it can turn around and say, like, hey, I was asked this question. I don't have this information. I reached out, got this information from this person. Now I need to know somebody who needs to validate that before I institutionalize that into my memory. Um, which makes sense. That's absolutely true. Because some person has to be responsible for that information. Because if it's wrong, there has to be a human that has something to lose that someone can go to and say, hey, how did this information get propagated?

Speaker C: Well, let me give you a thought experience. I was at a conference the other day, and I raised this question. So let's make believe that you have, uh, a terrible cough, and you go to your doctor who accesses your agentic medical profile, who then uses her agent to diagnose your cough, um, who then uses another agent to prescribe a medicine to you or which then is received by your agentic pharmacy. Maybe no humans involved now, by the way, um, who dumps out some pills into a little, uh, plastic container, which then an agentic, uh, drone picks it up from your pharmacy and flies it to you and drops it on your front door, and you take that medicine and it kills you. Whose fault is it, like, we are living in this world now, which we are not prepared for that problem. Right. That assumes that there is accountability and transparency in every one of those agents that, you know what information was passed from your. I mean, maybe your records were screwed up. Maybe you lied in what your conditions were. Right. Like, that issue is very, very complicated. Which is one of the reasons I don't believe that agents from an external perspective are going to move nearly as fast as we think they are because of these types of challenges. Um, and we have mcp, but the reality is these issues are unknowable right now.

Speaker A: Yeah. MCP has nothing. Yeah. It doesn't solve the problem that agents don't have anything to lose. You have to have an entity that has something to lose in that chain that, in our case, we put people with paychecks. Right? That's right. And we pay them lots of money. And we say, okay, if this happens, you'll be held responsible and you'll lose your paycheck. There has to be something to lose. Agents don't have anything to lose in this. In this whole Equation, Um, and our just society depends on a human being accountable. Otherwise your whole point is like, yeah, if it's the government that gave you the pills and you're like, you can't sue the government, then you're screwed.

Speaker C: No, but this is why I think in, like our kids will be flexing in 10 years in a bar saying that they've got, you know, they own and are responsible for the top agents at their company. Right. That's going to be exactly. You're going to manage people and you're going to manage an agents. But if your agent, you know, to your point, Rob, if you're running the agent that does prescription management at the pharmacy and it comes out that your agent sent the wrong drugs out to Josh, you're going to lose your job.

Speaker A: Yeah. And somewhere in that chain is going to be you swiping. Right. And not reading it.

Speaker C: Well, that's a whole thing.

Speaker A: It's gonna be like, are you gonna let me ready to deliver? Do you approve? And you're like, yep, you know what? I'm too busy to read it. Yep. And then, and then, boom. You're gonna be like, shit, next time I'm gonna read that thing.

Speaker C: Well, this is what I, I, I wrote something about this recently. Um, about that. What, what's gonna, it's not gonna be like one moment in time. We're going to start to seed our judgment over time and it's going to happen in very small increments.

Speaker A: Right.

Speaker C: And it will not be a problem until it is a problem. And I will predict, I will predict that we are not far away from some very bad thing happening in our world. And the CEO is going to get on CNBC and said, well, we had an agentic problem. And the reality is, you know what? Your agenda problem is your problem. Right. Like, it's no, it's no different than if I hired a bunch of summer interns and gave them, gave them vodka on a Friday afternoon and they havoc. Right, Exactly. It's my job. Like, we can't. And I think there is a.

Speaker A: And as that manager, it's why you don't get paid seven bucks an hour. Because if you have nothing to lose, you will give them vodka. Like, whatever. Lose my job. I can get a better paying job than McDonald's.

Speaker C: That's right. Which is, I mean, the argument. And by the way, I don't know if I prescribe to this, but the argument is that we are going to create so much complexity in these agents that we're going to need better Higher quality people supervising all this because the job is going to be really, really hard. It's going to be a problem.

Speaker A: It's like lawyers that read the fine print, you know, like you've got vigilance. Right. You got to be like, uh, a lot of decisions really fast. We're almost like all going to become like Obama and his gray suit, blue suit thing. Like, I just less decision he has to make that day on something unimportant because he's got so many decisions to make in that day that matter.

Speaker C: And, and the problem is that the risk. And I think that's a good example right before, there was only so much damage that, that the three of us could do in a day. M. Right. When you're managing an agent that's going to send out all the tax notifications to your 4 million client base. Right. Or is going to process those prescription drugs. Right. Your ability to do harm at scale is so far so much bigger. Which again demonstrates why you need. I mean it's very likely that firms are going to have five or even six independent monitoring controls to include humans. Like you may have five agentic monitors and two people.

Speaker A: Yeah.

Speaker C: Overlooking this because if you don't give them the right dialysis approach, people are going to die. Right. Uh, and really putting in the level of controls, which gets back to the conversation we started with like good data foundations, good semantic layers, good controls. And again, like, we're good for 15 agents, we're not good for 15,000.

Speaker A: Yeah. And that context now to kind of bring it full circle is not. We realize this knowledge management isn't just for AI agents. It's for the people that have to make these decisions too. They need the same knowledge to know that the decisions they're making make sense because. Because the consequences now are so great that context matters more.

Speaker C: Well, and the good, the good news. The good news is you can, if you properly structure your data, if you have the right sources and they're of high quality, you can help your employee base make those better decisions. Right. But again, it comes down to all the stuff we've been talking about.

Speaker B: Well, that ties back to another aspect of use case zero that we've talked about. We trained an agent on all the knowledge in our book and that was exciting. And then we're just staring at another empty interface that's like, ask me anything that's interesting, but it's not terribly helpful. Like what, what that interface should really be doing is acting on behalf of the person staring at it. Like, who are you? Have you read the book? Did you just buy the book or did you read it all? Let me see how much you know. And now let me teach you everything I know about agentic AI and let me teach it to you in a way that you're going to understand and that you're going to grow and you're going to be engaged and interested and you'll learn it all. But we'll take you through it the way that you're going to like it or that you're going to learn the most from. And I think, you know, that also opens an opportunity where if, if agents inside of, uh, an organization are actively teaching people not just about how agentic AI works, but also the corpus of knowledge about that organization and bringing them up to speed, there's still an opportunity there too, for a feedback, uh, loop into the system. Because people are going to spot things that are wrong. People are going to educate the system as well. They're not just going to be taking the education.

Speaker C: Well, you make the point which I've said, I teach at Columbia. I was talking to students about this the other day and the comment someone said, well, someone sort of said, well, AI can make you really dumb. And, and I said, yes, AI can make you really dumb if you want to be dumb. It can also make you incredibly smart if you want to be smart. Right. Both of those things are factually true. Right. And I think to your point, AI can teach and inform and challenge and elevate your knowledge base in a way that is just extraordinary. Uh, on the other hand, if you want to be lazy and act smart, I can do that too. And I think one of the. And I know we're getting philosophical here, but one of the questions is how do we build environments and cultures where people are doing the former rather than the latter? And how can AI play a constructive, um, a constructive tool in that knowledge expanding? Right. So I know more stuff about the world and I'm more thoughtful and I'm asking better questions as opposed to. Just to your point, Rob, you're like, done, done, done, done, done. And I'm on my iPhone looking at espn.

Speaker A: Yeah. I think, you know, as I look through this, there's one lens to see that humans are going to be the bottleneck in all of this. Not machines, um, but it needed necessary bottleneck. And if you want to make them not more productive but more productive at making decisions, then you have to make information more accessible and available to them at the time they're making decisions. And this kind of ends up at Decision science being like, at this, you know, use case zero isn't just the system learning on its own. It's the system teaching what it learned back to humans so that they can make better and faster decisions. Which is. It's an interesting loop, right? Human in the loop is like, I'm going to teach the system and now the system's going to teach me.

Speaker C: But there's, I mean, again, not to go back to the human thing, but that's when we're at our best, right? We're at our best when the three of us are doing things, listening, being humble, being open, and then saying, no, I think you're wrong. No, I think you're wrong. I think you're right. That is the essence of learning and challenging. And that's why I personally, I don't know if you guys saw I built a bot over the weekend. Talk about wasting tokens. Um, that does. Um, I created a board of advisors and it's on MacMillinai.com, and it's a board of advisor. And what I did is I took my favorite 20 leaders in the world, like from Teddy Roosevelt to Richard Feynman, Steve Jobs, Sun Tzu. Right. And Nelson Mandela and experts. You can ask, by the way, it's about a penny Aquarius, so people can go crazy. I got plenty of tokens. Um, and what you can do is you can ask questions. And by the way, I, uh, have a look at it, it's called the Board of Advisors on my site. And, and what it will do is it will give you different perspectives. And by the way, some of them are actually different. Like Steve Jobs will say something that's different than fdr. Right. And I would argue that that process is really helpful. It's like, I can't call you guys and ask you your opinion all day long, but I can ask AI a lot. I can say, be Josh, be Rob, be somebody else and give me your point of view. And that elevates my game.

Speaker A: Yeah, we had this discussion, um, Josh, remind me. The genome project.

Speaker B: Dr. Lee Hood.

Speaker A: Yes, Dr. Lee Hood. So we had this with Dr. Lee Hood and we discussed this, right? And we were talking about the fact that the medical system isn't going to change at its core based on AI. What's going to happen is there's going to be AI abstraction layer that helps you manage the bureaucratic healthcare system that doesn't change. Um, and that's going to happen faster. It's just going to be a layer that will make appointments for you in the antiquated old system that they didn't get around to changing. And it makes a lot of sense. Um, but what we discussed is, yeah, it could be gps, right? Some people will be lazy and they'll just let the thing make decisions, but more than likely we're going to have different perspectives. And once you add that, once you add a couple of, um, agents, if you want to call them that, arguing with each other about whether you should eat the donut or not eat the donut, um, now you're forced to kind of read the arguments. And what's happening at that point is you're actually going from like the opposite. Now you're learning because the best way to learn is to hear different perspectives, right? So now you're like, should I eat the donut? Well, your psychologist says eat the donut, but your trainer says, don't eat the donut. Uh, now you decide. You get to decide on whether you'll eat the donut or not. Well, you've just learned something.

Speaker C: We're getting somewhat philosophical here, But I remember 15 years ago when I was working on machine learning stuff and I used to make the joke that smart people with better information make better decisions than smart people without it. And I think what you're highlighting is the fact that generative AI does this on steroids, right? Now, again, it's a question of what do you do with that information? Are you just doing whatever the AI tells you and says, okay, well, now tell me what the right answer is or are you applying critical thinking? And the power is that if you wish to apply critical thinking, you will produce better decisions without question. But if you do not, if you're lazy and do not apply critical thinking, you will make dumb decisions. Right? And I think this is where we as humans have to be very thoughtful about this issue. Right? It's a very existential question. And it's not letting the AI saying, oh, okay, well, just, it told me to eat the donut, I'm just gonna eat the donut.

Speaker A: Yeah. And that's just up to us, right? It can't force us to make choices. Right. I, I think, you know, we all, this is another anthropomorphic concept we, that we apply that it gets dangerous because, you know, we think of things in singular. We have like one thread in our consciousness and we talk and we have an opinion. We have an opinion, right? And so we think there's a right and wrong answer to a lot of things because, well, run from the lion or don't run from the lion. Really matters to us. But AIs are multi threaded. There's. There's 20, 100,000 right answers to that question, should I eat the donut? There's not one right answer and one wrong answer, there's a hundred. So now you have to pick the one right. One right answer out of a hundred right answers. There's no right answer. Um, and that's where knowledge gets complicated, is you're like, well, there is no answer. There's just the answer you want to pick.

Speaker C: Well, the other thing is, I mean, you could imagine a world in many years where you have a choice. It is a probabilistic curve of outcomes. You then experience that outcome and then it goes back into your algorithm that adjusts. So the next time, maybe you don't eat the donut, right, or maybe the next time you do because the factors have changed. But the idea of, I mean, what makes great smart people is they learn from experience, right? There's a reason, you know, 18 year olds do stu. Stupid things sometimes, right? Because they don't have that knowledge. But, you know, I think one of the things that's so powerful about AI is it really can capture those learnings in a way that, that we as humans, you know, we forget stuff that we don't remember. But being able to sort of build. And I ultimately think, you know, we will have these algorithms and they'll have the Rob algorithm, the Josh and the Jeff algorithm, right? And these will be telling us all day long. You know what? Um, maybe you can eat the donut today, uh, but maybe tomorrow you shouldn't for the following 50 million different reasons.

Speaker A: Yeah, and what's the likely thing? The likely thing is it says, you know what? Just give me the data. The data is going to say, you know what? Eat the donut. You'll feel good for 20 minutes and then you'll feel like shit in an hour. Your decision.

Speaker C: That's right.

Speaker B: That's right.

Speaker C: And we'll probably still eat the donut, but that's a different problem.

Speaker A: Yep. Well, cool. This was great conversation as usual. Never enough time. We barely crack it open. Um, but thanks for joining us.

Speaker C: No, my pleasure, guys, Always a pleasure.

Speaker B: Thanks for hanging out with us on Invisible Machines. Remember to like, subscribe and or follow us wherever you get your podcast media. Thanks to onereach AI for their support and much gratitude to the many people working behind the scenes to make this podcast great. Until next time, Awagi.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Structured Truth for Enterprise AI: PaligoSourceForge Podcast · on RAG (Retrieval Augmented Generation)86 / 100
  • Automatic Data Pipelining: One More Turtle AheadAdventures in DevOps · on RAG (Retrieval Augmented Generation)83 / 100
  • Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769The TWIML AI Podcast · on RAG (Retrieval Augmented Generation)82 / 100
  • How MLOps and LLMOps Drive Consistent Results (Kristen Kehrer)What’s the BUZZ? - AI in Business · on RAG (Retrieval Augmented Generation)82 / 100
  • Album 8 Track 17: The Mythology Behind a Great Marketer w/Louis MonoyudisBrands, Beats & Bytes · on Morgan Stanley81 / 100
  • From Sales to Spare Parts - Festo´s AI solutionsIndustrial AI Podcast · on RAG (Retrieval Augmented Generation)80 / 100

More from Invisible Machines podcast by UX Magazine

All episodes →
  • Nuclear Fusion, No Power Lines ft Jonathan Frankle
  • When Agents Have Wallets, Trust Is Currency
  • No Strategy Without Vision ft Brian Evergreen | Invisible Machines
  • The Confabulation Machine ft. Evan Ratliff of Shell Game | Invisible Machines Podcast
  • Crisis Is Your Opening | Marina Nitze | Invisible Machines
Explore the best B2B AI & Data podcasts →
All Invisible Machines podcast by UX Magazine episodes →