The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Disambiguation
Disambiguation artwork

The End of One Model to Rule Them All: Why Enterprise AI Is Going Small, Specialized, and Multi-Model

Disambiguation · 2026-06-24 · 42 min

0:00--:--

Key moments - from our scoring

Substance score

65 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality12 / 20
Guest Caliber15 / 20
Specificity & Evidence13 / 20
Conversational Craft11 / 20

The enterprise AI market is experiencing a fundamental shift away from the "one model to rule them all" philosophy toward specialized, smaller models trained for specific tasks. Calvin Cooper, co-founder and COO of Neurometric AI, explains that this represents a natural progression from pilot projects to production-scale implementations where cost becomes a critical priority. AT&T's public case study - cutting costs by 90% while scaling from 8 billion to 27 billion tokens daily - exemplifies this transition. Cooper argues that roughly 75% of enterprise AI tasks don't require frontier models with advanced reasoning; structured work like classification, extraction, routing, and summarization can be handled far more efficiently by small language models. Neurometric AI addresses this through their leaderboard benchmarking against MTEB, pricing models around fixed monthly costs per endpoint, and orchestration approaches like their "coding swarm" that coordinates multiple task-specific agents rather than relying on a single large model. The conversation covers why GPU utilization and inference efficiency barely registered as concerns last year but have become first-order priorities today, how companies can determine which tasks belong in small versus large model categories, and the organizational maturity stages required to transition from monolithic AI services to dynamic, task-matched approaches. Cooper also discusses how AI augmentation in software development has expanded rather than eliminated engineering jobs, with developers now managing swarms of coding agents rather than being replaced by them.

Key takeaways

  • →Small, specialized models fine-tuned for specific tasks deliver faster speed, lower costs, and better accuracy than large foundation models for 75% of enterprise AI use cases like classification, extraction, and summarization.
  • →As enterprises scale AI implementations from pilots to production, cost management becomes a critical priority, driving adoption of multi-model orchestration strategies and hybrid approaches that blend task-specific small models with fallback to frontier models.
  • →Neurometric AI's leaderboard and marketplace demonstrate that different models perform better or worse on different tasks, with ensemble approaches and inference-time tactics sometimes outweighing raw model choice.
  • →Organizational maturity in enterprise AI requires a mindset shift from proving immediate ROI to learning-by-doing and experimentation, with success dependent on top-level commitment and expertise rather than technology capability.
  • →Fixed monthly pricing per endpoint rather than consumption-based token pricing enables cost predictability, making AI economics sustainable for production deployments at enterprise scale.

In this episode

  1. 1The Journey from Venture Capital to AI Infrastructure
  2. 2Why Small, Specialized Models Beat Large Foundation Models
  3. 3Enterprise Cost Pressures Driving Multi-Model Architecture Shifts
  4. 4Task-Specific Models and Predictable Pricing Models
  5. 5Coding Swarms: Orchestrating Multiple Agents for Development
  6. 6AI Maturity Stages and the Path from Experimentation to Scale
  7. 7Organizational Mindset Shift: Learning Over Immediate ROI

Mentioned

Neurometric AICalvin CooperMichael FauscetteNCT VenturesRoveAnthropicOpenAIClaudeLlamaSalesforcePilot WaveAT&T

Guests

Calvin Cooper

Topics in this episode

Small language modelsClaude (Anthropic)Foundation modelsNeurometric AIMulti-model orchestrationMTEB benchmarkInference-time computeGPU utilization efficiencyTask-specific fine-tuningCoding swarms

Questions this episode answers

Why are companies moving away from large language models to smaller specialized models for enterprise AI?

Smaller, specialized models fine-tuned for specific tasks are faster, cheaper, and more accurate than large models that contain knowledge about all of world history and science. At scale, cost pressures force companies to optimize - AT&T cut costs by 90% while scaling by switching to specialized models for task-specific work like classification and extraction.

What percentage of enterprise AI tasks actually need frontier models like GPT-4 or Claude?

Approximately 75% of enterprise AI tests don't need frontier models or advanced reasoning; they involve structured and repetitive narrow tasks like classification, extraction, formatting, routing, and summarization that can be handled efficiently by small language models.

How does Neurometric AI's pricing model work compared to token-based pricing?

Neurometric AI charges a fixed monthly price per endpoint for task-specific models with a fallback to larger foundation models, providing cost predictability instead of variable consumption-based pricing, which aligns with how enterprises have already transitioned to fixed pricing for text messaging and data.

What is a coding swarm and how does it differ from using a single frontier model for development?

A coding swarm orchestrates multiple task-specific small language models each trained for different parts of the development lifecycle - testing, documentation, code generation - rather than relying on a single frontier model to handle all tasks, resulting in better efficiency and cost control.

What's the biggest failure mode companies encounter when adopting enterprise AI?

The failure mode is treating ROI as the first KPI; instead, successful companies prioritize learning-by-doing and experimentation, requiring top-level commitment and expertise to drive value through a multi-stage maturity progression.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode delivers a solid set of ideas around multi-model orchestration, cost-driven architecture shifts, and the move away from frontier-model-only strategies. However, much of the core thesis - that smaller specialized models beat large ones for narrow tasks - is intuitive rather than surprising, and the conversation often restates the same core argument (surgeon analogy, AT&T example) without drilling deeply into novel technical or business mechanisms. There's moderate novelty in the infrastructure efficiency angle and the Pareto law framing, but significant filler around venture mindset and crypto analogies that don't add operational insight.

why would you hire a surgeon to schedule an email? You're going to overpay for that. Why would you do that's just obvious, right?
75% of enterprise AI tests don't really need a frontier model

Originality

12 / 20

The core argument - small specialized models for narrow tasks, multi-model orchestration, efficiency over scale - is now widely circulating in AI infrastructure discourse. While Calvin frames it well through the surgeon analogy and systems thinking, the underlying contrarian position has been articulated by others (e.g., Andrej Karpathy's efficiency focus, others in MLOps). The venture capital lens on thesis-driven thinking and the venture mindset critique feel tangential rather than original to the AI architecture discussion.

the future is less like Mission Impossible if you're a fan and more like Tron
calling a bubble is is kind of like almost retarded

Guest Caliber

15 / 20

Calvin Cooper is a credible operator with genuine founder experience (Rove fintech, Nasdaq exit), venture capital background (NCT Ventures), and is currently building a company in the space (Neurometric AI, co-founder/COO). He has hands-on experience with enterprise customers and cost optimization problems. However, he is not a marquee name with massive-scale public company AI deployment experience, and the episode lacks the depth of a CTO or principal engineer who has driven large-scale model migrations at FAANG-scale companies.

I founded Rove as the CEO. Consumer Fintech products sold it about three years ago, with the plan to take it public via direct listing or the Nasdaq, which we did
I met Rob, the CEO, founder of Neural Network, last year. We all decided to go full time in August

Specificity & Evidence

13 / 20

The episode includes some concrete examples: AT&T scaling from 8B to 27B tokens daily while cutting costs 90%, one customer achieving 4x latency improvement and cost savings with a Llama model, 115 task-specific models under 20B parameters in their marketplace, and the Claude API pricing shift forcing prosumer developers to rethink architectures. However, most claims lack granular detail - company names are often vague ("one customer," unnamed enterprise team), financial figures are sometimes rounded or illustrative, and technical benchmarks are referenced but not deeply explored.

one example...they've got two models in production from different model providers. And then we ran analysis on dozens of different alternatives and showed where one of the llama models could perform and meet the accuracy requirements and the performance requirements, but also do so at a cost improvement in four x latency improvement
AT&T chief data science officer...as they scale from 8 billion tokens a day...to 27 billion tokens a day. They had to...cut costs by 90%

Conversational Craft

11 / 20

Michael asks solid foundational questions and does follow up on key points (cost pressure shifts, maturity models, the Coding Swarm product), but rarely pushes back or probe uncomfortable territory. The conversation is friendly and exploratory rather than adversarial. Michael mostly validates Calvin's framing and occasionally adds personal anecdotes rather than challenging claims. There are few moments of genuine tension or disagreement that would deepen the analysis. The host also allows Calvin to go on tangents (venture philosophy, crypto, Luddites) without redirecting firmly to enterprise substance.

Yeah, it's interesting. I mean, just looking at history. It is very typical in tech, especially to call oh, this is a bubble.
Well, and you know, over the last quarter or so, even the the stock market fallout reactions...a lot of pushback from the enterprise around. Well, yeah, we need cost predictability

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

calvin67michael46model37models22different21interesting16small13market12problem12language11started11task11first11data11starting10specific10

Episode notes

In this episode of the Disambiguation podcast, host Michael Fauscette talks with Calvin Cooper, Co-Founder and COO of Neurometric AI, about why the dominant narrative of scaling ever-larger frontier models is giving way to a more practical reality: smaller, specialized models fine-tuned for specific tasks that are faster, cheaper, and more accurate for the vast majority of enterprise AI workloads. Calvin started his career in early-stage venture capital at NCT Ventures in the Midwest, then founded Rove, a consumer fintech company he took public via a Nasdaq direct listing. Now he and Rob May have co-founded Neurometric AI, which builds task-specific small language model infrastructure. They went full time in August 2025, at a time when the dominant narrative was still "scale compute, scale larger models, AGI," because they were seeing something very different in the research and in practical enterprise deployments.

Full transcript

42 min

Transcribed and scored by The B2B Podcast Index.

00:00:09:09 - 00:00:10:00 Michael Welcome. 00:00:10:02 - 00:00:32:01 Michael Welcome to disambiguation. I'm your host, Michael Fauscette. Each week, we interview experts in artificial intelligence, generative AI, and business automation to help business leaders understand how to use these tools for the biggest business impact.

Today's episode is the end of one model to rule them all. Why Enterprise AI is going small, specialized and multi-modal and joined by Calvin Cooper, Co-founder and COO of Neurometric AI. Welcome, Calvin. 00:00:45:08 - 00:00:50:10 Calvin Hey, great to be connected.

I'm excited to dive in. Thanks. Great. 00:00:50:12 - 00:01:17:00 Michael Yeah, I think this will be a fun conversation.

It seems like it's one that has come up a lot lately, which that's a good thing. I guess. People are starting to talk about alternatives to the large language model fallacy. So before we get going, though, it's just you have a really interesting background and I know, you know, you venture capital at NCT ventures, founder, Rove fintech company, you know, through a Nasdaq exit.

00:01:17:00 - 00:01:37:13 Michael So you have some experience that some of us interpreters never actually get to interpret at Iowa State. Now co-founder neuro metric AI. So what did this journey look like through investing, building, exiting and now really leading you to focus on, you know, inference time compute and small model orchestration? 00:01:37:15 - 00:02:05:20 Calvin Yeah, I've always just followed curiosities to their logical conclusion.

And I can't just not do a thing. So, yeah, I started my career in early stage venture capital in the Midwest at NCT ventures. So help Raise fund two brought in a few of the LPs led deals across several thematic areas, and then went on to to found Rove as the CEO. 00:02:05:21 - 00:02:37:14 Calvin Consumer Fintech products sold it about three years ago, with the plan to take the choir public via direct listing or the Nasdaq, which we did, and.

Taking a look at the market with the venture hat on. Like what's changing in technology, what's changing in consumer behavior, and where where are the real new opportunities, what's in the adjacent possible. 00:02:37:14 - 00:03:22:28 Calvin And it became very clear that inference was going to be the largest technological market opportunity of our lifetime. And I met Rob, the CEO, founder of Neural Network, last year.

We all decided to go full time in August at a time when the dominant narrative was still scale, compute scale. Larger models, AGI. And we were seeing something very different in the in academia and the research and just, you know, reasoning to how do you build practical AI systems that scale and a cost effective way off problems for enterprises. 00:03:23:01 - 00:03:26:12 Calvin And so we started investigate that.

00:03:26:14 - 00:03:47:22 Michael Interesting. I mean, your tagline actually feels almost counter message to what we all sort of accept today or we hear a lot today anyway. And that's the small model specific jobs, paperwork done, not tokens. I mean, that's that's a bit of a challenge in a way to the to prevailing narrative.

Right? Bigger models are always better. Can you, can you break that down for us a bit? 00:03:47:26 - 00:03:54:03 Michael And, you know, what are you seeing in the enterprise market that convinced you this was the right pat?

00:03:54:06 - 00:04:22:24 Calvin I mean, so what is being seen now is validation of a hypothesis from before, right? So a lot of things become obvious when you just think critically. So they often say like contrarian contrarian thinking. I think people approach that the wrong way.

Often they're thinking they want to be a devil's advocate, which is just the inverse of a dominant narrative. 00:04:22:26 - 00:04:58:03 Calvin But if you just think about it critically, do you need to hire a surgeon to schedule an email for, you know, you're going to overpay for that? Why would you do that's just obvious, right? So if you think about AI systems similarly, why would any technology that knows all of world history and you know, different aspects of science and has to navigate to answer your question, why would that be the most efficient way to to answer a question or solve a problem?

00:04:58:07 - 00:05:34:00 Calvin So that's just obvious. And, you know, I'll leave the FCS to talk more about the technical limitations of that. And we published a lot of stuff on our Substack, but it just becomes obvious that a smaller, specialized model that's fine tuned for a specific task is going to be faster, cheaper, and more accurate. Right.

So when you're going up against the dominant narrative, that's well funded, right? 00:05:34:01 - 00:06:02:28 Calvin A lot of people have shortcuts in their brain. How are you going to be the anthropic or open AI? You have to think at the systems level.

Like, you know, you got to think about incentives. You got to think about the science, the trajectory, the JSON possible, and then you got to figure out what's the smallest first thing you can do to start to prove the thesis and learn as well, because some things you're wrong about, right. 00:06:03:00 - 00:06:26:18 Calvin And you expose that as quickly as possible. So you want to ship in a week or a month.

So the first thing we shipped was our leaderboard, right. So we had put up a leaderboard against CR Marina Benchmark which is a public benchmark. Salesforce. And you Marina.

It's because it was less academic and more useful for actual work, right? 00:06:26:19 - 00:06:58:20 Calvin Right. And so and you can break down to the task label like like named entity disambiguation and such. And so what we showed was is that different, different models perform better or worse at different tasks.

There's no like universal good model. And also often how you approach the model or different inference time tactics or algorithms paired with different models. 00:06:58:20 - 00:07:26:18 Calvin Those combinations can be as impactful or more than simple model choice, right? And so we started to show that and then started to show how ensembles, models and different approaches and sometimes small language models can perform better.

And so we started to prove that. And then a couple months later signed a few early adopter customers and then ran. 00:07:26:22 - 00:07:57:16 Calvin You know, they are a journey that most people are just get something to work at all costs, and then you start to scale and they need their, you know, first issue and getting it to work, maybe latency requirements meeting those. And then after that it becomes price.

Right. So that's just obvious following you know reason. And so you know one example I can think of it's they've got to my models in production from different model providers. 00:07:57:16 - 00:08:32:01 Calvin And then we ran analysis on dozens of different alternatives and showed where one of the llama models could perform and meet the accuracy requirements and the performance requirements, but also do so at a cost improvement in four x latency improvement.

So then now you're you're moving from hypothesis to data and research and publishing that to making an impact with the real customer. 00:08:32:03 - 00:09:04:22 Michael Yeah. Now that makes sense to me. I mean, I'll say eventually I came to this similar conclusion that size and the language model to the task is actually the correct approach.

But but it's funny because I sort of took a circular route, I guess, but because for me it was it was a realization as I started to look at some really specialized functions, that it just didn't make sense to have this large generic model when what I needed was something highly specialized and tuned to the exact either vertical or process or whatever. 00:09:04:24 - 00:09:31:19 Michael Right? So I can see a lot of a lot of the argument as I as I looked through this and went through my own process over the last year.

But I think a lot of people still, you know, kind of accept that narrative that, oh, big is always better, right? And, you know, when we were we were talking about this in our prep conversation, you know, you said you started last summer to, you know, that there was almost no awareness around GPU utilization or efficiency for inference. 00:09:31:19 - 00:09:51:26 Michael Compute. This whole idea of consumption based pricing, even, you know, leads to that.

And the conversations really shifted a lot this year. I mean, what what changed? And what do you think CTOs and CIOs are telling you now about cost pressure? That's really forcing them to rethink their model architecture?

00:09:51:28 - 00:10:20:25 Calvin I mean, many different things. I think one public article to read is there's one AT&T chief data science or chief data officer was in VentureBeat a few weeks ago or a month or two ago, and he was interviewed talking about how as they scale from 8 billion tokens a day, they had to, you know, to to 27 billion tokens a day. 00:10:20:26 - 00:10:48:13 Calvin They had to really rethink their architecture and cut costs by 90% even as they scaled. Right.

So they cut costs by 90% as they went from 8 billion tokens today to to almost 30. And that's what changed, right? Last year we we even trying to understand if AI was going to make an impact. I think people who are closest to are like, oh yeah, it's obvious.

00:10:48:13 - 00:11:30:07 Calvin But enterprises were failing to prove ROI on projects, and now you're starting to see real scaled implementations at the enterprise level that's driving impact. And when you move from cute pilot project to production to production, that's making an impact in scaling economics. Test the rest, right. And paradox applies here.

And so your Costco from minuscule negligible to first order concern priority. 00:11:30:07 - 00:11:59:22 Calvin This is a problem that needs to be solved and can be solved. Right. So to make it very tractable for listeners, open cloth you're aware came out since 24 seven agent tech runtime where it's not just an LLM chatbot.

You have an agent that is running 24 over seven is actually useful, and people who are vibe coding with cloud code can even get up an agent in a virtual machine or a mac mini. 00:11:59:22 - 00:12:38:02 Calvin And so let's say you're using it for content automation and meeting scheduling and something like that. Right. And your costs are going from, you know, tens of dollars to hundreds of dollars a month and more.

And you don't know that because you're using your code subscription. But anthropic knows it and it's it's running up their bill. It's messing up their economics because they're subsidizing the the costs that it takes to provide this intelligence to you with venture capital. 00:12:38:02 - 00:13:09:12 Calvin And and so they say pump the brakes.

You got to use the API. And we're going to charge you by the token. And all of a sudden your project that is delivering value to your your life and workflows as an open client, since now goes from your subscription payment to hundreds, thousands, $10,000 bill a month, it's $100,000 agent now, right? 00:13:09:12 - 00:13:39:07 Calvin So you got to do what you orchestration and multiple models and Claude for advanced reasoning.

But then task specific small language models for certain things or open source models that are cheaper or things that are self-hosted. So you have to start to explore the whole framework differently. So it's just obvious that that's the next step. And I think our government's even finding that out.

00:13:39:09 - 00:13:44:07 Calvin You know, the DoD saga is important to to watch. 00:13:44:09 - 00:14:05:01 Michael Well, and, you know, over the last quarter or so, even the the stock market fallout reactions to, you know, particularly in the SaaS narrative of, oh, you know, the seat pricing model is dead and we have to go to value based or we have to go to consumption based. But a lot of a lot of pushback from the enterprise around. 00:14:05:02 - 00:14:24:24 Michael Well, yeah, we need cost predictability in this doesn't lead there.

Right. So I get the I get the reaction. But you know, it is a it does really lead down this path that you're talking about. Okay.

Let's think about how we size this to make sense in in each of the process environments that we're really working towards. Right. 00:14:24:26 - 00:14:53:16 Michael So I mean that resonates. But in your marketplace I know I saw like 115 tasks, specific small models, something like that.

There was a large number under 20 billion parameters. And, and they're all fine tuned to a single job or specific job, you know, even some auto small language model creator that generates custom models on demand. I mean, can you talk a bit about how that works in practice? 00:14:53:16 - 00:15:06:06 Michael And, you know, if I'm an enterprise team, what, you know, with a with this specific workflow, what does the process look like to go from I have this problem to I'm ready to deploy this specialty model.

00:15:06:08 - 00:15:33:14 Calvin So that that this comes naturally from our last topic. Right. So as you're seeking more cost predictability, right. That's that's the next shift.

It's and even if you zoom out it makes sense. Like think about text messages you were paying per text and you have unlimited text. You have bandwidth and now you have unlimited bandwidth and data. Right.

00:15:33:15 - 00:16:08:00 Calvin Similar with AI, the progression is towards predictability. And so we see that paradigm shifting. And we decided to ship a solution into the market price that way. So you're paying a fixed monthly price per endpoint right.

Specific model job to be done. And then your fallback is to a larger language model, maybe a foundation model like provider like from anthropic or open AI. 00:16:08:03 - 00:16:38:12 Calvin And so when you blend the two you have more price predictability. And so we decided to ship that into the market and then ship a solution into the open claw ecosystem claw pack.

And we've got more coming soon that are going to, you know, target coding problems and tasks. And so instantly we saw tons of people in the timing was great because sometimes you don't want to be too early. 00:16:38:14 - 00:17:04:26 Calvin But right as we were doing this, that was when anthropic just so happened to ban the subscription use open call users. So this isn't just useful for open call users.

That's just one channel to ship into. And the timing worked out. But instantly you get a couple thousand API key requests. They were like, oh, let's turn it on paid and you turn on paid, and then people just start paying you.

00:17:04:26 - 00:17:34:06 Calvin And it's illustrative of what's possible. And so we're shipping into that prosumer alpha dev ecosystem, even as we're working directly with business customers that are early adopters. But proving that you can have a predictable cost around your agency workflow solution. 00:17:34:08 - 00:18:01:08 Michael I mean, it's an interesting argument that you guys have made around, you know, I think is like roughly 75% of enterprise AI tests don't really need a frontier model or, you know, with that depth of reasoning and, you know, and honestly, the large consumption of tokens.

But I mean, you know, structured and repetitive narrow tasks like classification and extraction, formatting, routing, summarization. 00:18:01:09 - 00:18:12:26 Michael I mean, how do you help companies figure out what task belong in the small territory or medium territory, and which ones really do need the big model, the frontier models? 00:18:12:27 - 00:18:52:25 Calvin Yeah, it's a heuristic, and I think Pareto law applies here. And I mean, if we're honest about things, the vast majority of things that humans do, not humans, but like companies do, is you move away from ambiguity to things that are you can create SOPs around right.

And as as you go from, you know, that ambiguity to more deterministic, more easier to eval, it's those things are things that are probably handle. 00:18:52:26 - 00:18:57:19 Calvin You can handle them with small language models. 00:18:57:21 - 00:19:16:01 Michael You know, one of the, one of the products that you guys have developed, you call the the coding Swarm. And it's a really interesting idea.

And in fact, even some of the times when I, you know, vibe coder play around with with Claude code, for example, it does spin up different agents to do different pieces, parts of the task. 00:19:16:02 - 00:19:50:08 Michael Right. But but your argument here is, you know that instead of using the single frontier model, you orchestrate a coordinated team of task specific small language models. Each one is trained for that specific part of the development lifecycle, you know, testing, documentation, etc.

and you know what led you to think that way? And how does it compare to kind of what we see when you're just using cloud code, you know, generally to, to, to code a project? 00:19:50:10 - 00:20:34:01 Calvin I mean, I think that's just where the market's going. You're seeing you're seeing that is the approach you have to take to get things done.

So there's a there's a profound shift happening. It's pretty incredible. And you mentioned that your vibe coding, I think one of the biggest what and sorry to take the question differently, but I think what's really interesting about what's happening with with AI augmented code development is the broader conversation and especially America, the fears that AI is going to take away our jobs. 00:20:34:03 - 00:20:49:09 Calvin But the first, best use case of AI that's proven to drive impact and scale is coding agents.

Right. And. 00:20:49:12 - 00:21:31:13 Calvin Copilots and things like that. And not just copilots now, like swarms of agents that can write code.

And you're now managing and overseeing those agents. And what happened to the software developers? Did they go right now they're in much higher demand. And companies that couldn't afford engineers now have software development capabilities in their organizations.

It's it's expanded the jobs in that category significantly and made and even made the people with the craftsmanship. 00:21:31:13 - 00:21:45:02 Calvin And, you know, as software developers or computer scientists even more valuable and sought after. And so I think that's just interesting. Yeah.

00:21:45:09 - 00:22:06:26 Michael I mean, this is really interesting. And I've, I've argued for a bit that, you know, you're not really talking at least at this stage, and I don't know what the future looks like, but but for now, it seems like we're not really talking about, oh, we're going to replace this large group of people. We in some cases we are, you know, changing their job somewhat. 00:22:06:27 - 00:22:34:15 Michael I mean, maybe a developer that used to spend all day coding, spends an awful lot of time doing modeling and, you know, thinking of of the outcome, the conceptual, you know, abstraction of what you want, right?

Versus I wrote the code to do that. But I mean, that's still that's still developer function, right? So in effect, we're you know, we're not really saying, oh, look, this large swath of jobs is going to go away next month because of this. 00:22:34:16 - 00:23:09:13 Michael Right.

Which definitely is hopefully that's comforting to some people anyway. But, you know, if you if you think about maturity and I've done a few maturity models around around organizational maturity, around technical maturity when it comes to to to AI, into AI particularly. And, you know, you talk about a stage for maturity in enterprise AI, which is like a transition from treating AI's a monolithic service, to starting to treat it as a dynamic resource where the right models match to the task.

00:23:09:13 - 00:23:27:24 Michael And a lot of companies are are just not there, right? I mean, they're much in much earlier phases, but what's the path look like if I'm a stage one company, I'm just now starting to, you know, experiment with this. What does it look like if I go from stage one to stage four and and where where are the pitfalls? 00:23:27:24 - 00:23:31:18 Michael Where do they often get stuck that you that you've seen?

00:23:31:20 - 00:24:03:15 Calvin Yeah. I mean, I think the commitment has to come from the top to see it through. Right. So the failure mode is the mindset that the first KPI is to prove some kind of ROI.

Now, the first KPIs to learn by doing. And so I think the first thing the level said on and I do some work with a private equity firm called Pilot Wave, which is purpose built for AI. 00:24:03:16 - 00:24:40:15 Calvin Rolex, founded by. I've seen the former chief data science officer of JP Morgan.

And the main gap is no longer is the technology capable. It's. Do you have the expertise of the the commitment to drive value and that process of closing that gap. And so I'll kind of make it personal a little bit like I grew up in the Midwest and Columbus, Ohio and started in the venture ecosystem there.

00:24:40:15 - 00:24:53:22 Calvin And there's a lot of great things that happen. But I remember one thing that used to bother me a lot, and my instinct was that I need to go to San Francisco once a quarter. So I started to do that. And.

00:24:53:25 - 00:25:16:00 Calvin Like you hear people say, oh, they can think that way because they have so much money to figure things out. Or and I'm like, no, maybe they have all the money because they think a certain way. Maybe we have it backwards and one of the frame works that's backwards is often that there's value to saying something won't work, or that it's all hype. 00:25:16:02 - 00:25:44:09 Calvin It's literally of no value to say that most of that will fail structurally.

Most venture deals fail. That's the venture capital portfolio construction model, right? Or to call a bubble seems like, oh, I called the bubble. This is a bubble that is not intellectually interesting at all.

Bubbles are the default mode. We always have bubbles. We had a railroad bubble, but that didn't mean that railroads weren't massively important. 00:25:44:10 - 00:26:15:04 Calvin Right?

So calling a bubble is is is kind of like almost retarded. And so like when you when you look at crypto as an example, it's like, oh, crypto didn't have any use case. Well it does. There's a multi-trillion dollar market now.

And just because it doesn't solve a problem for you in your country doesn't mean there isn't $100 billion a day in stablecoin transactions and making cross-border payments free almost. 00:26:15:06 - 00:26:45:02 Calvin You know, we hire engineers, and the best way to pay them globally is with stablecoins now. And so it's solving real problems. So similarly with AI, is the technological breakthrough real?

Absolutely. Is there a new adjacent possible 100% if you can't figure out how to drive ROI? That's a you problem. So I think the the first most important thing to answer this question is to change the frame of mind to experimenting.

00:26:45:02 - 00:26:53:07 Calvin And what did you learn? Did you ship something useful? That's it. As quickly as possible?

00:26:53:09 - 00:27:16:24 Michael Yeah. It's interesting. I mean, just looking at history. It is very typical in tech, especially to call oh, this is a bubble.

This is you know, this is really not real. There's, you know, all these other issues or whatever or even, you know, going back to industrial revolution, you know, we have the great Luddites to use as an example. 00:27:16:25 - 00:27:38:28 Michael You know, historically, people that push back on, on the on the technology change and didn't do well with the change management. But I, you know, I think it does make sense to me that, that that's our default.

But it's a narrative that is almost like an apology. It's like, oh, I'm sorry we created another bubble. Like, okay, that's it's it's not unreal. 00:27:38:28 - 00:28:12:01 Michael It's just, you know, it's perhaps growing faster than you want, but but yeah, that that resonates I think and you know, I know you guys have a you host a podcast inference time tactics podcast.

And you guys go really deep on the infrastructure side of this chips, accelerators, data center design. I mean, from those conversations, what are you seeing about how the physical infrastructure layer is adapting to support this shift towards a multi-model? 00:28:12:07 - 00:28:16:25 Michael You know, sizing the model to the task kind of approach to things? 00:28:16:27 - 00:28:47:25 Calvin Well, I think, you know, older GPUs become useful, other alternatives become useful.

So when you think about our AI industrial capacity in America and competing globally, we don't just need to build more data centers and advance the chips and have more of them, or have trade restrictions and things like that. I mean, scale is important. We still need to push the frontier. 00:28:47:25 - 00:29:22:03 Calvin Absolutely.

But we even explore this with a few people on our podcast, which you can find neural metric inference, time tactics that the GP use that we have in existing data systems centers are underutilized. So we're not even close to maximum efficiency on the capacity we already have built. So we have to have a multi-pronged approach. We need to both scale our capacity and build more data centers, and we need to make the ones that we already have more efficient.

00:29:22:03 - 00:29:31:13 Calvin And we need a more robust reseller market, and refurbishing and reusing. 00:29:31:15 - 00:29:59:22 Calvin Compute memory in a way that's useful. And when you think about alternatives and small language models and where the market's going to go next will be edge and compute some compute pushed to the edge, it's going to be really interesting. So we're very early in the innings and you don't see much talk at all in the media or in the narrative about AI efficiency.

00:29:59:25 - 00:30:04:22 Calvin But we're just starting to have that conversation in 2026. 00:30:04:25 - 00:30:32:16 Michael Yeah, I mean, I'm starting to hear some startups talk about it. I'm starting to hear some people kind of bandy around these different ideas around it. And some of that perhaps is, you know, the reaction to, oh, we don't have enough capacity.

But but I, you know, it seems to me that it that it does lead us down a path where the innovation needs to, to, to kick in that says we can do this with less resources, we have a more efficient way to deal with it. 00:30:32:22 - 00:30:57:16 Michael And the the size and the model to the problem seems like a really logical one in my mind that that does help us. You know, with resourcing, it also kind of opens up tech in a way to. Right.

Because you mentioned that some older chips, some older approaches work better in, in that, you know, kind of monitored world where we we do we are focused around efficiency. 00:30:57:19 - 00:31:14:26 Calvin Yeah. And and you know better is optimization problem like so it's not just always accuracy on everything or costs. It may be a latency issue you're solving for or privacy issue or compliance issue you're solving for.

Right. 00:31:14:28 - 00:31:40:26 Michael Yeah. I mean that's a that's an interesting point to that. Efficiency isn't always, you know, that specific to the fact that technology simply won't operate that way.

It's it could simply be the way you operate. Right? I mean, that's that's a choice. And, and the way you've approached this, by only using frontier models for problems that could be solved by smaller models, that, in effect, is you could contribute into the problem not solving it.

00:31:40:26 - 00:31:41:04 Michael Right? 00:31:41:06 - 00:32:11:28 Calvin Yeah. You may not want to give a corporation or an open source tool that can run rampant that's connected to that. Maybe you don't want to give this system level access to all of your data, right?

Maybe there's some things you want processed on your phone or a device, and then only certain things escalated to solve a problem. 00:32:11:28 - 00:32:52:22 Calvin Or or maybe there are compliance issues related to what objectives you need to achieve or latency issues. Right? Or maybe connect, like maybe you have AI enabled drones or something in the field, and you need to process intelligence on device when communications are lost.

Right? So there are all various reasons that you're going to need this kind of orchestration and that intelligence is going to be in the cloud, on the edge, on prem, and that you have to synthesize all of that for. 00:32:52:25 - 00:33:09:13 Michael I mean, there are a lot of different ways to solve a problem, I think is a really good point. Right.

And just because just because one approach has been popular doesn't mean it's the right approach. So taking a step back sometimes really does give you a different perspective on the whole problem. Yeah, yeah. 00:33:09:13 - 00:33:22:08 Calvin There's no one size fits all.

There's no one God model. You know, the future is less like Mission Impossible if you're a fan and more like Tron. That's just how. 00:33:22:10 - 00:33:42:13 Michael Yeah, that's fair.

I, you know, I mean, even for, for for analysts in my world, I've had this conversation a lot like, like they're like, I run a local model. Why do I run a local model? Well, I run a local model so I can air gap it so I can have it do things that I don't want to validate, you know, to violate an NDA over. 00:33:42:13 - 00:34:01:25 Michael But I still want to be able to summarize this presentation or, you know, track these different progress against these different types of, of, you know, activities.

And it's just things that, that I needed to protect. I thought I'll run a local model to do that. But in fact, what I've seen is that's a lot more efficient for certain types of tasks. 00:34:01:25 - 00:34:15:14 Michael And even even when it's connected, it's still, you know, it provides value in what I do.

Bye bye. Coming to that realization. So I think that's it's a learning that hopefully, you know, businesses are starting to experience. 00:34:15:16 - 00:34:17:08 Calvin Exactly.

Yeah. 00:34:17:10 - 00:34:46:19 Michael Yeah. I mean you've been on both sides of the VC table, you know, founder raising capital, venture capital, raising funds and, you know, working with different startups. How does that dual perspective really shape the way you think about the business model for near metric?

And maybe even more broadly, how do you see the economics of AI infrastructure evolving as the market moves from scale everything to more of an efficiency approach? 00:34:46:21 - 00:35:13:04 Calvin I think some of that came out and how I answer these questions, right. Reasoning through through things in a macro way. So often many VCs, many of the best ones that I admire, are thesis driven, right?

And asking these questions about the JSON possible. And then, you know, you roll up your sleeve and you get your hands dirty building a thing and trying. 00:35:13:08 - 00:35:40:04 Calvin And so how did that land to narrow metric? It was when I met Rob.

We were quick friends going to Knicks games and both serial entrepreneurs. And he's started his career ASIC chip designer and then built for companies and had some exits one nine figure exit. And he's been investing in AI companies, is an angel and then a VC for years. 00:35:40:04 - 00:36:03:06 Calvin And so we were talking about this thesis and he had this name, but we were talking about a bunch of different opportunities.

And this was more organic, that when he was introducing me to inference time algorithms, that I just became obsessed in thinking through this. And we were going back and forth and he's like, dude, you're like, doing the job. 00:36:03:07 - 00:36:21:28 Calvin Aren't you joined as a co-founder and CEO so often you can come to really interesting places if you ask interesting questions and follow your curiosity and don't stop there. Just like roll up your sleeve and try to do something and create.

00:36:22:01 - 00:36:48:19 Michael Yeah, I mean curiosity, that's a, that's a, that's a sort of a superpower to some people I think too. Right. I mean, you know, why did you do what you did? It's just because I'm inherently curious about everything.

And that's that's not a bad thing. I think, you know, if I'm a business leader listening right now and I'm running most of my AI on a single frontier model, which is, I'd say the most common approach right now anyway. 00:36:48:19 - 00:37:08:01 Michael And I'm starting to feel cost pressure, which again, I think is very common right now. What, you know, what should they do?

What's the first step they should take? How should they? You know what metrics are important here? What should they be measuring, evaluating so that they could make this decision that oh, multi-model makes a lot more sense for me.

00:37:08:03 - 00:37:16:25 Michael Or do you think it's universal? I mean, is it in general, should all companies be looking at size in the model to the task all the time? 00:37:16:27 - 00:37:38:21 Calvin Yeah. I mean if you are going from spending, you know, thousand to $10,000 a month and 100, I've got one simple piece of advice.

You should email Cooper at Neural Network AI, and I would love to help you cut your inference bill by 80 or 90%. You're on. 00:37:38:24 - 00:37:41:01 Michael Your inbox is going to blow up, I think. 00:37:41:07 - 00:38:07:16 Calvin And all jokes aside, if you're if you're prosumer vibe developer, you know, you can just try, one of our cells from the marketplace.

Go to our website. If you're if you've got an open for instance, it's just two lines of code. You just log in, it will give you an API key. And and you can you can do that.

00:38:07:18 - 00:38:40:07 Calvin We're not trying to replace your, you know, if you're using clod or OpenAI's API, you just would set fallbacks for advanced reasoning and, and and that's it. And so just start to send some of your tasks to smaller language models, and we can help them for you. And if you don't want to go through the process of figuring out how to distill a small English model yourself for the task and we can do custom builds is what we can. 00:38:40:08 - 00:38:46:04 Calvin We can create a custom small language model pack for you that works.

00:38:46:07 - 00:39:05:18 Michael Yeah, I mean, that makes sense. I think the, you know, we're seeing proof of concept all over the place. But, you know, as a part of that, evaluate size of the model to the problem. I feel like that should just be a basic part of all the, the proof of concept that you're, you're working through.

Right, exactly. 00:39:05:20 - 00:39:06:06 Calvin Yeah. 00:39:06:13 - 00:39:28:06 Michael Well, so that's really that's all the time we have today. But really interesting conversation.

I, you know, I feel like this is an area that's going to get a lot more attention this year. It's more and more companies start to realize that, you know, sizing things and thinking things through in a more optimized kind of framework is, is really important to to being successful. 00:39:28:06 - 00:39:42:27 Michael But before we let you go, one thing I always like to ask at the end of the episode is, you know, can you recommend somebody thought leader, an author, podcaster, somebody that you think the audience would really learn from and enjoy if they followed them?

00:39:43:00 - 00:40:14:19 Calvin Yeah, absolutely. I've already shamelessly plugged neural metrics. Substack. Yeah.

Rob May I'd say, yeah. Rob May, our CEO, but I'd say I've seen takes. I've seen f star who I mentioned earlier, the founder of Pilot Wave, working with him on some M&A. Right.

Bringing this technology into the real economy. And I've seems really interesting because he started his career. 00:40:14:21 - 00:40:46:24 Calvin He got an MD, PhD at Stanford, a medical doctor, and has a PhD in machine learning and AI. And then he was empty.

Goldman, chief data officer, science officer at JP Morgan, reporting to Jamie Dimon and then chief AI officer at Cerberus, the huge private equity firm. So he was the first chief AI officer at Wall Street History. 00:40:46:24 - 00:41:14:00 Calvin And then now he's doing it on his own. And he's the best kept secret.

And he's starting to do more content. So I'd say follow him too. So that would be at SNH f Onnx. Rob is at, Rob May at and I'm at Cooper underscore NYC underscore.

So I'd say like you're going to get different angles from those. 00:41:14:01 - 00:41:33:20 Calvin The macro lens the technical lens builder lens and AI. And you're going to get the AI in the real economy driving value and lower middle market companies. So that's a good set of things.

00:41:33:22 - 00:41:47:08 Michael Very good recommendations I appreciate that. And I know the the audience will check it out for sure. Well, Kelvin, thanks so much for joining today. Really interesting conversation.

And I feel like the audience certainly got a lot of value out of that. So thank you. 00:41:47:10 - 00:41:52:02 Calvin It thanks for having me on. 00:41:52:04 - 00:42:18:03 Michael And that's the show for this week.

Thank you all for joining us. Remember to like, share and subscribe to the show. If you enjoy the show, please leave us a review to help others find us. For more research on AI and other software, check out arionresearch.

com. If you're an expert in AI, generative AI, or business automation, either as a provider or an end user, email your information to disambiguation at arionresearch.com and don't forget to join us next week! Disambiguation is an Arion Research production.

I'm Michael Fauscette and this is the disambiguation podcast.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Spot That Vish!Simplifying Cyber · on Claude (Anthropic)90 / 100
  • A Conversation about the Human - AI Teaming Landscape: Designing the Hybrid WorkforceListen & Lead: Team Articles in Your Ears · on Foundation models86 / 100
  • The Benchmark With No Instructions - ARC-AGI-3 (winning team!)Machine Learning Street Talk · on Claude (Anthropic)85 / 100
  • 304: Boom, bust, or bubble? Rock Health weighs in on digital health funding in 2026Radio Advisory · on Claude (Anthropic)84 / 100
  • AI for Engineering Is Leaving the Demo PhaseAI Across The Product Lifecycle Podcast · on Claude (Anthropic)83 / 100
  • Building Action1: Mike Walters on Patch Management, AI Vibe Coding, and the Power of FocusCult Products · on Claude (Anthropic)81 / 100

More from Disambiguation

All episodes →
  • When AI Does the Building: Innovation, Ideation, and the New Creative Advantage65 / 100
  • AI Meets the Mid-Market: How PE-Backed Companies Are Leapfrogging with AI89 / 100
  • Beyond Efficiency: Why AI Is Forcing Marketing to Rethink Everything, Not Just Cut Costs87 / 100
  • The Cognitive Revolution in Leadership: Why AI Demands a New Human Operating Model73 / 100
  • The Flight to Relationships: Why AI Is Making Trust the Ultimate Sales Advantage84 / 100
Explore the best B2B AI & Data podcasts →
All Disambiguation episodes →