The Enterprise AI Show · 2026-08-05 · 15 min
Key moments - from our scoring
Substance score
43 / 100
Five dimensions, 20 points each
The episode challenges the industry's focus on ever-larger frontier models by questioning whether enterprises truly need trillion-parameter models for most use cases. Aaron and Brian draw parallels to database and application evolution, noting that value typically moves up the stack from infrastructure to user experience and workflow integration. They argue that while benchmarking drives model development, most enterprise applications - customer outreach, employee productivity, routine automation - don't require state-of-the-art performance and can't justify the cost premium. The hosts contrast this with use cases like recommendation engines and ad serving, where staying frontier-class creates genuine competitive advantage. They predict the future belongs to intelligent model routing: a unified interface that automatically directs tasks to appropriately-sized models (small, mixture-of-experts, or frontier), potentially combining owned smaller models with rented frontier capacity. The discussion covers benchmarking games (SOTA status lasting days, models distilled from bigger models), GPU economics, and enterprise ROI constraints. Takeaway: enterprises should strategically reserve frontier models for core differentiators while using smaller, cheaper alternatives everywhere else - enabled by smart routing layers similar to how OpenClaw routes between local and remote models.
No; most enterprise applications like employee productivity, customer outreach, and routine automation don't need frontier-class models and can't justify the cost premium. Only use cases tied to core competitive differentiation (like recommendation engines or ad serving) warrant frontier models.
The industry faces a 'race to the top' on benchmarks while simultaneously a 'race to the bottom' on business value, with SOTA status lasting only days and many models trained specifically to optimize narrow benchmarks rather than solve genuine business problems.
Intelligent model routing automatically directs tasks to appropriately-sized models (small, mixture-of-experts, or frontier) based on requirements, functioning like a unified interface that hides model selection from users - similar to how OpenClaw routes between local and remote models.
The future likely involves owning smaller, cheaper models for routine tasks while renting frontier models for high-value use cases, mediated by intelligent routing that optimizes both cost and performance.
After GPT-4, subsequent model releases haven't delivered obvious jumps in capability; industry observers can detect differences but most enterprise tasks - question-answering, agent workflows - see diminishing returns from newer, larger models.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode covers genuinely useful territory - questioning whether trillion-parameter models are necessary, discussing model routing, and drawing enterprise analogies - but much of the discussion remains abstract and speculative rather than grounded in concrete findings. The speakers repeat their core thesis multiple times without introducing novel data or surprising claims that a B2B operator would learn for the first time.
do you even need that trillion parameter model
the models is a race to the top, meaning benchmarks, but that race to the top is leading to a race to the bottom all at the same time
The core argument - that not every use case needs frontier models and that companies should optimize for ROI rather than chasing benchmarks - is sensible but circulating widely in AI discourse. The analogy to airlines prioritizing fuel costs and seating optimization is accessible but not particularly fresh. The discussion of model routing is somewhat more forward-looking but remains underdeveloped.
is that where we now start going
there will be a small set of stuff that the frontier level things probably make the most sense for, because you look at your competition
This is a peer discussion between two hosts rather than an external guest interview, which immediately limits caliber assessment. Neither host demonstrates deep operational evidence of having built or deployed AI systems at scale; they speak theoretically about enterprise needs and model economics without citing personal implementation experience or results.
I live in the model world
you know, coming from these enterprise backgrounds
The episode lacks concrete data, named examples, or specific metrics to ground its claims. References are vague ("a couple years into this," "two days" for SOTA models, no actual pricing). The Oracle and database analogy is mentioned but never developed. No specific enterprise case studies, benchmarks, or dollar figures are provided to support the thesis.
if you're soda, you're soda for like two days
there's lots of accusations of just, hey, we're just distilling off the big frontier models anyway
The conversation flows naturally with some back-and-forth, but lacks incisive follow-ups or genuine friction. When claims are made (e.g., about model distillation accusations, SOTA training to benchmarks), the co-host largely affirms rather than probes deeper. Questions tend to be rhetorical or confirmatory rather than challenging. The hosts agree frequently without testing each other's assumptions.
Yeah, yeah. And for me, like, okay, if I kind of look at the present state
Well, and the other part of it, it's going to come back to
Computed from the transcript - who did the talking, and the words that came up most.
SUMMARY: We continue our Models and Money series. In this episode, Brian and Aaron explore the current state and future of AI models, focusing on model size, model harnessing, and intelligent model routing. They discuss whether bigger models are always better, the economics of AI, and how enterprise applications can benefit from tailored AI solutions. SHOW: 1051 SHOW TRANSCRIPT: The Enterprise AI Show #1051 Transcript SHOW VIDEO: SHOW SPONSORS: Nasuni - Activate your data for AI and request a demo Topic: When will the models/harnesses be good enough? Why now? Benchmark Maxxing - cost to build/host/maintain 1+ trillion parameter model Past: The “wow” moments in versions really stopped around GPT4… (maybe?) Present: Race to the top/bottom, millions spent to gain SOTA for a few days Future: Will the pendulum swing back? Will bigger/faster always rule? FEEDBACK? Email: show @ the enterprise ai show dot come Bluesky: @TheEntAIShow.bsky.social Twitter/X: @TheEntAIShow Instagram: @TheEntAIShow
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: Uh, good evening ovr and welcome back to the Enterprise AI uh show. This is your host, Aaron. And today we have another episode in our Models and Market series. Brian and I asked the question, do you even need that trillion parameter model? We talk about the current and future state of frontier class models. We focus on model size, model harnessing and intelligent model routing. And we discuss whether bigger models are always better, the economics of AI, and how enterprise applications can benefit from tailored AI solutions. And all of that is coming up right after this break.
Speaker A: Today's show is sponsored by Nasuni. There's a growing gap in AI right now between what's possible in theory and what successfully works at scale inside an enterprise. The difference comes down to unstructured file data. Many AI initiatives struggle because the file data they depend on is scattered, unstructured and disconnected from where and how work actually happens. Nasuni changes that. It brings your unstructured file data into a single secure foundation so AI, both generative and agentic, can access it with the context, governance and performance it needs in production. Bring AI to where your unstructured data lives. See what it takes to activate your data for AI and request a demo@nasuni.com AI Aaron, you and I back again doing another one of these short shows that we have not yet come up with a name for. But the first couple we did, the last couple we did, were kind of ownership and cost and control focused. What do you have on your mind today? Are we going to get outside of that harness, if you will, of cost and control?
Speaker B: Yeah, uh, yeah. So the idea was there is like, okay, we kind of want to mix in some business shows and some future in India industry kinds of stuff, some macro level shows, if you will. But I also want to dig into some more technology focused and a little bit of like, okay, state of tech in the industry and then where it's going and we're going to do our usual like, hey, why are we talking about this right now? What's been in the past, what's the present state look like and where do we think this is going in the future? And so this topic that came uh, up with was like, when will the models and harnesses be good enough and when? Why am I saying that right now and why am I asking that question? Well, there's this concept out there of benchmaxing. So basically, you know, there's lots of big models out there and oh, by the way, there's lots of models that like they're, you kind of like building this and if you Build a trillion parameter model. I mean that thing costs a lot of money to build, it costs a lot of money to host, it costs a lot of money to maintain. And oh, by the way, do you even need that? Like we were talking on some of the other ones, like, hey, if I compare that to the Ferrari or you know, we compare it to like, okay, what kind of queries are we doing against? It's not. A lot of them aren't necessarily doing deep research. And so like the past piece I wanted to talk about was like, when was the last time we had a really big wow moment in versions? Um, like, uh, for me personally, it was like a GPT4 level maybe. Right. Like I noticed a jump from 3 to 4 but you know, from all of them since it wasn't a big like, yeah, like you kind of said, oh, you know, some of the folks in the industry can tell the difference between the models and it might have a different vibe or it might have a different this or my different, you know what if all you want to do is, is have it, you know, question, answer kind of thing or agent go off and do a task and tell me when you're done, hey, guess what? You don't need a trillion parameter model. Right? Yeah. Uh, and so like, are we reaching a state where everything's good enough and now we're just feeding in the industry of up and to the right?
Speaker A: Yeah, I mean, I think that's a discussion that I feel like it's been going on for now at least a year right now. The economics are still incredibly driven by the model being, the model being big, the model being frontier on however many levels. And again, this is, you know, I'm trying to think of the equivalent of this because it would be going back quite a long ways. But I mean, I'm sure back in the day when Oracle's first database or the first time series database or whatever it was came along, it was pretty revolutionary and it was like, oh, wow, we're accessing data really fast. We're able to do all sorts of uh, merges and queries and other stuff that was data centric. And then at some point the industry sort of said like, you know, the thing that really makes, you know, accessing this data or the information powerful is the user experience, the way it's blended into other workflows. Right. The value moved up the application chain. And you know, I think people are constantly questioning like, okay, is that where we now start going? Is that where we now start going? And I feel like a Little bit. We skipped it a little bit because we kind of went from chatbots were the first sort of application. Then people couldn't really figure out a whole lot of things other than chatbots. And then they've sort of moved on to agents and agentic and we really haven't built too, too many kind of classic applications on top. And maybe that's because that's not the way we're necessarily going to do these things. But it does feel like you're right, that what you get out of the model is as good as it is. And I think people still expect it to sort of. It's going to evolve and it's going to have more recent data. But. But it does feel like all the attention these days has been focused on the harness, on narrowing the experience, on perfecting the experience of trying to reduce hallucinations or make it seem more deterministic. It feels like that's where people feel like they can differentiate themselves, especially if they don't own the model, for example. I think they're trying to. And then I think they're trying to figure out a way to make that harness portable so that, you know, if they want to move off of one model to another. I don't know how realistic that is yet. Uh, you probably know that a little bit better because you live in the model world. But I feel like that's another top of mind thing people are trying to figure out.
Speaker B: Yeah, yeah. And for me, like, okay, if I kind of look at the present state of all of this and this kind of goes to where I was m. The core message or the core idea behind the question was I almost feel like, okay, the models is a race to the top, meaning benchmarks, but that race to the top is leading to a race to the bottom all at the same time. And because at the end of the day, I mean, we have no way to prove it. But I mean a lot of the models, especially the OSS models, the Chinese OSS models in particular, there's lots of accusations of just, hey, we're just distilling off the big frontier models anyway. And oh, by the way, it probably costs millions of dollars to generate these things. And oh, by the way, uh, if you're state of the art or soda, ah, as it's referred to in the industry, a lot of times if you're soda, you're soda for like two days.
Speaker A: Yeah.
Speaker B: And oh, by the way, you probably trained it to specifically be soda against that benchmark. And so like, I don't know, like, I Feel like maybe I'm getting already, what, a couple years into this. I'm already old man yells at AI or old man yells at cloud here of like, you know, I kind of have to step back every once in a while and go, but does this solve enterprise problems? And you know, with us, like, uh, kind of us, but coming from these enterprise backgrounds, it's like, what's the business outcome? You know, scoring just, you know, a couple points more on a benchmark really
Speaker A: do well, you know, and I think if we're, if we're trying to, with these, these little short shows, if we're trying to draw some analogies to things, I mean, if you think about like consumer types of applications, right? The, the ones that are super, super finely tuned and always sort of have to be latest to greatest are things like ad serving, right? Like, you know, recommendation engines and ad serving that are very, very tied to like your most core business activity. But you know, for a lot of the other stuff that you do, you don't need latest and greatest of stuff. And if you think about most like enterprises, there's probably only a few things that they do. You know, maybe it's, you know, if you're an airline, you're putting your best and brightest on like fuel cost and like seating optimization so you can, you know, get people to pay you more money for a better seat or a better. But like, you're not putting that on the other 89, 95% of the applications. And I think we're going to see probably a similar thing with AI is there will be, uh, a small set of stuff that the frontier level things probably make the most sense for, because you look at your competition or you look at the thing that's creating your moat for you and you're like, yep, I see the value of staying on frontier or state of the art, because it's something I can continue to create great margin with. But for everything else, if it's something that's just doing employee productivity or it's doing, you know, customer outreach, how much differentiation can you really do there? And even with that, how much are you willing to pay to differentiate that next thing, right? You know, if you can't guarantee that like that you can, you can get somebody to click back and buy something with one email as opposed to six emails, you know, reminding them to do stuff you don't need AI to do, you know, AB testing on those six emails, you're just like, no, I just know the system takes like six repetitive tasks and then they finally buy Something from me. So.
Speaker B: Well, it reminds me of. There was. Oh gosh, I think it was, uh. By the way, everyone go subscribe to the podcast this week in AI. It's a really good news podcast, but they also do some good breakdowns of things. But it was really interesting. They had this theory once upon a time and I don't know if it was their theory or somebody else's, but it was basically like, hey, the only reason we still need to go up into the right is for like these Frontier labs should honestly just take their models and go do like day trading with them. Because like that way you can at least make some money back on them and it goes back to your core differentiator. And like, what's the, what's the idea of where AI is going to move the needle? The most kind of thing? And it just, it kind of made me like laugh but at the same time go, hmm, Hm. Yeah, like, where do you really, really need all of these things? And uh, you know, at the end of the day, the enterprise a lot of times is like, okay, I need to solve a problem, right? But I need to solve a problem in a way that's economically viable for the company, gets the best ROI and is the fastest, right? It's the. Your typical project management constraints, right? Quality, time and cost. And what do you want? And you can only have two of the three kind of thing. And uh, is AI, uh, helping that or hurting that in something like this?
Speaker A: Well, and the other part of it, it's going to come back to, you know, if, you know, going back to one of the shows we did a couple of weeks ago, you know, if you feel like you need a little more control over your environment and that ends up being that you're either, you know, buying or renting GPUs, you don't necessarily want to start with a trillion parameter model on, you know, what might be limited GPUs or limited size GPUs. You want something that's gonna fit on the GPUs, be able to do a query within that one GPU, not have to ship it all over the place, do all sorts of load balancing and stuff like that. So there's you, there's going to be both quality of the work decisions, but also economic decisions that have some ripple effects as to what you choose.
Speaker B: Ah, by the way too, I want to go into the future state for a little bit here. Here's how I see the future state of all of this working out is, yes, like we just said, hey, Guess what? There is use cases for these really, really big models. But oh, by the way, there's lots of use cases for small models. There's lots of use cases for a mixture of experts and dents and all the other different technologies that are. But I uh, think this is where the model world really goes into. This is where agentic and agentic workflows and more specifically models and model routing comes into play. There's going to be one central interface I really firmly believe we're going to get to a point in the industry. It's like, do you remember when you had to think about the model you used? You just have a task and you put it in the thing and the thing decides where to route it and how to route it. And maybe it's kind of a culmination of our last couple shows of like, oh, by the way, that wraps up the finops and uh, it wraps up their own versus renting your AI of like, hey, guess what, you might own the small models but you might rent the big models. Like I feel like there's this future we haven't even begun to scratch the surface on which is intelligent model routing. So that's my idea of where all this gives.
Speaker A: Well, and that's going to be an area that'll be interesting because that's the kind of thing where what would normally happen in the marketplace is yes, there would be some sort of semantic routing capability, but it would be owned by the model provider itself. Right. It would be owned by an AWS or an Azure and they would be routing you between, you know, their five or six models at some tiered pricing. And then maybe there's this last resort of oh, okay, somebody said they need to be able to try something else. Or you know, the open and open source community would come along and build the equivalent of like Engine X Right. So that you could, you could do load balancing that you could have better control over that, you know, was a lower cost price point or something like that. Because the reality is if you offer customers something that you know, kind of mediates between what your profit looks like, you're going to want to try and get some extra profit out of that, that mediated thing. Right. Because it's, you know, it's like the equivalent of somebody being like, well, you could pick your own stocks for your portfolio, but if you have an expert do it for you, then for just a small cut. Yeah, you should pay them a 1% fee for basically doing so. Yeah, I mean, uh, that'll be an interesting thing to See that evolve. I think you're right. There's going to be a demand for that. The question becomes how does that show itself in the marketplace? Yeah.
Speaker B: Uh, and by the way, for everyone out there that actually uses OpenCall on a regular basis and knows the technology inside it out, I may completely mess this up because I've been researching some openclaw stuff here recently for actually for some show related stuff. But the idea is, yeah, okay, uh, openclaw runs on a harness, it runs on a workstation. But oh, by the way, when it needs to do stuff, it will route out and call bigger models for bigger things based off of certain decisions and connectors. And I almost see that kind of happening of like, you know, AI's calling AI's kind of thing and just trying to break it down into something that's easy and consumable and folks have access to today.
Speaker A: So yeah, I think we're, we're, we're still a long way from like in a perfect world, this is what it would look like for me. And you know, this is, this is the reality of it. But we can, we can dream, we can, we uh, can hope.
Speaker B: Agreed.
Speaker A: So. All right. M man, you want to wrap this one up there and uh, yeah, absolutely
Speaker B: everyone out there, thank you so much for listening and as always, reach out to us at any time if you enjoy these shows. We'd love to hear your ideas of what you want to hear as well. So definitely send us an email, reach out on the socials. We'd love to hear from you. Thank you.
Speaker A: Great. All right, talk to you next week folks. Thanks for listening. Check us out@theenterpriseaishow um.com for past shows, newsletters and all things enterprise.
Speaker B: AI.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.