
Beyond The Pilot: Enterprise AI in Action · 2026-07-08 · 15 min
Key moments - from our scoring
Substance score
48 / 100
Five dimensions, 20 points each
The enterprise AI landscape is fragmenting as companies grapple with competing priorities: reliability, cost efficiency, and access to cutting-edge capabilities. While OpenAI's guardrails dominate 51% of enterprises and hyperscaler bundles remain the default choice, recent disruptions - particularly Fable's 19-day outage - have prompted teams to reconsider dependency on single-provider models. Chinese models like GLM 5.2 and Tencent's newly released Hi3 are capturing developer mindshare by delivering frontier-level performance on specialized tasks (agents, RL environments, coding) at lower computational costs. Meanwhile, only 4.5% of enterprises fully trust automated evaluation tools, forcing manual validation of production deployments. The episode explores how organizations from Shopify to Intuit are developing hybrid routing strategies: leveraging frontier models (OpenAI, Anthropic) for high-stakes reasoning while deploying efficient open or alternative models for standard workloads. Infrastructure adoption remains early - Neo4j at 2% despite 45% planning investment - reflecting the market's immaturity. Guest insights reveal how companies use direct user feedback (agents rejecting routing decisions) to collect higher-quality training data than traditional preference collection, fundamentally reshaping how teams architect agent systems and make real-time model selection decisions.
Chinese models achieve frontier-level performance through aggressive reinforcement learning environment generation, competing directly with Anthropic's Opus while running at half the size and cost, making them attractive as production fallbacks after the Fable outage exposed single-provider risk.
Only 4.5% of enterprises fully trust automated evaluation platforms; most rely on manual validation and hyperscaler-provided tools despite recognizing their limitations, creating a significant gap between shipped agents and real-world customer performance.
Leading companies like Shopify and Intuit use hybrid approaches: frontier models (OpenAI, Anthropic Claude) for development and high-complexity reasoning, and cost-efficient alternatives for standard production workloads, with real-time routing based on cost and reliability requirements.
The outage prompted enterprises to reconsider dependency on single-provider models and invest in fallback options and open-source models, accelerating interest in alternatives like Chinese models and infrastructure tools despite previous inertia toward hyperscaler bundles.
Neo4j is used by only 2% of enterprise respondents but 45% plan infrastructure investment in the next 12 months, reflecting the early-stage market maturity and openness to disruptive tooling as enterprises move beyond basic deployments.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains some genuinely useful data points (51% using OpenAI guardrails, 66% moving toward zero human deployment, only 4.5% trusting automated evals, 2% using Neo clouds but 45% planning to invest) alongside substantive discussion of routing decisions and model selection strategies. However, significant portions are devoted to conference promotion, speculation about future guest panels, and repeated emphasis on the Fable outage without deep analysis. The conversation lacks granular case studies or detailed operational insights beyond high-level assertions about hybrid approaches.
Half of our respondents are reporting that they've actually shipped, uh, an agent that has passed their evals but have failed a customer in some way.
66% already allow or are engineering towards zero human deployment
The framing around model routing and the tension between frontier vs. efficient models is topical but not particularly novel. The discussion of Chinese models (GLM 5.2, Tencent) offers some fresh context, but much of the analysis relies on conventional wisdom about hyperscaler lock-in, cost pressures, and reliability concerns. The 'get it right, get it fast, get it cheap' framing is repackaged but not original. Limited contrarian or first-principles thinking.
people don't get fired for picking Google or Microsoft. The safe bet is to go with one of those big companies.
the difference, uh, you know, with this is you're not having to guess what the customer is doing or guess what the customer is thinking because they're actually telling you by telling the agent
This episode features no direct guests - only two hosts (Speaker A and Speaker B, identified as Matt Marshall and Sam) discussing enterprise AI trends and previewing upcoming conference panels. While they reference conversations with Farhan Tawar (Shopify engineering head) and Intuit engineers, these are indirect secondhand accounts, not direct testimony. The episode is essentially a research presentation and event promotion, not a substantive interview with a practitioner who has built and shipped at scale.
When we were talking with Farhan together, he was saying that just as much as they're using open source
I was talking with one of the people that's coming from Intuit
The episode cites concrete research data (51%, 66%, 4.5%, 2% adoption, 45% planned investment) and names specific models (GLM 5.2, Tencent model, Opus 4.8) and companies (Shopify, Intuit, LinkedIn, Pinterest, Amazon, Databricks, Brex). However, the evidence is largely presented as survey findings without detailed breakdowns, timelines, or dollar figures. Examples from Shopify and Intuit are mentioned but not explored with specificity. The Fable outage is referenced repeatedly without concrete impact data.
51% of enterprise companies. Half of our respondents are reporting that they've actually shipped, uh, an agent that has passed their evals but have failed a customer in some way.
Neo clouds, right? As popular as they are, they're only being used by like 2% of our respondents. But you know what, 45% of our respondents are saying that over the next 12 months, that's going to be their number one area of planned investments.
The hosts engage in some back-and-forth, with Speaker A pushing back on claims about model diversification and hyperscaler dependency (e.g., 'I would agree with that. Um, but I'm also seeing that the market is just wide open'). However, the conversation often meanders into conference promotion and future panel previews rather than drilling deeper into disagreements or testing assertions. Follow-ups tend to be soft, restating points rather than genuinely probing. The discussion lacks the sharpness needed to challenge vague claims or extract operational detail.
Right, so you're saying there's this, there's this diversity and hybrid approach. I agree with that. Um, but I'm also seeing that the market is just wide open.
I'm going to come back to the seam and keep pushing back. Is that for every point you're making about the need for efficiency and lowering costs, there's this, there's this inertia where a lot of companies maybe for simplicity reasons, for reliability reasons, they're relying on these hyperscalers.
Computed from the transcript - who did the talking, and the words that came up most.
66% of enterprise AI teams are actively engineering toward zero human oversight - but only 4.5% fully trust their automated evals. Half have already shipped an agent that passed evals and still failed a real customer. VentureBeat's latest research exposes a dangerous gap between deployment ambition and evaluation maturity. The most-used eval platform is OpenAI's native tooling. The second most common answer: no dedicated eval platform at all. Meanwhile, 51% of enterprises are relying on a hyperscaler's built-in guardrails as their primary AI security layer - not a deliberate security architecture. Fable's 19-day market absence changed the conversation. Enterprises that had built production workflows on frontier models got a wake-up call about single-vendor dependency. The result: GLM 5.2 is gaining serious developer traction as an open alternative competitive with Opus 4.8, and Tencent's new Hai 3 - less than half the size - is drawing attention for agentic workloads that don't require cutting-edge coding performance. Sam and Matt break down why companies like Shopify, LinkedIn, and Pinterest were already on open models, and why others are now catching up fast.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Enterprises are naming a hyperscaler solution as their primary security layer. Right. So OpenAI's guardrails are used by 51% of enterprise companies. Half of our respondents are reporting that they've actually shipped, uh, an agent that has passed their evals but have failed a customer in some way. The most used eval platform is OpenAI's native evals. And you know what the next answer is in terms of popularity among unrespondents? It's no dedicated agent eval tool at all. No one trusts these evals. Right. So 66% already allow or are engineering towards zero human deployment, meaning that they're really trying to move to get the agents working. Without humans in the loop, only 4.5% fully trust automated evaluation. It's moving so quickly that enterprise companies are essentially relying on these hyperscalers still. So tell me, how are these Chinese models breaking in? Is this just a developer fad or is it real?
Speaker B: Well, you think about it, people don't get fired for picking Google or Microsoft. The safe bet is to go with one of those big companies. But we've seen now, if you're trying to have the cutting edge AI and you were trying to go with anthropic and hey, that perhaps things didn't work out as well as people thought. And don't forget this is on the back of tokens suddenly costing a lot of money as well. So over the past few, few months we've had a lot of guests on the POD talking about, you know, the whole sort of token maxing versus token budgeting, the token economy and that kind of thing. Look, the Chinese models, uh, over the, you know, the past month or so we've seen GLM 5.2, uh, release huge sort of step, uh, forward for the open ecosystem there, that finally there's a model out there that's pretty close to Opus 4.8. Just today as we're recording this, we've seen high three from Tencent drop a new model that's about less than half the size of GLM 5.2. So literally people could run this in their office. And I think this is one of the things I'm looking forward to about transform. Right, right. We're going to be talking to the people who actually are making these decisions. And we know from the pod, when we've been talking to your director of engineering at Shopify, that they are doing open models. We know from LinkedIn they're doing open models. We know from Pinterest they're doing open models. The raw engineering talent have woken up to this a lot earlier, but I think Fable going down has made a lot of other people wake up to this as well.
Speaker A: Yeah, I'd push back a little bit, Sam. When Fable goes down for me, I give. Sure. I was using it for a few days. I liked it, it went down. But I just went back to 4.8 what the next offering was. I got to use Anthropic, um, uh, continued to use a Frontier model that was reliable. It was in my workflow. And then when it came back, I was up and using it again. This is Venture Beats beyond, the pilot podcast. It's about enterprise AI in action. I'm Matt Marshall. Today's episode is presented by Outshift by Cisco, Cisco's emerging tech incubation engine and driver of agentic AI quantum next gen, infra and beyond. SAM, GLM 5.2 and then this recent model from Tencent that just dropped today are grabbing mindshare. Walk us through. Right? You gave us a lot of good statistics, but walk us through exactly what it is that they are doing that is really innovative. Can you talk through GLM5 for a second?
Speaker B: Look, the GLM5.2 clearly is a very competitive model with something like Opus 4. 8. It's become a, uh, very popular model on open router already. You know, lots of people are using this, and I think a lot of companies are now starting to check that out. The big thing that they're doing is they're actually getting the whole idea of RL environments and how they're actually training the models. They're very much on top of that. We've seen a number of the Chinese labs really go into this in depth, and it does seem that what they have lacked in compute, or what they've lacked in perhaps sort of frontier knowledge, for some of them, they've been able to make up for through just generating tens of thousands of these RL environments. Uh, GLM 5.2 is really interesting one because so many developers are starting to say that this is the open model that they would happily use for their coding. In fact, we've seen this more and more on X that lots of people who only relied on, uh, Claude, or only relied on OpenAI, the Tencent model. I don't think that's the model for coding. Uh, when you look at the sort of stats for it, it's still quite far behind when it comes to coding, but it's half the size and it can do pretty much everything else really well, including the agent and the agentic stuff. Again, you know, bringing it back to transform for a second. We know that like a lot of the guests that we've got are, uh, talking about really interesting things that they've been doing with agents. And as they're sort of putting this stuff into production, if they're going to suddenly realize, gee, the M model that we were building on is being taken away from us, they're going to want to, you know, have something in reserve that they can use on their own GPUs. And at the moment that seems to be some of these Chinese models. 19 days. 19 days changed the world in AI, right? This is how long Fable was basically off the market for. And it turned out that it wasn't just, uh, anthropic. You know, OpenAI has done the same thing. They're supposedly going to release 5.6 this week, but they've kind of announced it and said that it was too dangerous to release something. We hear quite often from OpenAI, but we suddenly saw a lot of people wake up and realize that, okay, maybe they can't trust these big companies as much as they thought. The world's changed, right? These 19 days have changed the world. And then next week we're going to get to hear from a lot of these people that were directly affected. I suspect that one of the things we're going to see is that you've got the sort of three things right. Get it right, uh, get it fast, then get it cheap. And that some of them are still trying to work out the get it right bit and they're tending to use the biggest, best models. Uh, other people that ah, are perhaps like Shopify, are now looking at the get it fast, get it cheap element of the scale. And I think it's fascinating to hear that because I know just sort of leading up to it, some of the pre calls that I've been doing with some of the people I'm going to be talking to on stage. It is really interesting how different organizations are approaching even just the building of agents. I, uh, won't reveal sort of too much, but it's sort of fascinating that some people are using, uh, the anthropic clawed managed agents for prototyping. But once they've sort of got, you know, get it to work, they're then basically, you know, putting it going fully bespoke parts of sort of LangChain parts of, you know, some things like that. But again, they're heavily customizing it when they go to production. Uh, I was talking with one of the people that's coming from Intuit, uh, and we're going to be talking about how they actually bootstrapped their agents, uh, and how they sort of pivoted on a dime to change it when it wasn't working. One of the things that they were making, the point, uh, which is really interesting, is that the difference, uh, you know, with this is you're not having to guess what the customer is doing or guess what the customer is thinking because they're actually telling you by telling the agent, no, I don't want that. I want this. Right. So she was talking about some really interesting stuff about, about, you know, how in some ways this is getting much higher quality data than trying to get people to go back and remember why they liked something about a product or why they didn't like, you know, something. And I think that there's a lot of lessons to be learned from what some of these guests and what the companies that they're working at are actually doing here.
Speaker A: Right, so you're saying there's this, there's this diversity and hybrid approach. I would agree with that. Um, but I'm also seeing that the market is just wide open. And I think it's partly because of this chaos, right, because of this dynamism in the industry. So for example, we'll take, you know, take a look at infrastructure, right? Neo clouds, right? As popular as they are, they're only being used by like 2% of our respondents. But you know what, 45% of our respondents are saying that over the next 12 months, that's going to be their number one area of planned investments. So I think that's just a great example of how while there is this inertia and so kind of freeze in action because things are changing or fable being pulled, et cetera, that the market is wide open, still on a lot of these areas. And you're seeing that in the switches, right? OpenAI doing really well. Anthropic coming along and knocking them off the pedestal as the leader and enterprise. So this hasn't calcified yet. That's what's interesting.
Speaker B: I'm looking forward to transform next week and I'm looking forward to talking to a lot of the guests that we're going to have on stage because at the moment everything's just sort of crazy. You know, we had this, this sort of 19 day blackout of fable. A lot of companies have started to realize now that if they, if they're not sure if the model's actually going to be there, how do you actually think about building products on something like that, uh, you know, it's been a crazy couple of weeks.
Speaker A: You talk about Fable being pulled as though that was a big problem. You know, companies are looking for options, um, and alternatives. Absolutely. There's this vacuum that was created. But at the same time, Sam, what we're finding, you know, in venture breach research, and we're doing quite a bit of research of companies that have 100 employees or more in survey after survey, in area after area. Right. Orchestration Agent. Security. Security. The market leader is basically the provider bundle. Right. So whether it's OpenAI, Microsoft, Google, Anthropic or AWS, it's one of these players. Enterprise companies by and large are relying on the offerings of these providers. Right. Or it's nothing at all. And I could go down the list. So you talked about Farhan Tawar, the head of engineering at ah, Shopify, using closed models. When we were talking with Farhan together, he was saying that just as much as they're using open source and you know, maybe less frontier models for a lot of their workloads that don't need that high intelligence, they're always going to need that frontier intelligence for a good part of their workflows. Right. And so the big question is how does that shake out? And I think there's a lot of interesting research that's happening right now around how do you route to the better agent, the smarter agent, for what, um, workloads and when? That's the big question mark. And that's what we're also going to be talking about. Transforms. How are these leaders, uh, making these decisions right now when there is so much concern around costs?
Speaker B: Yeah, look, definitely this is not an either or thing. Right. Nobody was deploying Fable to customers because of the 30 day retention and that basically went against every, you know, all their sort of agreements. But in a world where models are being taken away from, people don't just want to rely on those models. That's what I'm saying. Uh, with this they want to have you know, those cutting edge models for development, for the creation of new products and stuff. But for day to day production they want things that are reliable, that are cheap, that they, they know that they can have in lots of different data centers, they know that they've got different fallbacks, all those sorts of things. Matt, what are you most looking forward to at Transform?
Speaker A: So we have uh, Innovation Showcase where we're going to be hearing from some, a lot of uh, a lot of new products. Looking forward to revealing our latest research. Sam, on these big questions, I think what we're going to see is I'm going to come back to the seam and keep pushing back. Is that for every point you're making about the need for efficiency and lowering costs, there's this, there's this inertia where a lot of companies maybe for simplicity reasons, for reliability reasons, they're relying on these hyperscalers.
Speaker B: Yeah, look, it's going to be a fascinating conference. I'm certainly looking forward to it. Matt, do you want to tell people quickly just where it is, where they can get last minute tickets, et cetera?
Speaker A: Yeah, that's right. It's, uh, a hotel near, in Menlo park, so not far from SFO Airport if you're flying in or if you're local. But yes, it's going to be an exciting two days. It's our flagship event. Several hundred people who are all enterprise tech decision makers, conversations, Amazon, DataBricks, Brex, Instacart, MasterCard, all, you know, we got VP and C level folks from all these companies and, you know, a good couple of dozen more. The networking is going to be just as fun as the content itself. I'm really looking forward to it. Looking forward to seeing you there, Sam.
Speaker B: I'm looking forward to it a lot. So, yes, if you see us, uh, come up and say hi. This series is brought to you by Outshift, Cisco's incubation engine. By creating an open interoperable infrastructure, Outshift is enabling agents and humans to share intent, context and reasoning. The cognitive evolution for agents is here. Explore the Internet of cognition@outshift.com for more stories about the AI revolution like and subscribe to the podcast and check out venturebeat.com to sign up for our newsletters.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.