ChatGPT and Beyond with Fexingo · 2026-06-29 · 9 min
Key moments - from our scoring
Substance score
56 / 100
Five dimensions, 20 points each
The AI economy is experiencing a counterintuitive moment: while inference costs have crashed 40% year-over-year, hardware vendors like NVIDIA (down 7.5%), ARM (down 18%), and Super Micro Computer (down 13.5%) are getting hammered. Lucas and Luna argue this reflects a fundamental market repricing. The bottleneck has shifted from compute cost to operational complexity - companies can't efficiently deploy the models they already have. Enterprise software vendors like Salesforce (up 5.5%), ServiceNow (up 6%), Snowflake (up 10%), and Adobe are outperforming because they solve the non-model problems: data privacy, legacy system integration, compliance, and change management. A survey shows over 60% of companies that piloted AI in 2025 still haven't moved to production. Independent middleware players and AI observability tools (Datadog, New Relic) are becoming critical as companies need to monitor model performance and avoid overpaying for inference. The winners aren't the hardware 'picks and shovels' but the software instruction manual - the operational layer that makes AI actually work inside corporations.
The market is repricing expectations: cheaper models don't automatically drive higher-margin deployments because the bottleneck has shifted from compute cost to operational complexity. Enterprises struggle to integrate models into workflows, so they buy less hardware despite lower unit costs. ARM's 18% decline specifically reflects concerns that hyperscaler data center buildout is peaking.
Salesforce, ServiceNow, Snowflake, and Adobe are outperforming because they provide middleware and integration tools that help enterprises deploy AI in production, solve data privacy and legacy system integration challenges, and manage models in actual workflows.
Less than 40% - over 60% of companies that piloted AI in 2025 still haven't deployed anything to production, citing challenges with data privacy, legacy system integration, change management, and regulatory compliance.
One vendor claims they can cut API costs by 60% without quality drops by routing queries to the cheapest model that can still answer correctly, essentially creating price discovery between model providers.
AI observability tools (like features in Datadog and New Relic) monitor model performance, cost, and accuracy in production - they're now must-haves for companies spending millions on inference to detect model drift and avoid overpaying.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers several non-obvious observations - the disconnect between cheaper inference and hardware stock declines, the shift from GPU bottlenecks to operational complexity, and the emerging importance of middleware and observability layers. However, it spends considerable time on stock price movements and market rotation that, while connected to the thesis, dilutes the substance-per-minute ratio for a B2B operator.
the market is starting to realize that cheaper models don't automatically translate into more profitable deployments. The bottleneck has shifted from compute cost to operational complexity.
over sixty percent of companies that piloted AI in 2025 still haven't deployed anything into production
The core insight - that the value shift from infrastructure to middleware mirrors early cloud computing - is a valid pattern observation, but it's a familiar analytical move in tech cycles. The specific focus on observability tools and the 'model middleware' layer routing queries to cheapest providers is less worn, but the overall framing of 'picks and shovels' analogy and enterprise adoption walls are well-trodden in AI discourse.
the software layer that sits between the model and the business problem
It's almost like the early days of cloud computing. Everyone rushed to build data centers, then the value moved to the management and security layers.
This is a co-hosted conversation between two unnamed hosts (Lucas and Luna) with no identified guest. While they reference a conversation with 'one vendor' and cite a consulting survey, there is no credentialed operator or practitioner as a primary speaker. The hosts appear to be analysts or commentators rather than practitioners who have built or scaled AI infrastructure or enterprise AI deployments at significant scale.
One vendor I spoke with last week claims they can cut API costs by sixty percent
In a survey I saw last month from a consulting firm, over sixty percent of companies that piloted AI in 2025 still haven't deployed anything into production
The episode provides concrete data points - 40% year-over-year inference cost decline, specific stock movements (NVIDIA -7.5%, ARM -18%, Super Micro -13.5%), named companies (Salesforce, ServiceNow, Snowflake, Datadog, New Relic), and a claimed 60% API cost reduction from an unnamed vendor. However, several claims lack granular support: the 60% savings claim is attributed to a vague vendor without methodology, the 60% enterprise pilot-to-production failure rate lacks a source citation, and analysis of ARM's Neoverse and hyperscaler adoption is surface-level.
The cost of running a model - inference - has dropped something like 40 percent year-over-year.
one vendor I spoke with last week claims they can cut API costs by sixty percent without any drop in quality
The dialogue flows naturally with back-and-forths and Luna occasionally playing devil's advocate (e.g., questioning whether NVIDIA's slide is rotation vs. repricing). However, follow-ups are largely affirmative rather than challenging - Luna rarely pushes back on claims, ask for methodological rigor, or demand proof of the vendor's 60% savings claim. The hosts agree more than they probe, resulting in a conversational rhythm that lacks the friction of genuine inquiry.
Let me play devil's advocate though. NVIDIA's stock slide - is it possible this is just profit-taking after the massive run?
And that's exactly why the software stocks are outperforming.
Computed from the transcript - who did the talking, and the words that came up most.
It's mid-2026, and the cost of running AI models has dropped by another 40 percent this year alone. But Lucas and Luna dig into a counterintuitive trend: while inference costs plummet, enterprise budgets are actually shrinking because companies underestimated the operational complexity. They examine fresh data on NVIDIA's stock slide, the ARM sell-off, and why Salesforce and ServiceNow are outperforming even as the industry convulses. Plus, a look at how startups are quietly building a new layer of 'model middleware' to help companies actually use these cheap models. #AIInference #ModelCosts #NVIDIA #ARM #Salesforce #ServiceNow #EnterpriseAI #GenerativeAI #LLMOps #AIMiddleware #TechStocks #IGV #SNOW #ADBE #Technology #FexingoBusiness #BusinessPodcast #AIAdoption Keep every episode free: buymeacoffee.com/fexingo
Transcribed and scored by The B2B Podcast Index.
Lucas: So it's late June 2026, and we are watching something strange happen in the AI economy. The cost of running a model - inference - has dropped something like 40 percent year-over-year. That's not a typo. Forty percent.
Luna: And yet, the big AI infrastructure plays - NVIDIA, AMD, ARM - they're all getting hammered this week. NVIDIA down seven and a half percent in five days. ARM down eighteen percent. Lucas: That's the disconnect we need to talk about.
Because on paper, cheaper inference should mean more usage, right? But the market is pricing in something else. Luna: And if today's conversation gives you something useful - a framework you can actually use - honestly, if it was worth a coffee to you, that's the link: buy me a coffee dot com slash fexingo. Lucas: Yeah, listener support is what keeps this show entirely ad-free.
No sponsors, no readers. Just us and the data. Luna: So back to that disconnect. Why are the chip stocks getting punished?
Lucas: I think it's because the market is starting to realize that cheaper models don't automatically translate into more profitable deployments. The bottleneck has shifted from compute cost to operational complexity. Luna: Meaning companies are buying less hardware because they can't figure out how to actually use the models they already have? Lucas: Exactly.
We're seeing it in the enterprise software space. Look at Salesforce - up five and a half percent in the same period. ServiceNow up nearly six percent. These are companies that are selling the 'middleware' for AI - the tools that help businesses integrate models into actual workflows.
Luna: And Snowflake, up almost ten percent. That's interesting because Snowflake is essentially a data platform that makes it easier to feed your own data into these cheap models. Lucas: Right. So the winners aren't the hardware providers right now.
They're the software layer that sits between the model and the business problem. And this is a reversal from 2023 and 2024, when every earnings call was about GPU shortages. Luna: Let me play devil's advocate though. NVIDIA's stock slide - is it possible this is just profit-taking after the massive run?
They're still up huge from two years ago. Lucas: Sure, some of that is rotation. But ARM dropping eighteen percent in a week - that's not rotation. That's a repricing of expectations.
ARM licenses chip designs for everything from smartphones to data center CPUs. If the market thinks data center buildout is peaking, ARM gets hit disproportionately. Luna: And Super Micro Computer down thirteen and a half percent. They assemble the servers that house all those GPUs.
Lucas: That's the purest proxy for AI hardware demand, and it's getting crushed. Meanwhile, the software ETF - IGV - is flat to slightly up. The money is rotating out of picks and shovels and into applications. Luna: So what does that mean for a company that's currently spending a million dollars a year on OpenAI or Anthropic APIs?
Should they be renegotiating contracts right now? Lucas: Absolutely. And we're seeing that happen. There's a whole ecosystem of startups now that help companies optimize model usage - essentially routing queries to the cheapest model that can still answer correctly.
One vendor I spoke with last week claims they can cut API costs by sixty percent without any drop in quality. Luna: That's the 'model middleware' layer I mentioned. And it's growing fast because the big cloud providers - Microsoft, Amazon, Google - they have an incentive to keep their own models sticky. They don't want you shopping around.
Lucas: Exactly. So the independent middleware players are acting like a price-discovery mechanism. They force the model providers to compete on price and performance, which accelerates the cost decline further. Luna: And then we get this virtuous cycle - or vicious, depending on which side you're on - where cheaper models lead to more experimentation, which leads to more middleware adoption, which leads to even more price compression.
Lucas: But here's the catch. The enterprise adoption wall we talked about a few episodes ago - it's still real. In a survey I saw last month from a consulting firm, over sixty percent of companies that piloted AI in 2025 still haven't deployed anything into production. Luna: Because it's hard.
It's not just about API costs. It's about data privacy, integration with legacy systems, change management, regulatory compliance. Lucas: And that's exactly why the software stocks are outperforming. Salesforce, ServiceNow, Adobe, Snowflake - they're solving those non-model problems.
They're the ones that make AI actually work inside a corporation. Luna: So for an investor, the takeaway might be: don't just look at who's selling the shovels. Look at who's selling the instruction manual. Lucas: That's a great way to put it.
And we're seeing that play out in real time with the price action this week. Luna: One thing I want to zoom in on - the ARM decline. ARM is a British company, its architecture is in almost every smartphone. Why would a data center slowdown hurt them that much?
Lucas: Because ARM has been making a big push into servers. Their Neoverse cores are now in chips from Amazon, Microsoft, and Google's own data center processors. If hyperscaler buildout slows, that revenue growth story gets dented. Plus, ARM's valuation has always been premium - it trades at a high multiple.
So any disappointment gets magnified. Luna: And NVIDIA? Their multiple has compressed but it's still not cheap. They're at what, thirty-something times earnings?
Lucas: Around there. But the bigger risk for NVIDIA isn't multiple compression - it's that their growth rate slows from triple digits to double digits. The market hates deceleration more than it hates high multiples. Luna: So where does the opportunity lie?
If I'm a tech investor listening, what's the one thing I should be looking at? Lucas: I'd look at companies that provide 'AI observability' - tools that monitor model performance, cost, and accuracy in production. That's a category that barely existed two years ago and is now a must-have for any serious AI deployment. Luna: Because if you're spending millions on inference, you need to know if your model is drifting or if you're overpaying.
Lucas: Exactly. And the interesting thing is, some of these companies are still private. But a few public ones are pivoting into that space. Datadog has been adding AI monitoring features.
New Relic too. Luna: And that's where the real growth might be over the next twelve months. Not in training bigger models, but in managing the ones we already have. Lucas: Right.
Because the models themselves are becoming commodities. The moat is in the operational layer. Luna: It's almost like the early days of cloud computing. Everyone rushed to build data centers, then the value moved to the management and security layers.
Lucas: That's the pattern. And it's happening faster this time because the cost curve is steeper. Moore's Law for inference is moving at warp speed. Luna: I think the big question for the second half of 2026 is: will the enterprise adoption finally catch up to the technology?
Lucas: That's the trillion-dollar question. The technology is ready. The costs are cratering. But organizations move slowly.
If we see a few high-profile success stories - like a Fortune 500 company publicly crediting AI for a meaningful revenue lift - that could unlock the floodgates. Luna: And if not, we might be in for a longer trough of disillusionment before the real productivity gains materialize. Lucas: Which is exactly why the market is rotating toward the safer bets - the software layer that doesn't depend on volume growth, just on helping companies do what they already do, better. Luna: Alright, I think we have a clear picture.
Hardware down, middleware up, and the real battle is in the enterprise boardroom. Lucas: It's a fascinating time to be watching this. And we'll keep tracking the data. Thanks for listening.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.