The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/ChatGPT and Beyond with Fexingo
ChatGPT and Beyond with Fexingo artwork

Why AI Model Costs Are Crashing Faster Than Anticipated

ChatGPT and Beyond with Fexingo · 2026-06-29 · 8 min

0:00--:--

Key moments - from our scoring

Substance score

52 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality10 / 20
Guest Caliber6 / 20
Specificity & Evidence12 / 20
Conversational Craft11 / 20

The collapse in AI inference costs is rewriting the software economics equation in real time. Token pricing from OpenAI, Anthropic, and Google has plummeted from roughly 20 cents per million tokens at the start of 2025 to under a dime, with some providers pricing below five cents - a shift that unlocks previously uneconomical use cases like customer service bots and document summarization for small businesses. This cost crash is creating new market dynamics: training-focused hardware companies like Super Micro Computer are losing momentum, while inference-infrastructure players like Snowflake and AMD gain ground. NVIDIA remains dominant but faces pressure as the market shifts away from massive training clusters toward distributed inference. Enterprise adoption patterns are changing too - companies are discovering that off-the-shelf models deliver 90% of results at a fraction of the fine-tuning cost, commoditizing the core model product. This forces model labs like Anthropic into volume plays (discounting to government agencies) and free-tier strategies (Gemini's free image generation) to lock in users. The real defensibility in AI increasingly comes not from the models themselves, but from proprietary data, distribution platforms (GPT Store, Vertex AI Model Garden, Hugging Face), and embedded workflows - explaining why Adobe and ServiceNow are outperforming pure model providers.

Key takeaways

  • →Inference costs have fallen 40% year-over-year and pricing continues dropping below a nickel per million tokens, fundamentally changing which AI use cases are economically viable.
  • →The market is shifting from training infrastructure (hurting companies like Super Micro) to inference infrastructure (benefiting AMD and distributed inference platforms), creating new winners and losers in hardware.
  • →Model providers are being commoditized - the defensible moat is no longer the model itself but proprietary data, distribution channels, domain expertise, and embedded workflows in platforms like Adobe and ServiceNow.
  • →Enterprise adoption is stalling on custom fine-tuned models because cheap off-the-shelf models now deliver 90% of the performance at a fraction of the cost, driving volume but squeezing margins.
  • →Memory supply and high-bandwidth memory costs for inference could become the next bottleneck, offsetting efficiency gains as demand shifts toward inference at scale.

Topics in this episode

OpenAIAnthropicGoogleNvidiaSnowflakeARMSuper Micro ComputerAMD MI series chipsGPT StoreVertex AI Model Garden

Questions this episode answers

How much have AI inference costs dropped in the past year?

Inference costs have fallen approximately 40% year-over-year, with per-million-token pricing from major providers like OpenAI, Anthropic, and Google dropping from around 20 cents at the start of 2025 to under a dime, with some models pricing below a nickel.

Why are training infrastructure companies like Super Micro Computer declining while inference companies like Snowflake are rising?

The market is shifting demand away from massive training clusters (which require huge centralized server farms) toward distributed inference infrastructure, which can be more widely distributed across different systems and doesn't require the same capital-intensive server setup.

What is creating the competitive moat in AI now that models are becoming commoditized?

With model costs collapsing, defensibility comes from proprietary data, distribution platform control, embedded workflows in enterprise tools, and domain expertise - exemplified by Adobe's AI integration in creative tools and ServiceNow's enterprise workflow embedding.

Why are model providers like Anthropic offering discounted pricing to government agencies?

It's a volume-lock strategy to establish sticky, long-term contracts before prices fall further, since government procurement cycles are lengthy and create lasting vendor relationships.

What downstream effect could prevent inference costs from falling further?

High-bandwidth memory supply constraints and costs - referred to as 'RAMageddon' by South Korean tech giants investing $550 billion to address it - could become a bottleneck that offsets efficiency gains in inference at scale.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode packs substantive economic claims (40% YoY inference cost decline, token pricing drops from 20 cents to sub-nickel) and traces meaningful downstream effects (training-to-inference shift, hardware market winners/losers, enterprise ROI realization). However, it relies heavily on stock price movements as proxy evidence and spends significant time on already-circulating themes (moat = data/distribution, commoditization of models) without drilling deeply enough to feel packed. The final 2 minutes devolve into fundraising language and soft speculation ('ambient AI') that dilutes density.

the per million token price for their smallest models has dropped from around 20 cents at the start of 2025 to under a dime today. Some are below a nickel.
Companies are realizing that the ROI on custom fine-tuned models doesn't pencil out when a cheap off-the-shelf model can do 90 percent of the job.

Originality

10 / 20

The core observation - cost crashes reward inference/distribution moats over raw model access - is sound but not novel; it's been discussed across enterprise AI discourse for months. The training-to-inference shift, the Adobe/ServiceNow examples, and the emphasis on data/workflows as defensible positions are all conventional wisdom by early 2025. The Anthropic-Newsom deal and stock-market framing offer modest freshness, but the underlying analysis largely rehashes established frameworks without genuine counterintuitive push or first-principles work.

the defensible moat in AI isn't the model anymore - it's the data, the distribution, the domain expertise.
Anthropic's deal with California Governor Newsom - they're offering Claude to state agencies at half price. It's a volume play.

Guest Caliber

6 / 20

This is a two-person co-hosted format with no external guest. Lucas and Luna appear to be the hosts/producers of Fexingo, but the transcript provides no biographical detail on either - no titles, track records, or evidence they have built or operated AI-backed businesses at scale. They speak knowledgeably about market trends and stock moves, but they read more as intelligent analysts/commentators than seasoned practitioners. Without demonstrated operator credentials, this falls below the caliber expected for substantive B2B learning.

We've been hearing about cost declines for a while, but the numbers this quarter are wild.
And that's a theme we keep coming back to: the defensible moat in AI isn't the model anymore

Specificity & Evidence

12 / 20

The episode grounds claims in concrete data: 40% YoY cost decline, token prices from 20 cents to under a dime, Super Micro down 15.5%, Snowflake up 9%, NVIDIA down 2.5%, AMD up 4%, stock moves for Adobe and ServiceNow. Named companies and ticker moves lend surface specificity. However, many claims lack sourcing ('a report that inference costs have fallen'), the stock correlations are asserted without detailed causal evidence, and no dollar figures are given for actual AI workload costs or adoption patterns. The Anthropic-Newsom deal and 'RAMageddon' are mentioned but not quantified beyond the 550 billion SK investment.

the per million token price for their smallest models has dropped from around 20 cents at the start of 2025 to under a dime today.
Super Micro Computer, for example, is down 15 and a half percent over the past five days.

Conversational Craft

11 / 20

Lucas and Luna trade observations smoothly and occasionally build on each other's points (e.g., Luna connecting cost crashes to 'ambient AI'), but the dialogue lacks genuine challenge or skepticism. No claim is pressure-tested; when Luna floats the 'counterintuitive' ARM stock decline, Lucas affirms it immediately rather than pushing back. There are few probing follow-ups on data sources, no pursuit of counterarguments, and the two rarely disagree. The tone reads as collaborative analysis rather than rigorous inquiry, and both defer to market price action as truth without questioning its validity as evidence.

Luna: That makes a lot of sense. And it's a smart move because government procurement cycles are long
Lucas: That's a good read. The other big story tied to cost drops is enterprise adoption stalling

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

lucas18luna17model14cost12inference9percent9data6keep6models5training5building4costs4seeing4market4cheap4crash4

Episode notes

In this episode, Lucas and Luna dive into the latest data showing AI model inference costs dropping over 40% in the first half of 2026, with some small models now running for pennies per million tokens. They examine the implications for startups, enterprise adoption, and the broader shift from training to inference. Using real market data - including Super Micro Computer's 15.5% drop and Snowflake's 9.2% gain - they connect the cost crash to hardware spending changes and the rise of model marketplaces. The conversation also touches on Anthropic's deal with California to offer Claude at half price, signaling how government adoption could accelerate. A must-listen for anyone tracking AI's economic inflection point. #AIInference #ModelCosts #EnterpriseAI #AICrash #GenerativeAI #LLMs #AIEconomics #CloudComputing #NVIDIA #Snowflake #SuperMicro #Anthropic #AIMarketplaces #AISpending #Technology #BusinessPodcast #FexingoBusiness #ChatGPTandBeyond Keep every episode free: buymeacoffee.com/fexingo

Full transcript

8 min

Transcribed and scored by The B2B Podcast Index.

Lucas: If today's technology conversation gave you one usable takeaway, I think it should be this: the cost to run an AI model is dropping so fast that the economics of building software around it are basically being rewritten in real time. Luna: We've been hearing about cost declines for a while, but the numbers this quarter are wild. I saw a report that inference costs have fallen something like 40 percent year-over-year. Lucas: At least.

And it's accelerating. If you look at the token pricing from the major model providers - OpenAI, Anthropic, Google - the per million token price for their smallest models has dropped from around 20 cents at the start of 2025 to under a dime today. Some are below a nickel. Luna: That changes the math on a lot of use cases that were previously too expensive.

Customer service bots, document summarization for small businesses - things that needed a really efficient model to justify the cost. Lucas: Exactly. And one of the more interesting downstream effects is what we're seeing in the hardware market. Super Micro Computer, for example, is down 15 and a half percent over the past five days.

That's a company that was riding high on server demand for AI training. The market is pricing in a shift. Luna: From training infrastructure to inference infrastructure. And inference is a lot more distributed - it doesn't need the same massive clusters that training does.

Lucas: Right. So the companies that bet big on building out giant training factories are seeing that demand soften, while the ones focused on making inference cheap and accessible are thriving. Snowflake is up over 9 percent in the same period, and that's not a coincidence - they've been leaning hard into ai powered data workloads. Luna: So the cost crash is actually accelerating the shift from training to inference that we've been talking about for months.

Lucas: And it's creating a new set of winners and losers. On the one hand, you have NVIDIA, still the dominant player, but down two and a half percent this week. On the other, you have AMD up nearly 4 percent, because they're gaining ground in the inference market with their MI series chips. Luna: And ARM is down over 6 percent, which seems counterintuitive because their architecture is everywhere in mobile and edge inference.

But maybe the market sees them as more exposed to the smartphone cycle than the data center buildout. Lucas: That's a good read. The other big story tied to cost drops is enterprise adoption stalling - or rather, shifting. We talked about that in episode 78, but the data keeps coming in.

Companies are realizing that the ROI on custom fine-tuned models doesn't pencil out when a cheap off-the-shelf model can do 90 percent of the job. Luna: So the cost crash is actually a double-edged sword for the big AI labs. It drives volume, but it commoditizes their core product. Lucas: Exactly.

And that's why we're seeing moves like Anthropic's deal with California Governor Newsom - they're offering Claude to state agencies at half price. It's a volume play. Lock in government contracts now, before the price falls even further. Luna: That makes a lot of sense.

And it's a smart move because government procurement cycles are long - if you get in now, you're sticky for years. Lucas: Meanwhile, the model marketplaces we've talked about in previous episodes are becoming the primary distribution channel. OpenAI's GPT Store, Google's Vertex AI Model Garden, Hugging Face - they're all competing to be the platform where developers discover and deploy these cheap models. Luna: And the margins for the model providers themselves are getting squeezed.

Which is maybe why we're seeing more of them offer free tiers, like Gemini's personalized image generation going free for US users. Lucas: Right. That's a land-grab strategy. Get users hooked on free, then upsell them to paid tiers later.

But if the cost keeps dropping, the paid tiers have to keep getting cheaper too. Luna: So where does that leave the startups that are building on top of these models? They're in a pretty good spot, I think - their input costs are plummeting. Lucas: For the ones that have a clear value-add, yeah.

But if your entire business is just wrapping an API with a slightly better UI, you're going to get squeezed as the underlying model becomes a commodity. The winners will be the ones with proprietary data or workflows that can't be replicated. Luna: And that's a theme we keep coming back to: the defensible moat in AI isn't the model anymore - it's the data, the distribution, the domain expertise. Lucas: One specific example: look at Adobe, up 4.

6 percent this week. They're integrating generative AI directly into their creative tools. The model itself is almost irrelevant - what matters is that designers are already in the Adobe ecosystem. That's their moat.

Luna: And then there's ServiceNow, up 4.2 percent - they're embedding AI into enterprise workflows. Same idea. Lucas: So the cost crash is good for adoption, good for consumers, good for startups with differentiated data - but it's brutal for anyone selling raw model access.

And I think that's going to drive more consolidation in the next twelve months. Luna: We might see some of the smaller model providers get acquired or just fold. The capital requirements to keep up with the frontier models are enormous. Lucas: And the hardware side is consolidating too.

The South Korean tech giants committing over 550 billion dollars to ease what they're calling 'RAMageddon' - that's a signal that memory supply is going to be tight even as demand shifts. Luna: Memory is a huge part of inference costs. If the cost of high-bandwidth memory doesn't come down, it could offset some of the efficiency gains. Lucas: That's the wrinkle.

But overall, the trend is clear: inference is getting dramatically cheaper, and that's unlocking use cases that were science fiction two years ago. Luna: And it's happening faster than almost anyone predicted. Lucas: Yeah. And honestly, one reason we can keep digging into these shifts every week is that a small group of listeners supports the show directly.

We don't run ads, we don't have sponsors - it's just listener-funded through buy me a coffee dot com slash fexingo. Luna: It's true. And that support means we can stay independent and follow the stories that actually matter, rather than chasing clickbait. Lucas: So for anyone who finds value in these conversations and wants to keep them going, that's the way.

But back to the cost crash - the question I keep coming back to is: what happens when running a state of the art model costs less than a cup of coffee? Because we're getting close. Luna: I think we see an explosion of ambient AI - always-on assistants, real-time translation, personalized tutors. Things that don't exist today because the cost was prohibitive.

Lucas: And regulators are going to have to catch up fast. When AI is that cheap and ubiquitous, the surface area for misuse expands exponentially. Luna: That's a conversation for another episode. Lucas: Definitely.

For now, the takeaway is: if you're building something with AI, check your cost assumptions every quarter because they're probably outdated. And if you're investing, pay attention to who owns the distribution - not just the model.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • AI CEO Series: Dr. Varun SivaramNext in Tech · on Nvidia86 / 100
  • 512. Is SpaceX Over or Undervalued, Why Consensus Kills, How Chewy Beat Amazon, and the GameStop Saga from a Board Member (Larry Cheng)The Full Ratchet (TFR) · on Anthropic86 / 100
  • DeepSeek's $50B Round, OpenAI's Delayed IPO, and the GP Stakes Market with CAZ Investmentstrading places · on OpenAI86 / 100
  • The New American Dream: Democratising InvestingThe Master Investor Podcast with Wilfred Frost · on OpenAI84 / 100
  • Agentic Engineering for Testers: How to Automate Your Way to the Top with Amit RawatTestGuild Automation Podcast · on OpenAI82 / 100

More from ChatGPT and Beyond with Fexingo

All episodes →
  • Why Palantir Stock Is Surging While AI Peers Decline72 / 100
  • How AI Model Marketplaces Are Becoming the New Operating Systems68 / 100
  • Why AI Model Marketplaces Are Becoming the New Operating Systems70 / 100
  • Neoclouds Are the New AI Infrastructure Battlefield72 / 100
  • How AI Chips Are Shifting From Training to Inference76 / 100
Explore the best B2B AI & Data podcasts →
All ChatGPT and Beyond with Fexingo episodes →