The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/ChatGPT and Beyond with Fexingo
ChatGPT and Beyond with Fexingo artwork

Why AI Model Costs Are Crashing Faster Than Expected

ChatGPT and Beyond with Fexingo · 2026-06-26 · 7 min

0:00--:--

Key moments - from our scoring

Substance score

37 / 100

Five dimensions, 20 points each

Insight Density9 / 20
Originality7 / 20
Guest Caliber3 / 20
Specificity & Evidence12 / 20
Conversational Craft6 / 20

The AI chip market is experiencing a sharp downturn - NVIDIA down 7%, Broadcom 8%, ARM 21% - but this isn't evidence of AI hype collapse; rather, it reflects a fundamental market realignment driven by rapidly falling inference costs. Cost per token has dropped 43% year-over-year (from roughly 70 cents to 40 cents per million tokens), driven by hardware efficiency gains from both NVIDIA's next-gen chips and emerging competitors like AMD, plus architectural improvements in models from OpenAI, Google, and open-source developers. Lucas and Luna explore the Jevons paradox at work in AI: as inference becomes cheaper, demand explodes, offsetting price declines and unlocking previously uneconomical use cases - mid-market logistics companies moving from $50K/month chatbot quotes to under $10K exemplifies this shift. The real winners are end users and software companies (ServiceNow showing resilience) that can pass savings to enterprises; the losers are chip makers betting on margin compression and companies like Palantir facing unrelated headwinds. Most enterprises should rent models via API rather than build custom systems, while hyperscalers diversify into custom silicon (Google TPU, Amazon Trainium) to reduce NVIDIA dependency, creating a two-tier market. The episode targets investors, enterprise software buyers, and founders deciding between proprietary vs. API-based AI strategies.

Key takeaways

  • →Inference cost per token dropped 43% in one year (from ~$0.70 to ~$0.40 per million tokens), driven by hardware efficiency gains and competitive pressure from AMD and custom silicon, not declining demand.
  • →The cost crash unlocks previously uneconomical use cases - a logistics company's AI chatbot quote fell from $50k to under $10k monthly, making projects viable that were rejected months earlier.
  • →Hyperscalers are building custom silicon (Google TPU, Amazon Trainium) to diversify away from premium NVIDIA chips, creating a two-tier market and compressing margins rather than indicating demand collapse.
  • →Most companies should use off-the-shelf API models rather than building custom models; only hyperscalers and large enterprises have the scale to justify custom training.
  • →Total AI spending continues rising despite per-token price drops, as volume growth from expanded use cases offsets deflation - a classic Jevons paradox dynamic.

In this episode

  1. 1AI Chip Stocks Crash Amid Inference Cost Collapse
  2. 2The Paradox: Prices Drop But Total AI Spending Rises
  3. 3Hardware Competition and Margin Compression in the Chip Market
  4. 4How Cost Reductions Unlock New Enterprise Use Cases
  5. 5Strategic Implications: APIs vs. Custom Model Training
  6. 6Market Winners and Losers in the Cost Deflation
  7. 7Adoption Metrics and Investment Outlook

Mentioned

OpenAIGoogleMicrosoftNVIDIAAmazonAMDBroadcomARMPalantirServiceNowGPT-4Gemini

Guests

Luna

Topics in this episode

GPT-4Jevons ParadoxNvidiaGoogle GeminiBroadcomGoogle TPUAmazon TrainiumInference cost per tokenARMMicrosoft Azure AI

Questions this episode answers

Why did NVIDIA stock fall if AI demand is still growing?

NVIDIA's decline reflects margin compression from competitors like AMD and custom silicon from hyperscalers (Google TPU, Amazon Trainium), not demand collapse. Hyperscalers are diversifying away from NVIDIA's premium chips to cost-optimized alternatives as inference costs drop.

What has the cost per token for AI inference dropped to in 2026?

Cost per token has fallen to around 40 cents per million tokens, down from roughly 70 cents a year ago - a 43% decline driven by hardware efficiency and model architecture improvements.

How is the Jevons paradox affecting enterprise AI adoption?

As inference costs fall, companies that previously couldn't justify AI projects financially now can; for example, a logistics company moved from a $50K/month chatbot quote to under $10K and greenlit the project, exemplifying how volume growth offsets price declines.

Should mid-size companies build custom AI models or use API services?

Most companies should use off-the-shelf models via API; only hyperscalers and the largest players have sufficient scale to justify custom model training. The strategic focus should be on application-layer value, not model development.

Are software companies faring better than chip makers in this market shift?

Yes - software automation companies like ServiceNow show relative resilience (down 5.5%) compared to chip stocks, because they can pass inference savings directly to enterprises, creating a new value capture opportunity.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

9 / 20

The episode delivers a handful of genuinely useful data points about inference cost declines and chip stock dynamics in a short runtime, but the density is hampered by repetitive affirmations ('Exactly', 'Right') and surface-level analysis that rarely goes deeper than a tech finance newsletter summary.

The inference cost per token has dropped from roughly seventy cents per million tokens a year ago to around forty cents today. That's a forty-three percent decline.
NVIDIA's stock is reflecting margin compression, not demand collapse.

Originality

7 / 20

The Jevons paradox framing is the episode's one intellectual anchor, but it is a well-circulated concept applied without extension or nuance; the rest of the takes - hyperscalers building custom silicon, application-layer value over custom training - are standard AI commentary recycled from mainstream tech coverage.

it's the Jevons paradox for AI. As something becomes cheaper, people use more of it.
For most companies, the smart move is to use off-the-shelf models via API. Only a handful of the largest players - the hyperscalers, some big banks, maybe a few others - have the scale to justify custom training.

Guest Caliber

3 / 20

There is no guest; the episode is a co-hosted chat between Lucas and Luna, neither of whom establishes any practitioner credentials, and the only first-hand evidence offered is a secondhand anecdote about a friend's company, which signals limited domain authority.

A friend at a mid-size logistics company told me they were evaluating an AI chatbot for customer service. Six months ago the quote was like fifty thousand a month just in inference costs. Now it's under ten.

Specificity & Evidence

12 / 20

The episode earns credit for citing real stock moves with percentages, a concrete cost-per-token trajectory, and named companies across the stack, but the logistics anecdote is secondhand and the broader cost decline figures are attributed only to 'some estimates', limiting evidential weight.

NVIDIA off seven percent in five days. Broadcom down almost eight. Even ARM, which for months seemed invincible, down twenty-one percent.
Palantir is down over sixteen percent this week - that's a different story, more about government contract delays. But ServiceNow, which is all about enterprise automation, only down five and a half percent.

Conversational Craft

6 / 20

The hosts function as an echo chamber - Lucas makes a claim and Luna affirms it, then reverses - with no genuine pushback, no probing follow-ups, and a mid-episode donation solicitation that further breaks analytical momentum; the counterargument about why anyone would invest in custom models is raised but immediately resolved without tension.

Lucas: Exactly. And that's what we're seeing.
Lucas: Exactly. And the data backs that up.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

lucas13luna12percent8cost8nvidia6inference5forty5enterprise5market5chip4costs4prices4stock4token3price3google3

Episode notes

In this episode of ChatGPT and Beyond with Fexingo, Lucas and Luna explore the dramatic drop in AI inference costs and what it means for the industry. With NVIDIA stock down 7% in a week and inference costs falling 40% year-over-year, the hosts discuss how cheaper models are reshaping business adoption. They examine the paradox of falling prices and rising investment, using real data to explain why companies are spending more on AI even as token prices plummet. A must-listen for anyone trying to understand the real economics of generative AI in 2026. #AIInferenceCosts #GenerativeAI #NVIDIA #OpenAI #TokenEconomics #EnterpriseAI #AIAdoption #Technology #TechPodcast #FexingoBusiness #BusinessPodcast #IndustryTrends #CostCrash #PriceParadox #AIInvestment #ModelEfficiency #June2026 #AIMarket Keep every episode free: buymeacoffee.com/fexingo

Full transcript

7 min

Transcribed and scored by The B2B Podcast Index.

Lucas: So it's June 26, 2026, and if you've been watching the AI chip stocks this week, you've seen a bloodbath. NVIDIA off seven percent in five days. Broadcom down almost eight. Even ARM, which for months seemed invincible, down twenty-one percent.

But right alongside that, we have this headline: inference costs are crashing - some estimates say forty percent year-over-year. Luna: And a lot of people are connecting those two dots and saying, see, the AI hype is over. But I think that's too simple. Lucas: Way too simple.

Look, the cost per token - basically the price of generating one unit of output from a large language model - has been dropping like a stone. OpenAI has cut prices on GPT-4 multiple times. Google's Gemini is cheaper. Open-source models are practically free to run if you have the hardware.

But here's the paradox: total spending on AI is still going up. It's not a contraction. It's a realignment. Luna: Right, it's the Jevons paradox for AI.

As something becomes cheaper, people use more of it. So the demand explosion offsets the price drop. Lucas: Exactly. And that's what we're seeing.

Enterprise AI adoption was stalling earlier this year - we did an episode on that. But the cost crash is actually unlocking new use cases. Things that were too expensive to run at scale six months ago are now viable. So the question is: are we in a deflationary spiral for AI, or is this just the normal maturation curve?

Luna: And who wins and loses? Because if NVIDIA is down on falling prices, but companies like Microsoft are still spending billions on AI infrastructure, something doesn't add up. Lucas: Let's break it down. The inference cost per token has dropped from roughly seventy cents per million tokens a year ago to around forty cents today.

That's a forty-three percent decline. And it's driven by two things: hardware efficiency and model architecture improvements. NVIDIA's next-gen chips are more efficient, but also competitors like AMD and startups are eating into the market. So NVIDIA's stock is reflecting margin compression, not demand collapse.

Luna: So the market is pricing in that the hyperscalers - Microsoft, Amazon, Google - are going to diversify away from NVIDIA's most expensive chips. Lucas: Exactly. And the data backs that up. Microsoft's stock is down seven percent this week too, but that's more about the broader tech sell-off.

The bigger story is that Azure AI revenue is still growing. The cloud giants are building out their own custom silicon - Google's TPU, Amazon's Trainium - and that's creating a two-tier market. One tier for the premium NVIDIA chips, another for the cost-optimized alternatives. Luna: Right.

So if you're a startup building an AI product, your margins just got a lot better. But if you're a chip company betting on infinite demand at high prices, you might need to adjust. Lucas: And that adjustment is exactly what we're seeing in the stock moves. But here's the thing - the cost crash is actually accelerating enterprise adoption.

We had that headline a few weeks ago about enterprise AI stalling. Well, the reason was partly that projects were too expensive to justify the ROI. At forty cents per million tokens, the math changes. Luna: Yeah, I've seen that firsthand.

A friend at a mid-size logistics company told me they were evaluating an AI chatbot for customer service. Six months ago the quote was like fifty thousand a month just in inference costs. Now it's under ten. So they pulled the trigger.

Lucas: That's the Jevons paradox in action. And it's not just chatbots. Companies are embedding AI into workflow automation, data analysis, even code generation. The cost per token is becoming negligible for many tasks.

So the total addressable market expands. Luna: But there's a counterargument: if the cost is dropping that fast, why would anyone invest in building their own models? Why not just rent the cheapest API? Lucas: That's where the strategic play gets interesting.

For most companies, the smart move is to use off-the-shelf models via API. Only a handful of the largest players - the hyperscalers, some big banks, maybe a few others - have the scale to justify custom training. The rest should focus on application-layer value. Luna: Honestly, if today's conversation gave you a clearer picture on where AI costs are heading, that's the kind of insight that makes the show worth a coffee to you.

There's a link - buy me a coffee dot com slash fexingo. It's a tiny gesture that keeps this ad-free and independent. Lucas: Yeah, and it really does make a difference. Keeps us focused on the numbers and the stories that matter, not on selling ads.

So, back to the cost crash - Lucas: - I think the real winner from all this is the end user. Consumers and businesses are getting access to AI capabilities at prices that were unthinkable two years ago. But the question for investors is: will volume growth compensate for price declines? And so far, the evidence suggests yes.

Luna: But we need to watch the data. If inference costs drop another forty percent next year, the market might reprice again. That's a risk for chip stocks, but an opportunity for software companies that can pass on savings. Lucas: Right.

And we're already seeing that divergence. Palantir is down over sixteen percent this week - that's a different story, more about government contract delays. But ServiceNow, which is all about enterprise automation, only down five and a half percent. That's relatively resilient.

Luna: So the takeaway is: don't look at the chip stock sell-off and assume AI is dying. Look at the adoption metrics, the API call volumes, the enterprise surveys. Those tell a different story. Lucas: Exactly.

And we'll keep tracking that. Next episode, let's look at the energy side - because more inference means more power consumption, and that's its own set of bottlenecks. For now, I'm Lucas. Luna: And I'm Luna.

Thanks for listening.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • How Idempotency-Key Design Prevents Payment DisastersThe Developer Tools Podcast with Fexingo · features Luna98 / 100
  • How Kubernetes Topology Spread Constraints Create Scheduling HotspotsDevOps Daily with Fexingo · features Luna95 / 100
  • Why Pipeline Velocity Trumps Deal Size Every TimeThe Growth Operator with Fexingo · features Luna95 / 100
  • Why Enterprise Software Deals Now Include a Vendor AI Model Explainability MandateB2B SaaS Talks with Fexingo · features Luna94 / 100
  • How B2B Brands Wreck Pipeline with Unsyncroned CRM DataThe Marketing Operator Podcast with Fexingo · features Luna92 / 100
  • Why Marketing Attribution Misses the Seasonality PatternMarketing Analytics with Fexingo · features Luna91 / 100

More from ChatGPT and Beyond with Fexingo

All episodes →
  • Why Palantir Stock Is Surging While AI Peers Decline72 / 100
  • How AI Model Marketplaces Are Becoming the New Operating Systems68 / 100
  • Why AI Model Marketplaces Are Becoming the New Operating Systems70 / 100
  • Neoclouds Are the New AI Infrastructure Battlefield72 / 100
  • How AI Chips Are Shifting From Training to Inference76 / 100
Explore the best B2B AI & Data podcasts →
All ChatGPT and Beyond with Fexingo episodes →