The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/ChatGPT and Beyond with Fexingo
ChatGPT and Beyond with Fexingo artwork

Why AI Hardware Spending Is Shifting From Training to Inference

ChatGPT and Beyond with Fexingo · 2026-06-26 · 7 min

0:00--:--

Key moments - from our scoring

Substance score

44 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality7 / 20
Guest Caliber4 / 20
Specificity & Evidence13 / 20
Conversational Craft9 / 20

The AI hardware market is experiencing a structural shift that goes beyond typical sector rotation. While software stocks - including AI software names - held steady this week, pure hardware plays like NVIDIA (down 7.7%), AMD (down 5.4%), and Broadcom (down 7%) took significant hits. Lucas and Luna argue this reflects investor recognition that AI spending is pivoting from training infrastructure to inference deployment. Training happens once per model and requires enormous capital (tens to hundreds of millions in compute for frontier models using ten thousand H100s for months), but inference - every GPT query, image generation, or code completion - happens millions of times daily and represents the long tail of compute demand. Specialized inference chips from Groq (LPU), Cerebras (wafer-scale engine), and d-Matrix, plus custom silicon from hyperscalers like Amazon (Inferentia), Google (TPU), and Microsoft (Maia), can deliver 2-5x better cost-efficiency per token than NVIDIA's training-optimized GPUs. This threatens NVIDIA's dominance not through AMD or Intel competition in training, but through vertical integration of inference silicon. Winners include inference-focused software platforms like Snowflake and ServiceNow, whose margins improve as inference costs decline, and inference chip startups that gain market traction. NVIDIA's challenge is maintaining pricing power as its data-center revenue mix - currently 80%+ training - potentially flips toward inference over the next two years.

Key takeaways

  • →NVIDIA and other training-focused chip makers are facing stock pressure as the market prices in a shift from training to inference as the dominant AI compute workload, where specialized chips offer 2-5x better cost efficiency
  • →Inference-specialized hardware from companies like Groq, Cerebras, and d-Matrix, plus custom silicon from hyperscalers (Amazon Trainium/Inferentia, Google TPU, Microsoft Maia), pose a structural threat to NVIDIA's pricing power that goes beyond normal sector rotation
  • →Software platforms like Snowflake and ServiceNow that monetize inference will benefit from lower inference costs expanding their gross margins faster than expected
  • →Cheaper inference economics will unlock new AI applications like real-time customer support agents and autonomous workflows that are currently uneconomical to deploy at scale on NVIDIA hardware
  • →NVIDIA's software stack (TensorRT, Triton) is being deployed as a lock-in strategy for developers, but hardware economics are shifting regardless of software moats

In this episode

  1. 1Hardware Stock Rotation: Training to Inference Thesis
  2. 2Training vs. Inference: Definitions and Market Implications
  3. 3Specialized Inference Chips and Hyperscaler Custom Silicon
  4. 4Cost Economics and Performance Advantages of Inference Chips
  5. 5Software Platforms and Inference Monetization
  6. 6NVIDIA's Pricing Power and Competitive Threats in Inference
  7. 7Inference Shift Enabling New AI Application Categories

Mentioned

NVIDIAAMDBroadcomMicrosoftServiceNowSnowflakeGroqCerebrasd-MatrixGoogleAmazonOracle

Guests

Luna

Topics in this episode

NvidiaBroadcomAMDGroqGoogle TPUAmazon TrainiumCerebrasd-MatrixAmazon InferentiaMicrosoft MaiaTrainiumInferentia

Questions this episode answers

What's the difference between AI training and inference, and why does it matter for hardware spending?

Training builds the model by feeding it billions of data points over months at enormous cost (tens to hundreds of millions for frontier models). Inference runs the trained model millions of times daily - every ChatGPT query or code completion. Training happens once per model version, while inference is recurring, so the market is shifting spending from one-time training infrastructure to the long tail of inference compute.

How cost-effective are specialized inference chips compared to NVIDIA GPUs?

Specialized inference chips like Groq's LPU and Cerebras' wafer-scale engine can be 2-5x more cost-effective per token generated for large language model inference compared to NVIDIA's H100 and B200 GPUs, especially for real-time applications where latency matters.

Which companies benefit from the shift from training spending to inference spending?

Software platforms that monetize inference - like Snowflake (up 10% this week after launching Cortex AI inference service) and ServiceNow (up 5.7%) - benefit from cheaper inference improving their margins. Inference chip startups (Groq, Cerebras, d-Matrix) and hyperscalers' custom silicon efforts (Amazon's Inferentia, Google's TPU, Microsoft's Maia) are also positioned to win.

Why is NVIDIA's stock declining if it dominates the AI chip market?

The market is pricing in a shift where inference becomes the majority of AI compute demand. Since inference doesn't require NVIDIA's most expensive training-optimized chips, and hyperscalers are building custom inference silicon, NVIDIA faces margin compression and competitive pressure in what may become the larger segment of total AI compute.

What new AI applications could become viable if inference costs drop significantly?

Real-time AI agents for tasks like customer support bots that reason through complex policies during live interactions become economically viable at scale with cheaper inference, enabling the next wave of AI applications that change how businesses operate day-to-day.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode packs a coherent thesis and several concrete data points into 7 minutes, but the core claim - that AI compute is shifting from training to inference - is already widely circulated in tech media. The supporting observations (hyperscaler custom silicon, inference chip economics) are useful but not novel for anyone who follows the space.

training happens once per model version. Inference happens millions of times per day, per model, per user. It's recurring. It's the long tail of compute demand.
specialized chips can be two to five times more cost-effective per token generated

Originality

7 / 20

The training-to-inference rotation thesis is a narrative already circulating widely in financial media; the episode presents it coherently but adds no contrarian angle, first-principles reasoning, or genuinely counterintuitive claim. The framing around software stocks outperforming hardware is an observation, not a novel argument.

the easy money in just buying NVIDIA and AMD for the training buildout might be done
the AI hardware story is not dead, but it's rotating

Guest Caliber

4 / 20

There are no guests - just two unnamed hosts (Lucas and Luna) whose professional backgrounds, seniority, or practitioner credentials are never established. There is no indication either has built, deployed, or invested in AI infrastructure at scale; they present as commentators rather than operators.

Lucas: So NVIDIA is down seven-point-seven percent over the past five days.
Luna: Okay, let's step back. For someone who's not deep in the AI jargon

Specificity & Evidence

13 / 20

The episode earns credit for naming specific stock movements, chip vendors (Groq, Cerebras, d-Matrix), hyperscaler silicon programs (Trainium, Inferentia, TPU, Maia), NVIDIA software stack tools, and the Snowflake Cortex launch date - meaningful specificity for a 7-minute show. However, key numbers like NVIDIA's revenue mix are hedged as estimates rather than sourced figures.

A single training run for a frontier model might use ten thousand H100s for three months
Snowflake up nearly ten percent. They just launched their Cortex AI inference service in March

Conversational Craft

9 / 20

Luna's role as structured interlocutor produces some useful clarifying follow-ups and one concrete example prompt, but the dialogue feels largely scripted, with no genuine pushback on speculative claims (e.g., the 80% training revenue estimate goes unchallenged) and no productive disagreement anywhere in the episode.

Like what? Give me a concrete example.
Oracle dropped fifteen percent this week - is that related?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

inference36lucas19luna18nvidia13training13model8percent7software6chips6chip6five5hardware5shift5real5seven4point4

Episode notes

Lucas and Luna unpack a major inflection point in AI infrastructure: the shift from training-focused hardware spending to inference. They examine why NVIDIA's 7.7 percent weekly drop amid a broader AI hardware selloff may signal market recognition that the training buildout is peaking, while inference workloads - and the chips optimized for them - become the next growth frontier. The hosts walk through the economics: training a single large model can cost over $100 million, but inference - actually running that model millions of times for users - is where the recurring revenue lives. They cite Microsoft's 1.5 percent resilience this week as a sign that software platforms monetizing inference are outperforming pure hardware plays. The episode also explores how startups like Groq, d-Matrix, and Cerebras are challenging NVIDIA with inference-specialized chips, and why the hyperscalers (Amazon, Google, Microsoft) are designing their own inference silicon. A concrete look at why the AI chip narrative is shifting in mid-2026.

Full transcript

7 min

Transcribed and scored by The B2B Podcast Index.

Lucas: So NVIDIA is down seven-point-seven percent over the past five days. AMD off five-point-four. Broadcom down nearly seven percent. And the chatter I keep hearing is that this is just a rotation out of tech - but I actually think something more structural is going on.

Luna: You think it's not just a sector rotation? Because the software names held up fine - Microsoft actually up one-point-five, ServiceNow up five-point-seven, Snowflake up nearly ten percent this week. Lucas: Right, exactly. The software stocks are fine.

The AI software names are even up. But the pure hardware plays - the companies that sell the pickaxes in the AI gold rush - they're getting hit. And I think the market is starting to price in a shift from training infrastructure spending toward inference. Luna: Okay, let's step back.

For someone who's not deep in the AI jargon - what's the difference between training and inference, and why does it matter for which stocks go up? Lucas: Training is when you build a model. You feed it billions of data points, it learns patterns, it takes months and costs tens or hundreds of millions of dollars in compute. That's been the dominant demand driver for NVIDIA's data-center GPUs for the last three years.

Luna: And inference is when you actually use the trained model - every time someone chats with GPT-4.6 or generates an image or runs a code assistant, that's inference. Lucas: Exactly. And the thing is, training happens once per model version.

Inference happens millions of times per day, per model, per user. It's recurring. It's the long tail of compute demand. Luna: So the market might be waking up to the idea that the training buildout has peaked - or is peaking - and the next wave is inference, which may not need NVIDIA's most expensive chips.

Lucas: That's the thesis. And you can see it in the numbers. A single training run for a frontier model might use ten thousand H100s for three months. But inference for that same model, once deployed, could consume more total compute over its lifetime - just spread out, at lower margins per chip.

Luna: And there are chips designed specifically for inference that are way more cost-efficient than NVIDIA's training-optimized GPUs. Groq, Cerebras, d-Matrix - they're all building inference-specialized hardware. Lucas: Right. Groq's language processing unit, or LPU, is built for low-latency inference.

Cerebras has their wafer-scale engine. These chips don't train models well, but they run them faster and cheaper than an H100 or a B200 for inference. Luna: And then you have the hyperscalers - Amazon with Trainium and Inferentia, Google with TPU, Microsoft with their Maia chip. They're all designing their own silicon for inference.

Lucas: That's the real threat to NVIDIA's dominance. Not that AMD or Intel take share in training - they've tried and haven't really broken through. It's that the hyperscalers vertically integrate inference silicon, and inference becomes the majority of AI compute. Luna: So the market's looking at NVIDIA's 7.

7 percent drop this week and thinking: maybe the multiple expansion from training demand is behind us, and now we have to value the company on an inference future where they face more competition. Lucas: Exactly. And NVIDIA knows this. They've been pushing their inference software stack - TensorRT, Triton Inference Server - to try to lock developers into CUDA even for inference.

But the hardware economics are shifting. Luna: Let's talk about the economics. What's the actual cost difference between running inference on an NVIDIA chip versus a specialized inference chip? Lucas: It depends on the workload, but for large language model inference - like running a GPT-4-class model - specialized chips can be two to five times more cost-effective per token generated.

And for real-time applications like voice assistants or code completion, latency matters too. Luna: And that's why you see companies like Snowflake and ServiceNow - which are basically AI inference platforms - up this week. They benefit from cheaper inference because it improves their margins. Lucas: Right.

Snowflake up nearly ten percent. They just launched their Cortex AI inference service in March, and if inference costs come down, their gross margins could expand faster than expected. Luna: So the winners in this shift are the software platforms that monetize inference, and the inference chip startups - if they can gain traction. Lucas: And the big question is whether NVIDIA can maintain pricing power in inference.

Their data-center revenue last quarter was still overwhelmingly training - probably eighty percent or more. If that flips over the next two years, their revenue mix changes. Luna: Oracle dropped fifteen percent this week - is that related? They're a big buyer of NVIDIA chips for their cloud.

Lucas: It could be. Oracle's cloud is built on NVIDIA GPU clusters, and if the narrative shifts to inference, Oracle's competitive positioning - buying the most expensive GPUs - looks less attractive versus hyperscalers with custom silicon. Luna: So the takeaway: the AI hardware story is not dead, but it's rotating. The easy money in just buying NVIDIA and AMD for the training buildout might be done.

Now you have to understand who wins in inference. Lucas: And that's a more nuanced bet. It involves understanding software moats, hyperscaler strategy, and the unit economics of chip design. Not as simple as 'buy the pickax seller.'

Luna: Speaking of understanding - if today's conversation gave you something useful to think about, that's exactly why we keep this show ad-free. Listener support is what makes that possible. Lucas: Yeah, it's a small thing that makes a big difference for us. If you'd like to help keep it going, you can find us at buy me a coffee dot com slash fexingo.

Luna: No pressure, genuinely. We're just glad you're here. Lucas: Alright, back to the shift. One area I want to dig into is how this inference shift affects the startup ecosystem.

Because cheaper inference doesn't just help the Snowflakes of the world - it enables entirely new product categories. Luna: Like what? Give me a concrete example. Lucas: Real-time AI agents.

Think of a customer support bot that can reason through a complex refund policy while the customer is on hold. That requires low-latency, high-throughput inference. At today's costs on NVIDIA hardware, it's borderline uneconomical at scale. With inference-specialized chips, it gets viable.

Luna: So the shift from training to inference could unlock the next wave of AI applications that we've been waiting for - the ones that actually change how businesses operate day-to-day. Lucas: Exactly. The training phase was about building the models. The inference phase is about deploying them into workflows.

And that's where the real productivity gains - and the real investment opportunities - will emerge. Luna: I think that's a great note to end on. Thanks, Lucas. Lucas: Thanks, Luna.

Talk to you next time.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • How Idempotency-Key Design Prevents Payment DisastersThe Developer Tools Podcast with Fexingo · features Luna98 / 100
  • How Kubernetes Topology Spread Constraints Create Scheduling HotspotsDevOps Daily with Fexingo · features Luna95 / 100
  • Why Pipeline Velocity Trumps Deal Size Every TimeThe Growth Operator with Fexingo · features Luna95 / 100
  • Why Enterprise Software Deals Now Include a Vendor AI Model Explainability MandateB2B SaaS Talks with Fexingo · features Luna94 / 100
  • How B2B Brands Wreck Pipeline with Unsyncroned CRM DataThe Marketing Operator Podcast with Fexingo · features Luna92 / 100
  • Why Marketing Attribution Misses the Seasonality PatternMarketing Analytics with Fexingo · features Luna91 / 100

More from ChatGPT and Beyond with Fexingo

All episodes →
  • Why Palantir Stock Is Surging While AI Peers Decline72 / 100
  • How AI Model Marketplaces Are Becoming the New Operating Systems68 / 100
  • Why AI Model Marketplaces Are Becoming the New Operating Systems70 / 100
  • Neoclouds Are the New AI Infrastructure Battlefield72 / 100
  • How AI Chips Are Shifting From Training to Inference76 / 100
Explore the best B2B AI & Data podcasts →
All ChatGPT and Beyond with Fexingo episodes →