ChatGPT and Beyond with Fexingo · 2026-07-01 · 8 min
Key moments - from our scoring
Substance score
56 / 100
Five dimensions, 20 points each
The AI hardware market is undergoing a fundamental restructuring as inference workloads - running trained models in production - now dominate compute spending. Lucas and Luna trace how inference has grown from 40% to 60% of AI compute cycles year-over-year, driven by the proliferation of AI features across applications. This shift favors AMD's MI300 series and custom inference accelerators from companies like Groq (their language processing units) over NVIDIA's training-optimized GPUs, while creating headwinds for Super Micro Computer's liquid-cooled training clusters. Beyond hardware, the inference boom is reshaping cloud pricing - AWS, Azure, and GCP have moved to per-token billing models - and creating opportunities for orchestration platforms like Kubernetes and serverless inference providers (RunPod, Banana). Application-layer inference plays like Snowflake's AI-powered data pipelines and ServiceNow's ITSM copilot are capturing significant value. For builders, the key insight is that inference costs are falling per unit but exploding in volume, making inference optimization critical to product economics.
AMD's MI300 series are gaining traction for inference workloads where cost-efficiency matters, while Super Micro's liquid-cooled racks are optimized for dense training clusters; as spending shifts from training to inference, SMCI's product mix becomes less relevant.
Inference accounts for roughly 60% of AI compute cycles as of Q1 2026, up from about 40% a year prior, representing a massive shift in hardware demand away from training-optimized chips.
Inference workloads are highly variable - bursting to 10,000 tokens then dropping to a trickle - so per-token billing lets customers pay for actual usage rather than idle GPU time and forces cloud providers to adopt flexible capacity allocation.
Inference chips prioritize low latency and high throughput per watt (to handle variable, continuous queries), while training hardware needs massive memory bandwidth and matrix math performance for large batches of data during model optimization.
Both are selling inference as an application-layer runtime - Snowflake runs models on data warehouses, ServiceNow powers ITSM copilots - capturing value by monetizing inference execution rather than just the hardware underneath.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers several substantive claims about the infrastructure shift from training to inference, with specific data points (60% of compute cycles, 40% YoY price drops, stock movements). However, it relies heavily on inference from market movements rather than deep operational insight, and some arguments are asserted rather than unpacked - e.g., why inference requires less dense infrastructure or how per-token pricing actually changes capacity management.
inference now accounts for roughly 60 percent of AI compute cycles, up from about 40 percent a year ago
inference prices have fallen something like 40 percent year-over-year. But the volume is exploding.
The core observation - that inference is becoming the dominant AI workload - is increasingly mainstream by mid-2026 and not particularly contrarian. The split between training (high-end, low-volume) and inference (high-volume, cost-sensitive) is a useful framework but not novel. The edge cases (Apple M4, on-device inference, Groq LPUs) add some freshness, but the overall thesis lacks genuinely counterintuitive insight.
Training is the lab; inference is the factory floor.
AMD, and even companies like Broadcom with custom ASICs, are making inroads
This is a two-person co-hosted show with no external guest. Lucas and Luna appear to be analysts or commentators speaking from market observation and published reports rather than operators or practitioners who have built or deployed inference infrastructure at scale. Their credibility rests on synthesis, not lived experience in the space.
A report from a major cloud provider last quarter showed that inference now accounts for roughly 60 percent of AI compute cycles
We've talked on this show before about model costs crashing
The episode names specific companies (AMD, NVIDIA, Super Micro, Snowflake, ServiceNow, Groq, Apple M4, Qualcomm Snapdragon, RunPod, Banana) and provides some quantitative anchors (60% inference cycles, 40% YoY price decline, stock movements like AMD +12%, SMCI -10%). However, most claims about product capabilities and market dynamics are stated without concrete customer examples, deployment metrics, or financial data to back them up.
AMD is up nearly 12 percent in five days, while Super Micro Computer is down almost 10 percent
Their MI300 series accelerators have been gaining traction specifically for inference workloads
The dialogue is smooth and well-structured, with Luna and Lucas building on each other's points. However, there is little genuine pushback, skepticism, or probing. Both hosts agree throughout; there are no challenging follow-ups that test assertions (e.g., "How confident are you that SMCI's drop is really about cooling, not other factors?"). The conversation reads more as parallel exposition than rigorous interrogation.
That's a great point.
Exactly.
Computed from the transcript - who did the talking, and the words that came up most.
Episode 84 of ChatGPT and Beyond with Fexingo explores the massive shift in AI hardware spending from training to inference. Lucas and Luna break down why AMD surged 11.8% in the last five days while Super Micro Computer dropped 9.6%, and what that tells us about the real bottleneck in AI right now. They dig into new data showing that inference workloads now consume more than half of all AI compute, and why that's changing everything from chip design to cloud pricing. Plus, a look at how companies like Snowflake and ServiceNow are riding the inference wave. No fluff, just clear analysis on one concrete market shift. #AIHardware #Inference #Training #AMD #NVDA #SMCI #AVGO #SNOW #NOW #CloudComputing #ChipDesign #AISpending #Technology #FexingoBusiness #BusinessPodcast #TechTrends #ArtificialIntelligence #Semiconductors Keep every episode free: buymeacoffee.com/fexingo
Transcribed and scored by The B2B Podcast Index.
Lucas: So it's July 1st, 2026, and if you've glanced at chip stocks this week, you've probably noticed something strange: AMD is up nearly 12 percent in five days, while Super Micro Computer is down almost 10 percent. That divergence tells a pretty clear story about where AI hardware spending is actually going. Luna: It feels like the market is finally waking up to the fact that training AI models isn't the only game in town anymore. Inference - actually running those models in production - is where the real volume is.
Lucas: Exactly. And the numbers back that up. A report from a major cloud provider last quarter showed that inference now accounts for roughly 60 percent of AI compute cycles, up from about 40 percent a year ago. That's a massive flip.
Luna: And it explains why AMD is suddenly hot. Their MI300 series accelerators have been gaining traction specifically for inference workloads, where they offer competitive performance per dollar against NVIDIA's H100 and B200. Lucas: Right. NVIDIA still dominates training - their CUDA ecosystem is a moat - but inference is a different game.
It's more about latency, throughput, and cost efficiency at scale. That's where AMD, and even companies like Broadcom with custom ASICs, are making inroads. Luna: And Super Micro's drop? That seems tied to their heavy exposure to liquid-cooled racks for training clusters.
If demand is shifting toward inference servers, their product mix might be out of step. Lucas: That's exactly the read. SMCI rode the training build-out wave hard - their revenue tripled in two years. But inference infrastructure is often less dense, less power-hungry, and doesn't always need the same exotic cooling solutions.
So their growth narrative gets called into question. Luna: It's a reminder that in AI hardware, the tailwinds can flip fast. Look at Snowflake and ServiceNow - both up nicely this week, Snowflake up over 12 percent. They're not chip companies, but they're inference platforms.
They're selling the runtime. Lucas: That's a great point. Snowflake's recent push into ai powered data pipelines - running models on your data warehouse - is pure inference. Same with ServiceNow's ITSM copilot.
These companies are monetizing inference at the application layer. Luna: So what does this shift mean for someone building an AI product today? Should they be thinking more about inference costs upfront? Lucas: Absolutely.
We've talked on this show before about model costs crashing - inference prices have fallen something like 40 percent year-over-year. But the volume is exploding. So total inference spend for most companies is actually going up, not down. The unit cost drops, but usage grows faster.
Luna: Which is exactly why chipmakers are scrambling to optimize for inference. Lower cost per query means you can afford to run models on more data, more often. That virtuous cycle benefits the hardware guys who nail the inference price-performance curve. Lucas: And we're seeing new architectures emerge specifically for inference.
Apple's latest M4 chip, for example, has a neural engine that's really tuned for on-device inference. Qualcomm's Snapdragon X Elite is another. This isn't just data center stuff anymore. Luna: Edge inference is a whole other layer.
But even in the cloud, the design priorities are different. Training needs massive memory bandwidth and matrix math. Inference needs low latency and high throughput per watt. Lucas: And that's why you see companies like Groq building custom LPUs - language processing units - that are basically just inference engines.
They can't train models, but they can run a Llama 3 at blazing speed. And they're getting traction. Luna: It's almost like the AI hardware market is splitting into two distinct ecosystems. Training is the high-end, low-volume, cutting-edge race.
Inference is the high-volume, cost-sensitive, everywhere race. Lucas: Exactly. And the inference race is bigger in terms of total addressable market. Every app that integrates an AI feature - search, customer support, code generation - runs inference constantly.
Training happens intermittently. Luna: If today's conversation gave you something usable - a framework for thinking about AI hardware, or even just a stock idea to dig into - that's exactly why we do this show ad-free. And listener support is what keeps it that way. Lucas: Yeah, it's a small thing that makes a big difference.
If you're finding value here, you can support the show at buy me a coffee dot com slash fexingo. No pressure, but every bit helps us keep the conversation going. Luna: Alright, back to the hardware shift. One thing I find fascinating is how this inference boom is changing pricing models at cloud providers.
They're moving away from per-hour GPU pricing to per-token pricing. Lucas: Right. AWS, Azure, GCP - all of them now offer inference as a service with per-token billing. It's a direct response to the fact that inference workloads are much more variable than training.
You might need a burst of 10,000 tokens one minute and just a trickle the next. Luna: And that's actually better for customers, because you're not paying for idle GPU time. But it also means the cloud providers have to manage capacity differently. They need more flexible, spot market style allocation.
Lucas: Which plays into the hands of companies with good orchestration software. That's where something like Kubernetes with GPU sharing comes in. Or even startups like RunPod and Banana that focus purely on serverless inference. Luna: So the inference shift isn't just about hardware - it's reshaping the entire software stack around deployment.
And that creates opportunities for new companies, but also risks for incumbents who are slow to adapt. Lucas: Take Oracle, for example. Down 7 percent this week. They've been pushing hard into cloud AI, but their inference offerings aren't as mature as AWS or Azure.
The market might be penalizing that gap. Luna: Or it could be unrelated - there are a lot of factors moving stocks. But the trend is clear: inference is the new battleground. Lucas: And it's still early.
We're maybe two years into the inference-first era. The hardware roadmap for the next 18 months is packed with inference-optimized chips - from NVIDIA's Blackwell Ultra to AMD's next-gen CDNA to custom ASICs from every hyperscaler. Luna: It feels like the AI industry is growing up. Training is the lab; inference is the factory floor.
And the factory floor is where the real economic value gets produced. Lucas: Well said. And that factory floor is only getting bigger. Every chatbot, every code assistant, every automated customer service call - that's all inference.
The volume is going to keep compounding. Luna: So for the listener trying to figure out where to focus their attention - whether as an investor, a developer, or a business leader - the inference layer is probably the most important piece of the AI stack to understand right now. Lucas: Agreed. And we'll keep tracking it.
Next time, I want to dig into one specific inference company - maybe Groq or Cerebras - and look at their actual customer traction. But for today, I think the big takeaway is: the hardware spend is moving, and the chips that win inference will define the next phase of AI. Luna: Sounds like a plan. Thanks for listening, everyone.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.