ChatGPT and Beyond with Fexingo · 2026-07-09 · 6 min
Episode 100 of ChatGPT and Beyond with Fexingo marks a milestone by exploring the quiet structural split happening inside today's large language models: the separation of reasoning from knowledge retrieval. Lucas and Luna break down why companies like NVIDIA and ARM are seeing divergent stock moves tied to this architectural shift, and how a new generation of 'hybrid RAG' systems is changing enterprise AI deployment. They discuss a specific case - a mid-sized logistics firm that cut inference costs by 40 percent by separating its model's reasoning and retrieval layers - and what that means for the broader AI stack. With live market data from July 9, 2026, the episode connects the dots between falling model costs, rising inference demand, and the companies positioned to win in a world where intelligence and memory are no longer bundled together. #AI #LargeLanguageModels #Reasoning #Retrieval #RAG #EnterpriseAI #NVIDIA #ARM #SMCI #AIInference #ModelArchitecture #HybridAI #AICosts #TechInvesting #GenerativeAI #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo
Other episodes covering the same guests and topics, from across The B2B Podcast Index.