The CTO Podcast with Fexingo · 2026-08-29 · 8 min
In this episode of The CTO Podcast, Lucas and Luna dig into how Anthropic, the company behind Claude, managed to cut inference costs by roughly 40 percent over the past year. They explore the technical levers: from optimizing transformer architectures and using speculative decoding to smarter batching and model distillation. They also discuss the trade-offs between cost, latency, and quality, and what this means for other AI teams trying to run large models economically. If you're building on top of large language models or just curious about the economics of AI, this episode gives you a concrete look at how one of the industry's leading labs is tackling the cost problem. #Anthropic #ClaudeAI #InferenceCosts #AIEconomics #LLMOptimization #SpeculativeDecoding #TransformerArchitecture #ModelDistillation #AIInfrastructure #CTO #TechLeadership #Engineering #BusinessAndTechnology #FexingoBusiness #BusinessPodcast #AIPodcast #TechPodcast #MachineLearning Keep every episode free: buymeacoffee.com/fexingo
Other episodes covering the same guests and topics, from across The B2B Podcast Index.