The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/The CTO Podcast with Fexingo
The CTO Podcast with Fexingo artwork

How Anthropic Cut Inference Costs by 40 Percent

The CTO Podcast with Fexingo · 2026-08-29 · 8 min

0:00--:--

Topics in this episode

Speculative decodinganthropic inference cost reductionclaude ai model optimizationcut llm inference coststransformer architecture optimization

Episode notes

In this episode of The CTO Podcast, Lucas and Luna dig into how Anthropic, the company behind Claude, managed to cut inference costs by roughly 40 percent over the past year. They explore the technical levers: from optimizing transformer architectures and using speculative decoding to smarter batching and model distillation. They also discuss the trade-offs between cost, latency, and quality, and what this means for other AI teams trying to run large models economically. If you're building on top of large language models or just curious about the economics of AI, this episode gives you a concrete look at how one of the industry's leading labs is tackling the cost problem. #Anthropic #ClaudeAI #InferenceCosts #AIEconomics #LLMOptimization #SpeculativeDecoding #TransformerArchitecture #ModelDistillation #AIInfrastructure #CTO #TechLeadership #Engineering #BusinessAndTechnology #FexingoBusiness #BusinessPodcast #AIPodcast #TechPodcast #MachineLearning Keep every episode free: buymeacoffee.com/fexingo

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysisTraining Data · on Speculative decoding95 / 100

More from The CTO Podcast with Fexingo

All episodes →
  • How Datadog Scaled Engineering Without Burning Out82 / 100
  • How Slack Scaled Engineering Without Adding Headcount
  • How Atlassian Tamed Technical Debt With Architecture Contracts
  • How Atlassian Tamed Technical Debt With Architecture Contracts
  • How GitHub Tamed Merge Conflict Chaos
Explore the best B2B Engineering & DevTools podcasts →
All The CTO Podcast with Fexingo episodes →