The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/cloud2030
cloud2030 artwork

Kubecon SC25 Debrief

cloud2030 · 2026-06-12 · 43 min

0:00--:--

Episode notes

In this episode, we debrief several industry events I went to last year, including Supercomputing, KubeCon, Stack, the AI Infrastructure Show, and the Red Hat AI Infrastructure Summit. We dive deep into some observations from the shows and what they tell us about the gaps and fractures in how we are working to build AI infrastructure. We focus on how observability is being used for evaluation, tuning, performance issues, GPU dropouts, and cluster management, while anomaly detection and root cause analysis remain less common, and we note that networking is still underserved. We also get into the shift from building clusters to observing and fixing them after deployment, especially for agentic systems, and we end by highlighting the need for observability across application, identity, networking, and infrastructure layers. Transcript:

More from cloud2030

All episodes →
  • AI UX Building60 / 100
  • Vibe Coding for Ops Project Restart [TechOps]
  • Kubernetes as Common Platform
  • AWS Outage
  • Back After a Break
Explore the best B2B Engineering & DevTools podcasts →
All cloud2030 episodes →