The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/DevOps Daily with Fexingo
DevOps Daily with Fexingo artwork

How Kubernetes Pod Disruption Budgets Cause Cascading Failures

DevOps Daily with Fexingo · 2026-07-12 · 7 min

0:00--:--

Topics in this episode

kubernetes pod disruption budgetpdb cascading failurespot instance migration kubernetescluster autoscaler pdbnode drain blocked pdb

Episode notes

In this episode, Lucas and Luna dive into a counterintuitive Kubernetes failure mode: correctly configured PodDisruptionBudgets that actually create cascading availability failures. Using a real-world example from a large e-commerce platform migrating to spot instances in July 2026, they explain how PDBs interact with cluster autoscaler taint-based evictions and node drain logic to deadlock workloads. Lucas breaks down the specific scenario - a 5-replica deployment with maxUnavailable=1 that should be safe but stalls an entire node pool - and walks through the underlying mechanism: how Kubernetes honors PDBs during voluntary disruptions but blocks evictions when nodes are cordoned, leading to hung pods and scaling failures. Luna brings data from a recent CNCF survey showing 40% of teams using spot instances have encountered this exact pattern. They cover mitigation strategies including pod-level priority, descheduler policies, and PDB monitoring, and discuss why this becomes more acute as GPU spot instances gain traction. Practical, specific, and immediately actionable for any SRE running Kubernetes at scale.

More from DevOps Daily with Fexingo

All episodes →
  • How Kubernetes ServiceAccount Token Expiration Breaks CI Workflows75 / 100
  • How Kubernetes StatefulSet PVC Resizing Causes Node Disk Failures94 / 100
  • How Kubernetes Topology Spread Constraints Create Scheduling Hotspots95 / 100
  • How Kubernetes CRD Versioning Breaks Controller Upgrades90 / 100
  • How Kubernetes Audit Logging Causes etcd Performance Degradation91 / 100
Explore the best B2B Engineering & DevTools podcasts →
All DevOps Daily with Fexingo episodes →