The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/DevOps Daily with Fexingo
DevOps Daily with Fexingo artwork

How Kubernetes Node Autoscaler Fails Under Spot Instance Churn

DevOps Daily with Fexingo · 2026-07-04 · 9 min

0:00--:--

Topics in this episode

Kubernetescluster autoscalerSpot instancesaws ec2node scaling

Episode notes

Kubernetes cluster autoscaler is a workhorse for dynamic workloads, but when you're running spot (preemptible) instances, the node autoscaler's default behavior can actually hurt reliability. In this episode, Lucas and Luna break down a real incident from a large ad-tech platform running 2,000+ nodes on AWS EC2 Spot Instances. When spot interruption rates spiked during a regional AZ outage, the cluster autoscaler's rigid scale-down logic kept terminating nodes that were seconds away from stabilizing, causing a cascading re-scheduling loop that nearly took down their real-time bidding service. They walk through the specific configuration failures: the unbound scale-down-delay-after-add, the lack of node disruption budgets for spot pools, and the missing mixed-instance-policy fallback. Plus, they share the three YAML changes that fixed it - including custom priority expander logic and a proactive node-drain webhook. If you're running spot nodes in production, this is the episode that will save you from a 2 AM pager.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • From Compliance Theater to GRC Infrastructure: Why AI Breaks Traditional GRC ft Jasmine Kaur, Principal of Security & Assurance Engineering @ CoreWeaveSecurity & GRC Decoded · on Kubernetes96 / 100
  • Ship It Conversations: Jake Warner on Cycle.io, Bare Metal’s Comeback, and Why Private Cloud Is Getting Interesting AgainShip It Weekly · on Kubernetes91 / 100
  • Kubernetes and retiring at the top with Kelsey HightowerThe Pragmatic Engineer · on Kubernetes90 / 100
  • Episode 121: Art Degrees, Sun Microsystems, and How Kubernetes Scales Contributions, with Josh BerkusSoftware Defined Interviews · on Kubernetes89 / 100
  • Rebooting Enterprise AI with MCP and KubernetesPractical AI · on Kubernetes88 / 100
  • Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and KubernetesData Engineering Podcast · on Kubernetes88 / 100

More from DevOps Daily with Fexingo

All episodes →
  • How Kubernetes ServiceAccount Token Expiration Breaks CI Workflows75 / 100
  • How Kubernetes StatefulSet PVC Resizing Causes Node Disk Failures94 / 100
  • How Kubernetes Topology Spread Constraints Create Scheduling Hotspots95 / 100
  • How Kubernetes CRD Versioning Breaks Controller Upgrades90 / 100
  • How Kubernetes Audit Logging Causes etcd Performance Degradation91 / 100
Explore the best B2B Engineering & DevTools podcasts →
All DevOps Daily with Fexingo episodes →