DevOps Daily with Fexingo · 2026-07-13 · 7 min
When a single misbehaving mutating admission webhook in Kubernetes can take down your entire cluster, it's not just a developer headache - it's a production emergency. In this episode, Lucas and Luna break down the real-world cascade failure at a major fintech that lost $2.3 million in trading revenue over 47 minutes. They walk through how webhooks work, why a simple config map change triggered a chain reaction across 1,200 pods, and what the team did to recover - and prevent it from happening again. You'll learn about webhook timeout defaults, failure policy nuances, and why monitoring webhook latency matters more than you think. If you run Kubernetes in production, this episode might save your weekend. #Kubernetes #MutatingAdmissionWebhook #CascadeFailure #K8sProduction #WebhookLatency #AdmissionController #CNCF #SiteReliabilityEngineering #IncidentResponse #CloudNative #K8sBestPractices #Fintech #PodSecurity #DevOps #Technology #FexingoBusiness #BusinessPodcast #Infrastructure Keep every episode free: buymeacoffee.com/fexingo
Other episodes covering the same guests and topics, from across The B2B Podcast Index.