DevOps Daily with Fexingo · 2026-08-08 · 12 min
In this episode of DevOps Daily, Lucas and Luna dig into a surprisingly common but often overlooked failure mode in Kubernetes: container runtime interface timeouts. When the kubelet's connection to the container runtime - typically containerd or CRI-O - goes stale or hangs, node drains can stall, pods get stuck in Terminating, and your carefully planned maintenance window turns into a fire drill. Lucas walks through a real-world scenario where a single slow CRI call blocked an entire node drain, leading to failed deployments and user-facing outages. They discuss why timeouts happen, how to spot them in kubelet logs, and what you can do to prevent them - from tuning the kubelet's node-status-update-frequency to setting proper runtime-level timeouts. If you've ever seen a pod stuck in Terminating with no obvious cause, this episode will give you the diagnostic path and the fix. #Kubernetes #CRITimeout #NodeDrain #ContainerRuntime #Containerd #DevOps #SiteReliabilityEngineering #Kubelet #PodTermination #CloudNative #InfrastructureAsCode #TechPodcast #BusinessPodcast #FexingoBusiness #BusinessPodcast #TechOps #CI #CD Keep every episode free: buymeacoffee.com/fexingo
Other episodes covering the same guests and topics, from across The B2B Podcast Index.