DevOps Daily with Fexingo · 2026-07-12 · 8 min
Episode 107 dives into a common misconfiguration: setting CPU limits on Kubernetes pods that unintentionally cap the Horizontal Pod Autoscaler's ability to scale. Lucas and Luna explore how the relationship between resource requests and limits affects the HPA algorithm, using a real-world example from a FinTech startup that saw scaling delays of over 40 seconds. They explain why the default CPU limit behavior can lead to throttling, how the HPA reads metrics from the `/metrics/resource/v1alpha1` endpoint, and why setting CPU limits too low can create a feedback loop where pods are throttled while the HPA waits for utilization to rise. The episode offers a concrete fix: using `targetCPUUtilizationPercentage` with appropriate request-to-limit ratios, and when to use `averageValue` instead of `averageUtilization`. No fluff, just one actionable lesson. #Kubernetes #HPA #HorizontalPodAutoscaler #CPUThrottling #K8sScaling #DevOps #FinTech #ResourceManagement #CloudNative #CI/CD #PodAutoscaling #KubernetesLimits #ProductionTroubleshooting #SiteReliability #FexingoBusiness #BusinessPodcast #TechOps #ContainerOrchestration Keep every episode free: buymeacoffee.com/fexingo