DevOps Daily with Fexingo · 2026-08-24 · 9 min
In this episode of DevOps Daily, Lucas and Luna dig into a surprisingly common Kubernetes misstep: setting CPU and memory requests too high. Using a real-world example from a mid-sized fintech that saw its cloud bill jump 40 percent after a migration, they explain how requests and limits work, why overprovisioning sneaks in, and the hidden costs that follow. They walk through the mechanics of the scheduler, the role of Quality of Service classes, and why the Kubernetes vertical pod autoscaler isn't a cure-all. Lucas shares a practical approach to right-sizing requests using metrics like the 99th percentile and the concept of the request-to-limit ratio. Luna brings up the tension between performance risk and cost, and they close with a forward-looking question about whether the industry will ever settle on standard right-sizing tools. If you've ever wondered why your cluster feels overprovisioned or why your FinOps dashboard looks scary, this episode gives you a clear framework to fix it.