The CTO Podcast with Fexingo · 2026-07-30 · 6 min
Netflix deploys thousands of times daily, but a few years ago, deploy-related incidents were too common. Their solution? A sophisticated canary deployment system built on their internal tool Spinnaker and the open-source Kayenta project. In this episode, Lucas and Luna break down how Netflix uses automated canary analysis to detect rollouts that degrade performance or reliability before they affect many users. They walk through the architecture: how Spinnaker manages multi-cluster deployments, how Kayenta compares metrics from control and experiment groups using statistical tests like Mann-Whitney U, and how the system automatically rolls back bad releases. The result: a 90 percent reduction in deploy-related incidents, saving hours of engineering time and improving customer experience. The hosts discuss the key design decisions, including how to set appropriate thresholds, the importance of automated rollback, and how this approach fits into Netflix's broader chaos engineering culture. They also connect it to practices that any CTO can adopt, even at smaller scale.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.