The CTO Podcast with Fexingo · 2026-07-28 · 5 min
In July 2024, a cascading failure in Datadog's own monitoring pipeline took down their dashboard for 47 minutes. That event triggered a top-to-bottom redesign of their incident command system. In this episode, we break down the three key changes: the new Incident Commander role, mandatory game-day drills, and a post-incident review process that cut mean time to resolve by 40 percent. We also explore what Datadog's CTO learned about balancing feature velocity with reliability, and how these practices have influenced the broader observability industry. If you're leading an engineering organization, this episode offers a concrete playbook for improving incident response without sacrificing speed. #Datadog #IncidentResponse #CTO #EngineeringLeadership #IncidentCommand #Postmortem #Reliability #MTTR #GameDay #SRE #DevOps #Observability #CloudInfrastructure #TechnicalLeadership #EngineeringCulture #SiteReliability #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo
Other episodes covering the same guests and topics, from across The B2B Podcast Index.