The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/The CTO Podcast with Fexingo
The CTO Podcast with Fexingo artwork

How Postmortems at Google Cut Downtime 50 Percent

The CTO Podcast with Fexingo · 2026-07-27 · 8 min

0:00--:--

Topics in this episode

Root cause analysisSite reliability engineeringIncident managementgoogle postmortemsblameless culture

Episode notes

This episode of The CTO Podcast dives into Google's incident postmortem culture. Lucas and Luna explore how the tech giant's blameless postmortems, systematic root cause analysis, and actionable follow-ups have cut repeat incidents by 50 percent. They discuss the cultural shift needed to move from blame to learning, the role of leadership in modeling transparency, and how small engineering teams can adopt similar practices. Drawing on internal studies and decades of incident data, the hosts unpack why postmortems are more than just a retrospective - they're a strategic tool for reliability. The conversation also touches on the balance between speed and thoroughness, common pitfalls like scope creep, and how to turn postmortems into org-wide process improvements. Perfect for CTOs, engineering managers, and anyone responsible for system reliability at scale.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Continual Improvement, Nonconformities, and Corrective Actions | Interview with Carlos CruzSecure & Simple · on Root cause analysis85 / 100
  • DevOps and observabilityNext in Tech · on Root cause analysis79 / 100
  • Systems Over Signs: How James Ferrell is Engineering Out Workplace HazardsThe Canary Report: Safety & Risk Management · on Root cause analysis79 / 100
  • Make Mistakes Matter: Turning Setbacks into Growth with Michael EhlersInnovation and the Digital Enterprise · on Site reliability engineering79 / 100
  • CX Superheroes podcast - Series 15 Episode 1 - Leading CX at Sclae - Tina LiljeCustomer Experience Superheroes · on Root cause analysis79 / 100
  • Why Testers Are Safe Despite AI Hype - Mitko MitevSoftware Testing Unleashed · on Root cause analysis76 / 100

More from The CTO Podcast with Fexingo

All episodes →
  • How Airbnb Rebuilt Its Search for 100 Million Listings72 / 100
  • How GitHub Migrated 100 Million Repositories to a New Storage Engine76 / 100
  • How Dropbox Rebuilt Its Sync Engine for 700 Million Users85 / 100
  • How Skyscanner Migrated 300 Microservices to Event-Driven Architecture85 / 100
  • How Palantir Rebuilt Its Foundry Ontology for Government AI Deployments85 / 100
Explore the best B2B Engineering & DevTools podcasts →
All The CTO Podcast with Fexingo episodes →