
First Commit · 2026-02-12 · 37 min
This week, we’re joined by Anish Agarwal , CEO of Traversal , an AI-native site reliability platform helping teams detect, diagnose, and remediate incidents before they spiral into prolonged downtime. Anish shares how Traversal is tackling the full lifecycle of reliability work: identifying what caused an incident, determining which signals actually mattered during alert triage, and helping teams plan how their infrastructure should evolve over time. As modern systems grow more complex, spanning microservices, serverless, and multi-cloud environments, the surface area for failure continues to expand, especially as AI-generated code accelerates change faster than humans can reasonably keep up. We talk about why observability has produced some of the largest outcomes in software, why the traditional dashboard-first model is breaking down, and why the next generation of tooling needs agents that can search, reason over, and act on observability data, not just display it. Anish also breaks down where “self-healing systems” are real today, where expectations need to be reset, and why many AI-SRE products risk building faster horses instead of rethinking the experience entirely.