
How I AI · 2026-06-15 · 40 min
In this episode, I sit down with Ankur Goyal , founder and CEO of Braintrust, the AI evals and observability platform used by teams like Notion, Stripe, Vercel, and Zapier. This one is for the senior engineers, staff engineers, VPs of engineering, and CTOs in my audience. We get into how coding agents can take on deeply technical architecture and infrastructure work that no single human engineer could tackle before, and then we demystify evals so you can use them to make your AI products better without touching the implementation. What you’ll learn: How Ankur uses Codex to run week-long benchmark experiments across database indexes, column store formats, and execution engines to speed up slow queries Why he argues there’s no excuse to skip rigorous benchmarking now that agents can run them tirelessly The “agent line” framework: how to decide which decisions, directions, and interactions you can hand off to an agent How I think about the practical vs.