
The FIT4Privacy Podcast · 2026-06-04 · 31 min
How do you know if an AI system is trustworthy, compliant, ethical, and fit for purpose ? In this episode of the FIT4Privacy Podcast, Punit Bhatia is joined by Stella Liu, an AI evaluation expert and founder of AI Evals & Analytics, to unpack one of the most practical and overlooked challenges in AI today: how to evaluate AI systems before and after deployment. KEY MOMENTS 02:09 - AI Definition 03:02 - AI Evaluations 10:31 - Why AI Testing Is Hard 14:06 - Evals Plus Analytics 18:15 - Synthetic Data 23:47 - Protecting Privacy Ethical 29:05 - AI Evals as a Company 29:52 - How to reach Stella Liu Stella explains why AI behaves differently from traditional software, why testing code alone is no longer enough, and how AI evaluations (AI evals) help organizations assess real-world behavior, risk, and performance. From evaluation driven development to continuous monitoring in production, the conversation explores how teams can move beyond guesswork and hype toward repeatable, measurable AI governance.