
AI + a16z · 2025-05-30 · 1h 42m
LMArena cofounders Anastasios N. Angelopoulos, Wei-Lin Chiang, and Ion Stoica sit down with a16z general partner Anjney Midha to talk about the future of AI evaluation. As benchmarks struggle to keep up with the pace of real-world deployment, LMArena is reframing the problem: what if the best way to test AI models is to put them in front of millions of users and let them vote? The team discusses how Arena evolved from a research side project into a key part of the AI stack, why fresh and subjective data is crucial for reliability, and what it means to build a CI/CD pipeline for large models. They also explore: Why expert-only benchmarks are no longer enough. How user preferences reveal model capabilities - and their limits. What it takes to build personalized leaderboards and evaluation SDKs. Why real-time testing is foundational for mission-critical AI. Follow everyone on X: Anastasios N.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.