The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Data Brew by Databricks
Data Brew by Databricks artwork

SWE-bench & SWE-agent | Data Brew | Episode 44

Data Brew by Databricks · 2025-04-17 · 36 min

0:00--:--

Episode notes

In this episode, Kilian Lieret, Research Software Engineer, and Carlos Jimenez, Computer Science PhD Candidate at Princeton University, discuss SWE-bench and SWE-agent, two groundbreaking tools for evaluating and enhancing AI in software engineering. Highlights include: - SWE-bench: A benchmark for assessing AI models on real-world coding tasks. - Addressing data leakage concerns in GitHub-sourced benchmarks. - SWE-agent: An AI-driven system for navigating and solving coding challenges. - Overcoming agent limitations, such as getting stuck in loops. - The future of AI-powered code reviews and automation in software engineering.

More from Data Brew by Databricks

All episodes →
  • Reinforcement Fine-Tuning and the Future of Specialized AI Models76 / 100
  • Benchmarking Domain Intelligence | Data Brew | Episode 45
  • Enterprise AI: Research to Product | Data Brew | Episode 43
  • Multimodal AI | Data Brew | Episode 42
  • Age of Agents | Data Brew | Episode 41
Explore the best B2B AI & Data podcasts →
All Data Brew by Databricks episodes →