Latent Space: The AI Engineer Podcast · Latent.Space

⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data

·26 min·2 clips
Mia found over half of C-Bench Verified problems had unfair tests, making model failures misleading.
OpenAI researchers Mia Glaese and Olivia discuss why C-Bench Verified, a key coding benchmark, has become saturated and contaminated, leading to unreliable measurements. They advocate for moving to harder benchmarks like C-Bench Pro and explore future directions for evaluating AI coding capabilities, including real-world impact and OpenAI's preparedness framework.
Listen to the show on