Latent Space: The AI Engineer Podcast · Latent.Space

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

·18 min·1 clip
In CodeClash, AI models maintain their own code bases, edit them each round, and then face off in an arena to see whose code performs better.
John Yang, creator of the influential SweetBench coding benchmark, discusses the current state of AI coding evaluations. He covers SweetBench's evolution, the new competitive programming tournament benchmark CodeClash, and debates the future direction of AI agents between long-term autonomy and fast human-in-the-loop interaction.
Listen to the show on