Latent Space: The AI Engineer Podcast · Latent.Space

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay

·1 hr 32 min·9 clips
Yi Tay explains why Google DeepMind threw away AlphaProof to bet everything on Gemini for the IMO gold medal.
Yi Tay, a senior research scientist at Google DeepMind, rejoins the podcast to discuss his work on reasoning and reinforcement learning. The conversation centers on Google's Gemini project and the recent achievement of an International Math Olympiad gold medal. Tay explains his return to Google and his shift toward on-policy reinforcement learning research. He describes on-policy RL as training a model on its own generated outputs, which he analogizes to human experiential learning. The team decided to pursue the IMO gold using an end-to-end Gemini model, abandoning a previous specialized system called AlphaProof. The live competition required team members in Australia to run inference on problems as they were released. Tay highlights the collaborative, round-the-clock effort between captains in different time zones to train the final model checkpoint. A surprising insight is that Tay, with no IMO background, contributed to a gold-medal system he couldn't personally solve. He argues that most specialized tools will eventually be subsumed into a single model's parameters. The discussion touches on using benchmarks like Pokémon gameplay to test long-horizon planning and knowledge application. Tay distinguishes between applying existing knowledge and generating novel discoveries, which remains an open challenge. He notes that the definition of "reasoning" in AI has broadened to include any post-training that elicits capabilities through thinking traces. The conversation explores whether a model's internal "latent" reasoning must mirror human-like chain-of-thought. Tay is skeptical, believing model thoughts don't need to align with human cognition. The tone is conversational and technical, blending personal anecdote with deep research insights. Listeners interested in the practical challenges of frontier AI research and the philosophy of machine learning will enjoy this episode. Those seeking a highly structured debate or an introduction to basic AI concepts might find it less accessible.

As heard by us

A grounded look at AI's growing practical reach in research, coding, and image work.

This episode offers a brisk check-in on where AI is starting to feel genuinely useful. It moves through models turning messy spreadsheets into plots, AI coding becoming part of the routine, and image tools reaching a point where they are no longer just novelty.

Read the full review in PlayNext →

Why you'd press play

You want the useful edge of AI, not the hype reel.

Read the full recommendation in PlayNext →
Listen to the show on