Latent Space: The AI Engineer Podcast · Latent.Space

[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton

·28 min·1 clip
A depth change plus residual connections suddenly makes performance skyrocket in one environment.
1. Latent Space: The AI Engineer Podcast covers Kevin Wang and Princeton’s NeurIPS Best Paper on 1,000-layer networks for self-supervised RL. 2. Kevin Wang is the lead author and recent Princeton undergraduate, and Ben taught the independent work seminar that helped start the project. 3. The episode asks whether reinforcement learning can scale more like language and vision, instead of staying limited to shallow networks. 4. Kevin says the project began in an IW seminar and grew through collaboration with Ishan, Nicole, Ben, and later the HALT group. 5. Ben says he started skeptical because previous attempts at deeper RL networks had not worked. 6. Kevin says the team took the bet because the infrastructure from the previous year made large experiments cheaper to run. 7. Kevin describes deep RL as the third branch of deep learning, after NLP and vision, where scaling to massive networks had already paid off. 8. He says traditional value-based RL does not scale well, so the team explored self-supervised RL instead. 9. The self-supervised setup pushes representations of states, actions, and future states together on the same trajectory and apart across different trajectories. 10. Kevin says the method can solve goal-reaching tasks without a human-crafted reward signal. 11. He says the first deep-network experiments failed because performance degraded when depth increased. 12. He says residual connections and layer norm were part of the combination that eventually produced large gains. 13. Ben says the strongest result came from the interaction between architecture and objective, not from depth alone. 14. Kevin says width scaling helped, but depth scaling was more parameter-efficient because width grows roughly quadratically in parameters. 15. The team says one environment showed a dramatic jump when depth crossed a critical point with the right recipe. 16. They say many tasks reached near-perfect performance at around 64 layers, not necessarily 1,000. 17. The conversation stays interview-driven and technical, with back-and-forth questions about scaling, ablations, and tradeoffs. 18. The tone is analytical and collaborative, with repeated references to poster-session questions, JAX GCRL, and robotics. 19. Best for listeners interested in RL scaling, robotics, and deep learning architecture. 20. Probably skip if you want narrative storytelling or nontechnical discussion.
Listen to the show on