Latent Space: The AI Engineer Podcast · Latent.Space

[LIVE] Anthropic Distillation & How Models Cheat (SWE-Bench Dead) | Nathan Lambert & Sebastian Raschka

·52 min·5 clips
Anthropic accuses Chinese labs of distillation attacks, revealing a hidden AI geopolitics battle over model training.
This episode of Latent Space features hosts Nathan Lambert and Sebastian Raschka discussing AI model distillation and benchmark reliability. They are joined by guest Sean Swix, a writer joining the SAIL Coalition, who contributes to the technical conversation. Anthropic published a blog post detailing "distillation attacks" from prominent Chinese AI labs, specifically naming DeepSeek and MiniMax. The post describes how these labs used distributed accounts to generate synthetic data from Claude's API to train their own models, violating Anthropic's terms of service. Nathan Lambert notes the Chinese labs face a GPU shortage, making API usage for synthetic data generation an attractive strategy. Sebastian Raschka explains distillation as training a smaller model on the outputs of a larger one, a common practice for creating smaller model variants. The discussion covers how companies might detect distillation by analyzing patterns in API usage volume and question distribution. Lambert mentions Anthropic previously blocked U.S. companies, including xAI, from using its models for similar reasons. A surprising insight is the difficulty in distinguishing between legitimate large-scale benchmark evaluation and data collection for distillation. The hosts speculate that detection likely relies on analyzing the scale and repetitiveness of API requests. They find it notable that Anthropic singled out DeepSeek, a well-known Chinese lab in the U.S., within a geopolitical framing. The conversation shifts to SWE-bench, a coding benchmark where models fix bugs in open-source software, which OpenAI curated into a 500-task "verified" subset. The hosts reveal that OpenAI's recent audit found 59% of these verified tasks were unsolvable due to flawed test cases. One memorable example is a task requiring a model to output a specific magic string like "get_annotation" to pass, which is effectively a memorization test. The hosts discuss how models like GPT-4o can cheat on such benchmarks by using knowledge of future API versions seen in their training data. They estimate OpenAI spent millions of dollars curating and auditing SWE-bench Verified, only to find fundamental flaws. The episode has a conversational and exploratory tone, moving between technical definitions and broader industry analysis. The style is informal and opinionated, with hosts sharing personal experiences using APIs and conducting benchmarks. This episode would appeal to AI engineers and researchers interested in model training practices, benchmark design, and AI industry geopolitics. Listeners seeking a highly structured, single-topic deep dive might find the free-flowing discussion less engaging.
Listen to the show on