Latent Space: The AI Engineer Podcast · Latent.Space

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena

·24 min·1 clip
He responds to the 'leaderboard illusion' paper and says it contained 'factual mistakes' and 'blatantly unscientific and false' claims.
1. Latent Space: The AI Engineer Podcast focuses on LMArena's $1.7B vision with Anastasios Angelopoulos, who discusses the company's leaderboard, user base, and evaluation strategy. 2. Anastasios Angelopoulos appears as LMArena's co-founder and leader, and the hosts frame him as the person explaining how the platform scaled from Berkeley roots into a company. 3. The episode asks why LMArena became a startup, how it uses its capital, and how it keeps a leaderboard credible while serving real users. 4. He says the company began as LMSYS at Berkeley and kept "LM" at first because it wanted to broaden beyond the original language-model research group. 5. He credits Ange for incubating the company, forming the entity, giving grants, and supporting the team before they committed to starting a business. 6. He says the main reason to incorporate was that building a company was the only way to scale the platform and its mission around real-world AI usage. 7. He describes the platform as a way to "measure, understand and advance the frontier AI capabilities" using "real world users" and "organic feedback." 8. He says LMArena raised about $100 million and frames the money as extra options, not an obligation to spend everything. 9. He says the platform is expensive because LMArena funds all inference, pays standard enterprise rates, and supports free usage. 10. He gives scale figures including more than 5 million users and mid tens of millions of conversations per month, and says about 250 million conversations have happened on the platform. 11. He says roughly 25% of users do software for a living and that half of users are now logged in, which lets LMArena learn more about who is using the product. 12. He contrasts LMArena with Artificial Analysis, saying their arenas use public benchmarks and pre-generated video, while LMArena uses users' own prompts and use cases. 13. He says that difference matters because organic input gives LMArena more realism, and he notes that their video arena would not require users to wait for pre-generated content. 14. He says moving off Gradio and onto React was necessary for performance, hiring, and custom features like loading icons and video notifications. 15. He responds to the 'leaderboard illusion' paper by saying LMArena published a response online and that the paper had factual errors about pre-release testing and model sampling. 16. He says LMArena is supportive of open-source models and cites a more like 60/40 split rather than the paper's claim of 9% open-source and 60% closed-source. 17. He says the public leaderboard is a loss leader and a "charity," because models cannot pay to enter, improve, or remove themselves from it. 18. He says LMArena will keep expanding into occupational categories like medicine, legal, business, finance, accounting, creative, and marketing, while also moving toward multimodal and video evaluation. 19. The tone is technical and defensive in places, but also collaborative, with long explanations and back-and-forth from the hosts. 20. Listeners who follow AI evaluation, model benchmarks, and startup scaling will get the most from it, while people avoiding benchmark disputes may skip it.
Listen to the show on